Hyperspectral Image Classification Method Based on Agent Attention and VIT Architecture
By adopting a proxy attention and VIT architecture method in hyperspectral image classification, the problems of insufficient feature extraction capabilities and high computational complexity are solved, and an efficient and lightweight hyperspectral image classification model is realized.
Patent Information
- Application Number
- CN202510435730.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The prior art has problems of insufficient feature extraction capabilities and high computational complexity in hyperspectral image classification, especially in the development of lightweight models suitable for edge devices.
Using a hyperspectral image classification method based on proxy attention and Vision Transformer (VIT) architecture, feature extraction and fusion are performed through sliding window grouping encoding module and proxy attention encoder to reduce computational complexity and improve representation ability.
It realizes efficient hyperspectral image classification, significantly reduces the computational complexity, and maintains strong feature expression capabilities, which are suitable for edge devices.
Smart Images

Figure CN119942250B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image classification, and particularly to a hyperspectral image classification method based on proxy attention and VIT architecture. Background Art
[0002] Hyperspectral Remote Sensing Image (HIS), as an advanced remote sensing technology, has received extensive attention in recent years. HIS has both rich spectral information and spatial information at the same time. Each pixel contains hundreds of narrow spectral band information in the channel dimension. This unique data structure makes hyperspectral images show great application potential in many fields such as urban construction, ocean observation, precision agriculture, disaster prevention and reduction, and resource exploration.
[0003] Early hyperspectral image classification mainly relied on traditional machine learning methods such as Random-Forest algorithm and support vector machine. However, these methods have limitations in feature extraction ability and adaptability to complex scenarios. In recent years, deep learning technology, especially Convolutional Neural Network (CNN), has been widely used in hyperspectral image classification tasks due to its powerful feature learning ability. However, based on the data characteristics of HIS images containing hundreds of continuous and approximate spectral band information, the deep semantic features extracted by CNN models are often limited; and as the depth of the model increases, the computational cost of CNN models increases significantly. The model ViT (Vision Transformer) has brought new opportunities for HSI classification while achieving remarkable results in the field of computer vision for deep learning models Transformer. Compared with CNN using a sliding window mechanism to gradually achieve global information fusion, Transformer has the advantage of capturing global information at one time, which makes it more suitable for processing HSI data containing hundreds of continuous spectral bands. However, Transformer still faces some challenges when applied to HSI classification: in order to develop a lightweight hyperspectral classification model suitable for edge devices, it is necessary to reduce the number of model parameters and computational complexity, and how to achieve a good balance between computational efficiency and representation ability has become a difficult task. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides a hyperspectral image classification method based on proxy attention and VIT architecture with simple algorithm and strong representation ability.
[0005] The technical solution of the present invention to solve the above technical problems is: a hyperspectral image classification method based on proxy attention and VIT architecture, comprising the following steps:
[0006] S1, Data preprocessing: Preprocess the hyperspectral image data and divide it into a training set and a test set;
[0007] S2, Model establishment: Construct a hyperspectral image classification model based on surrogate attention and the VIT architecture;
[0008] The specific process of step S2 is as follows:
[0009] Step S21: Construct a sliding window grouping encoding module based on the sliding window mechanism and surrogate pixel blocks, and encode the original hyperspectral data into an initial feature encoding matrix , which contains the spatial-spectral joint representation of the hyperspectral data;
[0010] Step S22: Construct a surrogate attention encoder ATE based on surrogate attention Agent Attention, and perform hierarchical feature extraction through L cascaded surrogate attention encoders ATE to realize the layer-by-layer fusion and enhancement of the spatial-spectral information in the initial feature encoding matrix to obtain a feature map;
[0011] Step S23: Construct a classification head composed of an average pooling layer and a multi-layer perceptron MLP layer, interpret the feature map, predict the category to which the original pixel belongs, and obtain a classification prediction map;
[0012] S3, Model training: Use the training set to train the hyperspectral image classification model constructed in step S2. During the training process, calculate the loss value between the prediction result and the actual label, optimize the model parameters through an iterative method until the model converges, and record and store the training-optimized model parameters;
[0013] S4, Model inference: Input the hyperspectral data in the test set into the trained hyperspectral image classification model, infer the predicted category, and evaluate the obtained predicted category through evaluation metrics.
[0014] For the above hyperspectral image classification method based on surrogate attention and the VIT architecture, the specific process of step S1 is as follows:
[0015] S11, Dataset division: Randomly divide the original hyperspectral dataset D into a training set and a test set , where the proportion of the training set is , and the proportion of the test set is ;
[0016] S12, Surrogate pixel block sampling: Centered on the pixel coordinates , perform pixel-by-pixel cropping on the processed dataset, and the window size is , obtain a set of proxy pixel block samples centered on pixels , where each proxy pixel block , has a size of , is the number of spectral channels, is expressed as:
[0017] ;
[0018] Among them, represents the clipping operation function, represents the proxy pixel block centered at coordinates ;
[0019] S13, Normalization and Data Augmentation: First, normalize the sampled proxy pixel block , and normalize the value to the range [0, 1]. The processing method is:
[0020] ;
[0021] Among them, is the normalized data, and are the minimum and maximum values in the proxy pixel block respectively;
[0022] Then perform data augmentation operations; perform data augmentation on through random rotation and flipping operations, and obtain a new set of proxy pixel block samples after augmentation. Each new proxy pixel block sample , has a size of , is the number of spectral channels, is expressed as:
[0023] ;
[0024] Among them, represents the data augmentation function;
[0025] S14, Construct a Set of Training Sample Pairs: Match the corresponding true class label for the th new proxy pixel block sample , represents the serial number of the new proxy pixel block sample, and establish a training sample pair , where , is the total number of classes, and finally obtain a set of training sample pairs , Expressed as:
[0026] ;
[0027] Wherein is the total number of samples in
[0028] For the above hyperspectral image classification method based on proxy attention and VIT architecture, the specific process of step S21 is as follows:
[0029] S211, sliding window grouping;
[0030] First, flatten from to vectors of size ; then perform vector padding operation on the flattened vector group to meet the requirement of grouping with each vector as the center during sliding window grouping; finally, perform sliding window grouping operation with the window size of vectors and a sliding step of 1, obtaining groups, each group containing vectors of size ; the entire grouping process is expressed as:
[0031] ;
[0032] Wherein, represents the result after grouping, represents the flattening operation, represents the padding operation, represents the sliding window grouping operation;
[0033] S212, grouping encoding: Encode each group to finally obtain the initial feature encoding matrix ;
[0034] First, project the group containing vectors into a vector space of size through linear projection, represents the feature dimension, and add positional encoding to obtain the initial feature encoding matrix of size , and the projection process is expressed as:
[0035] ;
[0036] Wherein, represents the linear projection operation, represents the positional encoding; the positional encoding The calculation method is as follows:
[0037] ;
[0038] ;
[0039] Among them, represents the position index, represents the dimension index.
[0040] For the above hyperspectral image classification method based on proxy attention and VIT architecture, the specific process of step S22 is as follows:
[0041] S221, construct an information aggregation module IAM to aggregate information. From a of size to obtain a proxy token matrix of size , indicating the information aggregation degree of the proxy token matrix;
[0042] S222, construct proxy attention; the calculation of proxy attention is divided into two stages:
[0043] First, aggregate information to calculate the proxy feature . The proxy token matrix is used as the query to aggregate information from the key and the value . Specifically, calculate the similarity between the proxy token matrix and the key , and perform a weighted sum on the value to obtain the proxy feature . The whole process is expressed as:
[0044] ;
[0045] ;
[0046] ;
[0047] ;
[0048] Among them, is the query matrix, is the key matrix, is the value matrix, represents matrix multiplication, , , are the weight matrices of the query, key, and value respectively, represents the dimension of the key vector, represents the matrix transpose operation, and softmax is the activation function;
[0049] Then, broadcast the information to obtain the final output of the proxy attention ; Use the aggregated proxy features as the new values, the proxy token matrix as the key, as the query, and perform the second self-attention calculation; this process broadcasts the global information in the proxy features back to each query to obtain the final output of the proxy attention :
[0050] - ;
[0051] Among them, - represents the proxy attention calculation operation, represents the matrix transpose operation;
[0052] S223, construct the multi-head proxy attention mechanism; use the multi-head proxy attention mechanism, that is, calculate the outputs of multiple proxy self-attention heads in parallel, and then splice and linearly map the outputs of multiple proxy self-attention heads; assume there are h proxy attention heads, and the calculation of the m-th proxy attention head is , and the multi-head output is - , and the calculation process is expressed as:
[0053] - ;
[0054] - ;
[0055] Among them, and are both learnable parameter matrices, represents the matrix splicing operation;
[0056] S224, construct the feed-forward neural network FFN;
[0057] The feed-forward neural network includes two linear transformation layers and an activation function. Assume the input is Z, and Z passes through the feed-forward neural network to output as:
[0058] ;
[0059] Among them, is the activation function, and are both weight matrices of the FFN, 、 are all bias terms;
[0060] S225, construct an L-level cascaded proxy attention encoder ATE; ATE includes an information aggregation module IAM, a multi-head proxy attention mechanism, and a feed-forward neural network FFN;
[0061] The composition of each layer of ATE is exactly the same. The initial feature encoding matrix is aggregated by the information aggregation module IAM to obtain a proxy token matrix ; The four matrices are input into the multi-head proxy attention mechanism together; the output of the multi-head proxy attention mechanism is then passed through a residual connection and normalization and then fed to the feed-forward neural network FFN; the output of FFN is subjected to another residual connection to obtain the feature map of the th layer , , where is the cascaded layer number of ATE.
[0062] For the above hyperspectral image classification method based on proxy attention and VIT architecture, the specific operation of the step S221 is as follows:
[0063] S2211, reshape the initial feature encoding matrix into a three-dimensional tensor , , , , where
[0064] ;
[0065] where, is the feature tensor after downsampling, , represents the adaptive average pooling operation;
[0066] S2212, design a multi-scale convolutional layer, which includes k two-dimensional convolutional operations with different kernel sizes. The size of each convolutional kernel is , is an adjustable positive integer. Perform parallel convolutional operations on the downsampled feature tensor to obtain the output feature maps of k convolutional branches:
[0067] ;
[0068] where, , br represents the th convolutional branch; is the The output feature map of a convolutional branch, ; is the th convolutional kernel parameter; represents a parallel convolutional operation;
[0069] S2213, introducing a channel attention mechanism, calculates the attention weights for each channel:
[0070] ;
[0071] Among them, represents a global average pooling operation; and are the parameters of the fully connected layer, , , is the dimensionality reduction factor; represents activation function; represents activation function; is the channel attention weight, ;
[0072] S2214, applies the channel attention weights to the output feature maps of each convolutional branch and adds the results to obtain the fused feature map G:
[0073] ;
[0074] Among them, represents a broadcast multiplication operation in the channel dimension; ;
[0075] S2215, reshapes the fused feature map into a proxy token matrix , :
[0076] ;
[0077] Among them, represents a reshaping operation that reshapes the fused feature map of into the proxy token matrix of .
[0078] The above hyperspectral image classification method based on proxy attention and VIT architecture, the specific process of step S23 is as follows:
[0079] First, apply a global average pooling operation to to obtain the global feature vector :
[0080] ;
[0081] Among them, the global feature vector ;
[0082] Then, feature transformation and classification are performed through two layers of MLP to obtain the classification probability vector :
[0083] ;
[0084] Among them, the classification probability vector , and are weight matrices, , are bias terms;
[0085] Finally, through function, is converted into a class probability distribution, and a classification prediction map is generated;
[0086] ;
[0087] Among them, is the predicted class probability distribution, with a dimension of Class, and the predicted class is the class index with the highest probability.
[0088] In the above hyperspectral image classification method based on proxy attention and VIT architecture, in step S3, during the model training stage, first, the training set data is input into the constructed hyperspectral image classification model for forward propagation calculation to obtain the predicted class probability distribution; the cross-entropy loss function is used to calculate the loss value between the prediction result and the true label, and the loss function Loss is:
[0089] ;
[0090] Among them, is the number of training samples, is the total number of classes, represents the true label value that the rd sample belongs to the th class, represents the probability value that the model predicts the th sample belongs to the th class;
[0091] The Adam optimizer is used for parameter optimization, the initial value of the learning rate is set to a constant, and the cosine annealing strategy is used for learning rate adjustment:
[0092] ;
[0093] Among them is the base learning rate, is the adjusted learning rate, is the current training epoch, is the maximum training epoch;
[0094] After all epochs of training are completed, save the globally optimal model parameters.
[0095] In the above hyperspectral image classification method based on proxy attention and VIT architecture, in step S4, the specific process of inferring the predicted category is as follows:
[0096] Input the hyperspectral data to be predicted in the test set into the trained hyperspectral image classification model batch by batch. The hyperspectral data to be predicted extracts the spatial-spectral joint representation through the sliding window grouping encoding module; after layer-by-layer feature fusion of the proxy attention encoder ATE, a feature map is generated; global average pooling is performed on the feature map to obtain a global feature vector, and then two-layer MLP is used to perform feature transformation and classification on the global feature vector to generate a classification probability vector , through function to is transformed into a category probability distribution, and finally the category index with the largest value in the category probability distribution is used as the predicted category of the pixel to generate a classification prediction map.
[0097] In the above hyperspectral image classification method based on proxy attention and VIT architecture, in step S4, the evaluation metrics include the overall accuracy , average accuracy and coefficient. The coefficient is an evaluation metric of the hyperspectral classification model. The calculation formulas of and are as follows:
[0098] ;
[0099] ;
[0100] Among them represents the total number of categories, represents the number of correctly classified samples, represents the total number of samples, represents the number of correctly classified samples in the th category, represents the th category of the total number of samples.
[0101] The beneficial effects of the present invention are as follows:
[0102] 1. The hyperspectral image classification model of the present invention has a simple and efficient structure, reducing the computational complexity of the traditional attention mechanism from to , that is, from the computational complexity of the square order to the linear computational complexity, greatly improving the computational efficiency of the model.
[0103] 2. The present invention constructs a sliding window grouping encoding module based on the sliding window mechanism and proxy pixel blocks, encoding the original hyperspectral data into an initial feature encoding matrix , which contains the spatial-spectral joint representation of the hyperspectral data, effectively fusing the local spectral features of adjacent positions, and improving the model's representation ability of the spatial-spectral joint features.
[0104] 3. The present invention constructs a proxy attention encoder ATE based on proxy attention Agent Attention, and performs hierarchical feature extraction through the proxy attention encoder ATE cascaded by L levels, realizing the layer-by-layer fusion and enhancement of the spatial-spectral information in the initial feature encoding matrix , while significantly reducing the computational complexity while maintaining a strong representation ability, achieving a good balance between computational efficiency and feature expression ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0105] Figure 1 is the overall flowchart of the present invention.
[0106] Figure 2 is the structural schematic diagram of the proxy attention encoder of the present invention.
[0107] Figure 3 is the schematic diagram of the proxy attention of the present invention.
[0108] Figure 4 is the schematic diagram of the multi-head proxy attention mechanism of the present invention.
[0109] Figure 5 is the schematic diagram of the information aggregation module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0110] The present invention will be further described below with reference to the drawings and embodiments.
[0111] As Figure 1 shown, a hyperspectral image classification method based on proxy attention and VIT architecture includes the following steps:
[0112] S1, data preprocessing: preprocess the hyperspectral image data and divide it into a training set and a test set.
[0113] This invention uses the Pavia University dataset, which was collected by the Reflective Optics System Imaging Spectrometer (ROSIS-3) sensor over the city of Pavia, Italy. It contains 610×340 pixels and 103 spectral bands. The image is divided into 9 classes, with a total of 42,776 labeled samples, including asphalt, grassland, gravel, trees, metal sheets, bare soil, asphalt, bricks, and shadows.
[0114] The specific process of step S1 is as follows:
[0115] S11, Dataset division: Randomly divide the original hyperspectral dataset D into a training set and a test set at a ratio of 1:9, where the proportion of the training set is and the proportion of the test set is ;
[0116] S12, Proxy pixel block sampling: Centered on the pixel coordinates , perform pixel-by-pixel cropping on the processed dataset with a window size of to obtain a set of proxy pixel block samples centered on pixels , where each proxy pixel block , has a size of , is the number of spectral channels, is expressed as:
[0117] ;
[0118] Among them, represents the cropping operation function, represents the proxy pixel block centered on the coordinates ;
[0119] S13, Normalization and data augmentation: First, perform normalization on the sampled proxy pixel blocks , standardize the values to the range [0,1], and the processing method is:
[0120] ;
[0121] Among them, is the normalized data, and are the minimum and maximum values in the proxy pixel block respectively;
[0122] Then perform data augmentation operations; perform data augmentation on through random rotation and flipping operations, and obtain a new set of proxy pixel block samples after augmentation , where each new surrogate pixel block sample , has a size of , is the number of spectral channels, and is expressed as:
[0123] ;
[0124] where represents the data augmentation function;
[0125] S14. Construct a set of training sample pairs: For the th new surrogate pixel block sample match the corresponding true class label , represents the serial number of the new surrogate pixel block sample, and establish a training sample pair , where , is the total number of classes, and finally obtain the set of training sample pairs , and is expressed as:
[0126] ;
[0127] where is the total number of samples in, and thus complete the operation of data preprocessing.
[0128] S2. Model establishment: Construct a hyperspectral image classification model based on surrogate attention and VIT architecture.
[0129] The specific process of step S2 is as follows:
[0130] Step S21: Construct a sliding window grouping encoding module based on the sliding window mechanism and surrogate pixels, and encode the original hyperspectral data into an initial feature encoding matrix , which contains the spatial-spectral joint representation of the hyperspectral data.
[0131] The specific process of step S21 is as follows:
[0132] S211. Sliding window grouping;
[0133] First, flatten , from to pieces with a size of vectors; then perform a vector padding operation on the flattened vector group to meet the requirement of grouping with each vector as the center during the sliding window grouping; finally, perform the sliding window grouping operation with the window size of vectors and a sliding step of 1 to obtain groups, with each group containing vectors of size ; the entire grouping process is expressed as:
[0134] ;
[0135] where represents the result after grouping, represents the flattening operation, represents the padding operation, represents the sliding window grouping operation;
[0136] S212, Grouping Encoding: Encode each group to finally obtain the initial feature encoding matrix ;
[0137] First, project the group containing vectors onto a vector space of size through linear projection, represents the feature dimension, and add positional encoding to obtain the initial feature encoding matrix of size , and the projection process is expressed as:
[0138] ;
[0139] where represents the linear projection operation, represents the positional encoding; the calculation method of the positional encoding is:
[0140] ;
[0141] ;
[0142] where represents the position index, represents the dimension index.
[0143] Step S22: Construct an agent attention encoder ATE (Agent - Trans Encoder) based on agent attention, and perform hierarchical feature extraction through L cascaded agent attention encoders ATE to achieve the initial feature encoding matrix Layer-by-layer fusion and enhancement of spatial-spectral information to obtain a feature map.
[0144] The specific process of step S22 is as follows:
[0145] S221, as Figure 5 shown, construct an information aggregation module IAM to aggregate the information, and obtain a proxy token matrix of size from with a size of , indicating the information aggregation degree of the proxy token matrix;
[0146] The specific operation of step S221 is as follows:
[0147] S2211, reshape the initial feature encoding matrix into a three-dimensional tensor , , , indicating the dimensions of the matrix, and perform preliminary downsampling through an adaptive average pooling operation:
[0148] ;
[0149] Among them, is the downsampled feature tensor, , indicating the adaptive average pooling operation;
[0150] S2212, design a multi-scale convolutional layer, including k two-dimensional convolutional operations with different kernel sizes, and the size of each convolutional kernel is , is an adjustable positive integer, and perform parallel convolutional operations on the downsampled feature tensor to obtain the output feature maps of k convolutional branches:
[0151] ;
[0152] Among them, , br represents the th convolutional branch; is the output feature map of the th convolutional branch, ; is the th convolutional kernel parameter; indicates the parallel convolutional operation;
[0153] S2213, introduce a channel attention mechanism to calculate the attention weights of each channel:
[0154] ;
[0155] Among them, represents the global average pooling operation; and are the fully connected layer parameters, , , are the dimensionality reduction factors; represents the activation function; represents the activation function; is the channel attention weight, ;
[0156] S2214, apply the channel attention weight to the output feature maps of each convolutional branch, and add the results to obtain the fused feature map G:
[0157] ;
[0158] Among them, represents the broadcast multiplication operation in the channel dimension; ;
[0159] S2215, reshape the fused feature map into the proxy token matrix , :
[0160] ;
[0161] Among them, represents the reshape operation, and the reshape operation converts the fused feature map into the proxy token matrix .
[0162] S222, as Figure 3 shown, construct the proxy attention; the calculation of the proxy attention is divided into two stages:
[0163] First, aggregate information and calculate the proxy feature , using the proxy token matrix as the query, aggregate information from the key and the value ; specifically, calculate the similarity between the proxy token matrix and the key , and perform a weighted sum of the values to obtain the proxy feature , and the whole process is expressed as:
[0164] ;
[0165] ;
[0166] ;
[0167] ;
[0168] Among them, is the query matrix, is the key matrix, is the value matrix, represents matrix multiplication, , , are the weight matrices of query, key, and value respectively, represents the dimension of the key vector, represents the matrix transpose operation, and softmax is the activation function;
[0169] Then, broadcast the information to obtain the final output of the proxy attention ; Aggregate the obtained proxy features as the new value, the proxy token matrix as the key, as the query, and perform the second self-attention calculation; this process broadcasts the global information in the proxy features back to each query to obtain the final output of the proxy attention :
[0170] - ;
[0171] Among them, - represents the proxy attention calculation operation, represents the matrix transpose operation.
[0172] S223, as shown in Figure 4 , construct a multi-head proxy attention mechanism; use the multi-head proxy attention mechanism Multi-Head Agent Attention, that is, parallelly calculate the outputs of multiple proxy self-attention heads, and then splice and linearly map the outputs of multiple proxy self-attention heads; assume there are h proxy attention heads, and the calculation of the m-th proxy attention head is , and the multi-head output is - , and the calculation process is expressed as:
[0173] - ;
[0174] - ;
[0175] Among them, and are both learnable parameter matrices, represents the matrix concatenation operation.
[0176] S224. Construct a feed-forward neural network FFN;
[0177] The feed-forward neural network includes two linear transformation layers and an activation function. Assuming the input is Z, Z passes through the feed-forward neural network and outputs as:
[0178]
[0179] Among them, is the activation function, and are both weight matrices of FFN, , are both bias terms.
[0180] S225. Construct an L-level cascaded proxy attention encoder ATE; ATE includes an information aggregation module IAM, a multi-head proxy attention mechanism, and a feed-forward neural network FFN;
[0181] The composition of each layer of ATE is exactly the same. The initial feature encoding matrix is aggregated by the information aggregation module IAM to obtain the proxy token matrix ; The four matrices are input into the multi-head proxy attention mechanism together; the output of the multi-head proxy attention mechanism is then passed through a residual connection and normalization and then fed to the feed-forward neural network FFN; the output of FFN is subjected to another residual connection to obtain the feature map of the th layer , where is the cascaded layer number of ATE.
[0182] Step S23: Construct a classification head composed of an average pooling layer and a multi-layer perceptron MLP layer to interpret the feature map, predict the category to which the original pixel belongs, and obtain a classification prediction map.
[0183] The specific process of the said step S23 is as follows:
[0184] First, apply a global average pooling operation to to obtain a global feature vector :
[0185] ;
[0186] Among them, the global feature vector ;
[0187] Then, feature transformation and classification are performed through two layers of MLP to obtain the classification probability vector :
[0188] ;
[0189] Among them, the classification probability vector , and are weight matrices, , are bias terms;
[0190] Finally, through function, is converted into a class probability distribution, and a classification prediction map is generated;
[0191] ;
[0192] Among them, is the predicted class probability distribution, with a dimension of Class, and the predicted class is the class index with the highest probability.
[0193] S3. Model training: The hyperspectral image classification model constructed in step S2 is trained using the training set. During the training process, the loss value between the prediction result and the actual label is calculated, and the model parameters are optimized through an iterative method until the model converges. The optimized model parameters during training are recorded and stored.
[0194] In the model training stage, first, the training set data is input into the constructed hyperspectral image classification model for forward propagation calculation to obtain the predicted class probability distribution; the cross-entropy loss function is used to calculate the loss value between the prediction result and the true label. The loss function Loss is:
[0195] ;
[0196] Among them, is the number of training samples, is the total number of classes, represents the true label value that the -th sample belongs to the -th class, represents the probability value that the model predicts the -th sample belongs to the -th class;
[0197] The Adam optimizer is used for parameter optimization. The initial value of the learning rate is set as a constant, and the cosine annealing strategy is adopted for learning rate adjustment:
[0198] ;
[0199] Among them is the base learning rate, is the adjusted learning rate, is the current training epoch, is the maximum training epoch;
[0200] After all epochs of training are completed, save the globally optimal model parameters.
[0201] S4, Model Inference: Input the hyperspectral data in the test set into the trained hyperspectral image classification model, infer the predicted classes, and evaluate the obtained predicted classes using evaluation metrics.
[0202] The specific process of inferring the predicted classes is as follows:
[0203] Input the hyperspectral data to be predicted in the test set into the trained hyperspectral image classification model batch by batch. The hyperspectral data to be predicted extracts the spatial-spectral joint representation through the sliding window grouping encoding module; after the layer-by-layer feature fusion of the proxy attention encoder ATE, a feature map is generated; perform global average pooling on the feature map to obtain a global feature vector, and then use two-layer MLP to perform feature transformation and classification on the global feature vector to generate a classification probability vector , through function to is transformed into a class probability distribution, and finally the class index with the largest value in the class probability distribution is used as the predicted class of the pixel to generate a classification prediction map.
[0204] The evaluation metrics include overall accuracy , average accuracy and coefficient. The coefficient is an evaluation metric for the hyperspectral classification model. The calculation formulas of and are as follows:
[0205] ;
[0206] ;
[0207] Among them represents the total number of classes, represents the number of correctly classified samples, represents the total number of samples, represents the th class of the number of correctly classified samples, represents the th class of the total number of samples.
[0208]
[0209] Table 1 shows the accuracy of the method of the present invention in the Pavia University dataset. Transformer (VIT) is the Baseline of this model, SpectralFormer is a model improved based on the VIT architecture, and Ours is the hyperspectral image classification model provided by the present invention.
[0210] Using the overall accuracy and the average accuracy and coefficient to evaluate the accuracy of hyperspectral image classification. It can be seen from Table 1 that the classification result of the model of the present invention is significantly improved compared with the Baseline.
Claims
1. A hyperspectral image classification method based on proxy attention and VIT architecture, characterized in that: The following steps are involved: S1, data preprocessing: preprocess the hyperspectral image data and divide it into training set and test set; S2, model building: building a hyperspectral image classification model based on agent attention and VIT architecture; The specific process of step S2 is as follows: Step S21: Construct a sliding window group coding module based on the sliding window mechanism and the proxy pixel block to encode the original hyperspectral data into an initial feature coding matrix , Contains spatial-spectral joint characterization of hyperspectral data; Step S22: Construct an agent attention encoder ATE based on agent attention, perform hierarchical feature extraction through L layers of cascaded agent attention encoder ATE, and realize the initial feature encoding matrix The spatial-spectral information is fused and enhanced layer by layer to obtain the feature map; Step S23: construct a classification head consisting of an average pooling layer and a multi-layer perceptron MLP layer, interpret the feature map, predict the category to which the original pixel belongs, and obtain a classification prediction map; S3, model training: using the training set to train the hyperspectral image classification model constructed in step S2, calculating the loss value between the predicted result and the actual label during the training process, optimizing the model parameters in an iterative manner until the model converges, and recording and storing the training optimized model parameters; S4, model reasoning: The hyperspectral data in the test set is input into the trained hyperspectral image classification model, the predicted category is obtained by reasoning, and the obtained predicted category is evaluated by the evaluation index.
2. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 1 is characterized in that: The specific process of step S1 is as follows: S11, data set division: randomly divide the original hyperspectral data set D into training sets according to the proportion and test set , where the training set accounts for , the test set accounts for ; S12, proxy pixel block sampling: in pixel coordinates Centered on the pixel, the processed data set is cropped pixel by pixel, with a window size of , get the sample set of proxy pixel blocks centered on the pixel , where each proxy pixel block , The size is , is the number of spectral channels, It is expressed as: ; in, represents the trimming operation function, Represented by coordinates The proxy pixel block is centered; S13, normalization and data enhancement: First, the sampled proxy pixel blocks are Normalize it and The value of is standardized to the interval [0,1], and the processing method is: ; in, is the normalized data, and are the minimum and maximum values in the proxy pixel block, respectively; Then perform data enhancement operations; randomly rotate and flip the Perform data enhancement to obtain a new proxy pixel block sample set , where each new proxy pixel block sample , The size is , is the number of spectral channels, It is expressed as: ; in, represents the data enhancement function; S14, construct a set of training sample pairs: New proxy pixel patch samples Match the corresponding true category label , Indicates the serial number of the new proxy pixel block sample, and establishes the training sample pair ,in , is the total number of categories, and finally the training sample pair set is obtained , It is expressed as: ; in for The total number of samples in .
3. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 2 is characterized in that: The specific process of step S21 is as follows: S211, sliding window grouping; First, Leveled, by becomes The size is vector; then the vector filling operation is performed on the flattened vector group to meet the requirement of grouping with each vector as the center when performing sliding window grouping; finally, the sliding window grouping operation is performed, and the window size is vectors, the sliding step size is 1, and we get groups, each containing indivual The whole grouping process is expressed as: ; in, Represents the result after grouping. Indicates the leveling operation. represents the filling operation, Represents a sliding window grouping operation; S212, group coding: Encode each group and finally obtain the initial feature coding matrix ; First, linear projection is used to indivual The grouping of vectors is projected onto a In the vector space of size, Represents the feature dimension and adds position encoding to get a size of The initial feature encoding matrix , the projection process is expressed as: ; in, represents a linear projection operation, Represents positional encoding; positional encoding The calculation method is: ; ; in, Represents the position index, Represents a dimension index.
4. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 3 is characterized in that: The specific process of step S22 is as follows: S221, build information aggregation module IAM, aggregation The information of size of Get a size of The proxy tag matrix , represents the degree of information aggregation of the proxy labeling matrix; S222, construct agent attention; The calculation of agent attention is divided into two stages: First, aggregate information and calculate proxy features , the proxy labeling matrix As a query, from the key Sum Aggregate information in ; specifically, calculate the proxy label matrix With key , and perform weighted summation of the values to obtain the proxy feature , the whole process is expressed as: ; ; ; ; in, is the query matrix, is the key matrix, is the value matrix, represents matrix multiplication, , , are the weight matrices for query, key, and value, respectively. represents the dimension of the key vector, represents the matrix transposition operation, and softmax is the activation function; Then, broadcast the information and get the final output of the agent attention ; Aggregate the proxy features As new values, the proxy label matrix As a key, As a query, a second self-attention calculation is performed; this process broadcasts the global information in the proxy features back to each query, and obtains the final output of the proxy attention : - ; in, - represents the agent attention calculation operation, Represents a matrix transpose operation; S223, construct a multi-head proxy attention mechanism; use the multi-head proxy attention mechanism, that is, parallelly calculate the outputs of multiple proxy self-attention heads, and then concatenate and linearly map the outputs of multiple proxy self-attention heads; assuming there are h proxy attention heads, the calculation of the mth proxy attention head is , the multi-head output is - , the calculation process is expressed as: - ; - ; in, and are all learnable parameter matrices, Represents a matrix concatenation operation; S224, constructing a feed-forward neural network FFN; The feedforward neural network consists of two linear transformation layers and an activation function. Assuming the input is Z, Z is output through the feedforward neural network. for: ; in, is the activation function, and are all weight matrices of FFN, , All are bias terms; S225, constructing an L-layer cascaded agent attention encoder ATE; ATE includes an information aggregation module IAM, a multi-head agent attention mechanism and a feed-forward neural network FFN; The structure of each layer of ATE is exactly the same, the initial feature encoding matrix The proxy label matrix is obtained by aggregation through the information aggregation module IAM ; The four matrices are input into the multi-head agent attention mechanism together; the output of the multi-head agent attention mechanism is then passed to the feed-forward neural network FFN after residual connection and normalization; the output of FFN is residually connected again to obtain the first Feature map of the layer , ,in is the number of cascaded layers of ATE.
5. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 4 is characterized in that: The specific operation of step S221 is as follows: S2211, the initial feature encoding matrix Reshape into a three-dimensional tensor , , , Represents the dimensions of the matrix and performs preliminary downsampling through an adaptive average pooling operation: ; in, is the feature tensor after downsampling, , represents the adaptive average pooling operation; S2212, design a multi-scale convolution layer, which contains k two-dimensional convolution operations with different kernel sizes. The size of each convolution kernel is , is an adjustable positive integer, which is used to adjust the downsampled feature tensor Perform parallel convolution operations to obtain the output feature maps of k convolution branches: ; in, , br means the convolution branches; For the The output feature map of the convolution branch, ; For the convolution kernel parameters; Represents a parallel convolution operation; S2213, introduce the channel attention mechanism and calculate the attention weight of each channel: ; in, represents the global average pooling operation; and is the fully connected layer parameter, , , is the dimensionality reduction factor; express Activation function; express Activation function; is the channel attention weight, ; S2214, apply the channel attention weights to the output feature maps of each convolution branch, and add the results to obtain the fused feature map G: ; in, Represents a broadcast multiplication operation on the channel dimension; ; S2215, fusion feature map Reshape into a proxy-labeled matrix , : ; in, Represents a reshape operation, which will Fusion feature map Convert to The proxy tag matrix .
6. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 5 is characterized in that: The specific process of step S23 is as follows: First of all, Apply the global average pooling operation to obtain the global feature vector : ; Among them, the global eigenvector ; Then, feature transformation and classification are performed through two layers of MLP to obtain the classification probability vector : ; Among them, the classification probability vector , and is the weight matrix, , is the bias term; Finally, through The function will Convert to category probability distribution and generate classification prediction graph; ; in, is the predicted category probability distribution, the dimension is Class, and the predicted category is the category index with the highest probability.
7. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 6 is characterized in that: In step S3, in the model training stage, the training set data is first input into the constructed hyperspectral image classification model for forward propagation calculation to obtain the predicted category probability distribution; the cross entropy loss function is used to calculate the loss value between the predicted result and the true label, and the loss function Loss is: ; in, is the number of training samples, is the total number of categories, Indicates The samples belong to The true label value of the class, The model predicts The samples belong to The probability value of the class; The Adam optimizer is used for parameter optimization, the initial value of the learning rate is set to a constant, and the cosine annealing strategy is used to adjust the learning rate: ; in is the basic learning rate, is the adjusted learning rate, is the current training round, is the maximum number of training rounds; After all rounds of training are completed, save the global optimal model parameters.
8. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 6 is characterized in that: In step S4, the specific process of inferring the predicted category is: The hyperspectral data to be predicted in the test set are input into the trained hyperspectral image classification model batch by batch. The spatial-spectral joint representation of the hyperspectral data to be predicted is extracted through the sliding window group encoding module. The feature map is generated after layer-by-layer feature fusion by the agent attention encoder ATE. The feature map is globally averaged pooled to obtain the global feature vector, and then the global feature vector is transformed and classified using a two-layer MLP to generate a classification probability vector. ,pass The function will It is converted into a category probability distribution, and finally the largest category index in the category probability distribution is used as the predicted category of the pixel to generate a classification prediction map.
9. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 6, characterized in that: In step S4, the evaluation index includes the overall accuracy , average accuracy and coefficient, The coefficient is the evaluation index of the hyperspectral classification model. and The calculation formula is as follows: ; ; in represents the total number of categories, represents the number of correctly classified samples, represents the total number of samples, Indicates The number of correctly classified samples in the class, Indicates The total number of samples in the class.
Citation Information
Patent Citations
Coding and decoding structure semantic segmentation model based on position attention mechanism
CN115908793A
Hyperspectral image classification method based on high-order interactive convolutional network
CN118506112A