Hyperspectral image classification method based on proxy attention and VIT architecture
By introducing proxy attention and VIT architecture into the hyperspectral image classification model, the problems of existing models' computational complexity and large parameters are solved, and efficient hyperspectral image classification is achieved.
Patent Information
- Application Number
- CN202510435730.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
When processing hyperspectral remote sensing images containing hundreds of continuous spectrum bands, the existing hyperspectral image classification model faces the problems of computational complexity and large amount of parameters, and it is difficult to achieve a good balance between computational efficiency and representation ability.
Using a hyperspectral image classification method based on proxy attention and Vision Transformer (VIT) architecture, the spatial-spectral joint characterization and hierarchical feature extraction of hyperspectral data is realized through the sliding window grouping encoding module and the proxy attention encoder.
It significantly reduces the computational complexity of the model, improves the computational efficiency, and maintains strong representation ability, achieving the accuracy and efficiency of hyperspectral image classification.
Smart Images

Figure CN119942250A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image classification, and in particular to a hyperspectral image classification method based on proxy attention and VIT architecture. Background Art
[0002] As an advanced remote sensing technology, hyperspectral image (HIS) has received extensive attention in recent years. HSI has both rich spectral information and spatial information. Each pixel contains hundreds of narrow spectral band information in the channel dimension. This unique data structure makes hyperspectral images show great application potential in many fields such as urban construction, ocean observation, precision agriculture, disaster prevention and mitigation, and resource exploration.
[0003] Early hyperspectral image classification mainly relied on traditional machine learning methods, such as Random-Forest algorithm, support vector machine, etc. However, these methods have limitations in feature extraction ability and adaptability to complex scenes. In recent years, deep learning technology, especially convolutional neural network (CNN), has been widely used in hyperspectral image classification tasks due to its powerful feature learning ability. However, based on the data characteristics of HSI images containing hundreds of continuous and approximate spectral band information, the deep semantic features extracted by CNN models are often limited; and as the model depth increases, the computational cost of CNN models increases significantly. The model ViT (Vision Transformer) allows the deep learning model Transformer to achieve significant results in the field of computer vision, while also bringing new opportunities for HSI classification. Compared with CNN using a sliding window mechanism to gradually achieve global information fusion, Transformer has the advantage of capturing global information at one time, which makes it more suitable for processing HSI data containing hundreds of continuous spectral bands. However, Transformer still faces some challenges when applied to HSI classification: in order to develop a lightweight hyperspectral classification model suitable for edge devices, it is necessary to reduce the number of model parameters and computational complexity, and how to achieve a good balance between computational efficiency and representation capability has become a daunting task. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a hyperspectral image classification method based on proxy attention and VIT architecture with simple algorithm and strong representation ability.
[0005] The technical solution of the present invention to solve the above technical problems is: a hyperspectral image classification method based on proxy attention and VIT architecture, comprising the following steps: S1, data preprocessing: preprocess the hyperspectral image data and divide it into training set and test set; S2, model building: building a hyperspectral image classification model based on agent attention and VIT architecture; The specific process of step S2 is as follows: Step S21: Construct a sliding window group coding module based on the sliding window mechanism and the proxy pixel block to encode the original hyperspectral data into an initial feature coding matrix , Contains spatial-spectral joint characterization of hyperspectral data; Step S22: Construct an agent attention encoder ATE based on agent attention, perform hierarchical feature extraction through L layers of cascaded agent attention encoder ATE, and realize the initial feature encoding matrix The spatial-spectral information is fused and enhanced layer by layer to obtain the feature map; Step S23: construct a classification head consisting of an average pooling layer and a multi-layer perceptron MLP layer, interpret the feature map, predict the category to which the original pixel belongs, and obtain a classification prediction map; S3, model training: using the training set to train the hyperspectral image classification model constructed in step S2, calculating the loss value between the predicted result and the actual label during the training process, optimizing the model parameters in an iterative manner until the model converges, and recording and storing the training optimized model parameters; S4, model reasoning: The hyperspectral data in the test set is input into the trained hyperspectral image classification model, the predicted category is obtained by reasoning, and the obtained predicted category is evaluated by the evaluation index.
[0006] The hyperspectral image classification method based on agent attention and VIT architecture, the specific process of step S1 is as follows: S11, data set division: randomly divide the original hyperspectral data set D into training sets according to the proportion and test set , where the training set accounts for , the test set accounts for ; S12, proxy pixel block sampling: in pixel coordinates Centered on the pixel, the processed data set is cropped pixel by pixel, with a window size of , get the sample set of proxy pixel blocks centered on the pixel , where each proxy pixel block , The size is , is the number of spectral channels, It is expressed as: ; in, represents the trimming operation function, Represented by coordinates The proxy pixel block is centered; S13, normalization and data enhancement: First, the sampled proxy pixel blocks are Normalize it and The value of is standardized to the interval [0,1], and the processing method is: ; in, is the normalized data, and are the minimum and maximum values in the proxy pixel block, respectively; Then perform data enhancement operations; randomly rotate and flip the Perform data enhancement to obtain a new proxy pixel block sample set , where each new proxy pixel block sample , The size is , is the number of spectral channels, It is expressed as: ; in, represents the data enhancement function; S14, construct a set of training sample pairs: New proxy pixel patch samples Match the corresponding true category label , Indicates the serial number of the new proxy pixel block sample, and establishes the training sample pair ,in , is the total number of categories, and finally the training sample pair set is obtained , It is expressed as: ; in for The total number of samples in .
[0007] The hyperspectral image classification method based on proxy attention and VIT architecture, the specific process of step S21 is as follows: S211, sliding window grouping; First, Leveled, by becomes The size is vector; then the vector filling operation is performed on the flattened vector group to meet the requirement of grouping with each vector as the center when performing sliding window grouping; finally, the sliding window grouping operation is performed, and the window size is vectors, the sliding step size is 1, and we get groups, each containing indivual The whole grouping process is expressed as: ; in, Represents the result after grouping. Indicates the leveling operation. represents a filling operation, Represents a sliding window grouping operation; S212, group coding: Encode each group and finally obtain the initial feature coding matrix ; First, linear projection is used to indivual The grouping of vectors is projected onto a In the vector space of size, Represents the feature dimension and adds position encoding to get a size of The initial feature encoding matrix , the projection process is expressed as: ; in, represents a linear projection operation, Represents positional encoding; positional encoding The calculation method is: ; ; in, Represents the position index, Represents a dimension index.
[0008] In the hyperspectral image classification method based on proxy attention and VIT architecture, the specific process of step S22 is as follows: S221, build information aggregation module IAM, aggregation The information of size of Get a size of The proxy tag matrix , represents the degree of information aggregation of the proxy labeling matrix; S222, build proxy attention; the calculation of proxy attention is divided into two stages: First, aggregate information and calculate proxy features , the proxy labeling matrix As a query, from the key Sum Aggregate information in ; specifically, calculate the proxy label matrix With key , and perform weighted summation of the values to obtain the proxy feature , the whole process is expressed as: ; ; ; ; in, is the query matrix, is the key matrix, is the value matrix, represents matrix multiplication, , , are the weight matrices for query, key, and value, respectively. represents the dimension of the key vector, represents the matrix transposition operation, and softmax is the activation function; Then, broadcast the information and get the final output of the agent attention ; Aggregate the proxy features As new values, the proxy label matrix As a key, As a query, a second self-attention calculation is performed; this process broadcasts the global information in the proxy features back to each query, and obtains the final output of the proxy attention : - ; in, - represents the agent attention calculation operation, Represents a matrix transpose operation; S223, construct a multi-head proxy attention mechanism; use the multi-head proxy attention mechanism, that is, parallelly calculate the outputs of multiple proxy self-attention heads, and then concatenate and linearly map the outputs of multiple proxy self-attention heads; assuming there are h proxy attention heads, the calculation of the mth proxy attention head is , the multi-head output is - , the calculation process is expressed as: - ; - ; in, and are all learnable parameter matrices, Represents a matrix concatenation operation; S224, constructing a feed-forward neural network FFN; The feedforward neural network consists of two linear transformation layers and an activation function. Assuming the input is Z, Z is output through the feedforward neural network. for: ; in, is the activation function, and are all weight matrices of FFN, , All are bias terms; S225, constructing an L-layer cascaded agent attention encoder ATE; ATE includes an information aggregation module IAM, a multi-head agent attention mechanism and a feed-forward neural network FFN; The structure of each layer of ATE is exactly the same, the initial feature encoding matrix The proxy label matrix is obtained by aggregation through the information aggregation module IAM ; The four matrices are input into the multi-head agent attention mechanism together; the output of the multi-head agent attention mechanism is then passed to the feed-forward neural network FFN after residual connection and normalization; the output of FFN is residually connected again to obtain the first Feature map of the layer , ,in is the number of cascaded layers of ATE.
[0009] In the hyperspectral image classification method based on proxy attention and VIT architecture, the specific operations of step S221 are as follows: S2211, the initial feature encoding matrix Reshape into a three-dimensional tensor , , , Represents the dimensions of the matrix and performs preliminary downsampling through an adaptive average pooling operation: ; in, is the feature tensor after downsampling, , represents the adaptive average pooling operation; S2212, design a multi-scale convolution layer, which contains k two-dimensional convolution operations with different kernel sizes. The size of each convolution kernel is , is an adjustable positive integer, which is used to adjust the downsampled feature tensor Perform parallel convolution operations to obtain the output feature maps of k convolution branches: ; in, , br means the convolution branches; For the The output feature map of the convolution branch, ; For the convolution kernel parameters; Represents parallel convolution operation; S2213, introduce the channel attention mechanism and calculate the attention weights of each channel: ; in, represents the global average pooling operation; and is the fully connected layer parameter, , , is the dimensionality reduction factor; express Activation function; express Activation function; is the channel attention weight, ; S2214, apply the channel attention weights to the output feature maps of each convolution branch, and add the results to obtain the fused feature map G: ; in, Represents a broadcast multiplication operation on the channel dimension; ; S2215, fusion feature map Reshape into a proxy-labeled matrix , : ; in, Represents a reshape operation, which will Fusion feature map Convert to The proxy tag matrix .
[0010] The hyperspectral image classification method based on proxy attention and VIT architecture, the specific process of step S23 is as follows: First of all, Apply the global average pooling operation to obtain the global feature vector : ; Among them, the global eigenvector ; Then, feature transformation and classification are performed through two layers of MLP to obtain the classification probability vector : ; Among them, the classification probability vector , and is the weight matrix, , is the bias term; Finally, through The function will Convert to category probability distribution and generate classification prediction graph; ; in, is the predicted category probability distribution, the dimension is Class, and the predicted category is the category index with the highest probability.
[0011] In the above-mentioned hyperspectral image classification method based on proxy attention and VIT architecture, in step S3, in the model training stage, the training set data is first input into the constructed hyperspectral image classification model for forward propagation calculation to obtain the predicted category probability distribution; the cross entropy loss function is used to calculate the loss value between the predicted result and the true label, and the loss function Loss is: ; in, is the number of training samples, is the total number of categories, Indicates The samples belong to The true label value of the class, The model predicts The samples belong to The probability value of the class; The Adam optimizer is used for parameter optimization, the initial value of the learning rate is set to a constant, and the cosine annealing strategy is used to adjust the learning rate: ; in is the basic learning rate, is the adjusted learning rate, is the current training round, is the maximum number of training rounds; After all rounds of training are completed, save the global optimal model parameters.
[0012] In the hyperspectral image classification method based on proxy attention and VIT architecture, in step S4, the specific process of inferring the predicted category is: The hyperspectral data to be predicted in the test set are input into the trained hyperspectral image classification model batch by batch. The spatial-spectral joint representation of the hyperspectral data to be predicted is extracted through the sliding window group encoding module. The feature map is generated after layer-by-layer feature fusion by the agent attention encoder ATE. The feature map is globally averaged pooled to obtain the global feature vector, and then the global feature vector is transformed and classified using a two-layer MLP to generate a classification probability vector. ,pass The function will It is converted into a category probability distribution, and finally the largest category index in the category probability distribution is used as the predicted category of the pixel to generate a classification prediction map.
[0013] In the above-mentioned hyperspectral image classification method based on agent attention and VIT architecture, in step S4, the evaluation index includes the overall accuracy , average accuracy and coefficient, The coefficient is the evaluation index of the hyperspectral classification model. and The calculation formula is as follows: ; ; in represents the total number of categories, represents the number of correctly classified samples, represents the total number of samples, Indicates The number of correctly classified samples in the class, Indicates The total number of samples in the class.
[0014] The beneficial effects of the present invention are: 1. The hyperspectral image classification model of the present invention is simple and efficient, reducing the computational complexity of the traditional attention mechanism. from Reduce to , that is, the computational complexity is reduced from quadratic to linear, greatly improving the computational efficiency of the model.
[0015] 2. The present invention encodes the original hyperspectral data into an initial feature encoding matrix by constructing a sliding window group encoding module based on a sliding window mechanism and a proxy pixel block. , It includes the spatial-spectral joint representation of hyperspectral data, effectively integrates the local spectral features of adjacent positions, and improves the model's ability to represent the spatial-spectral joint features.
[0016] 3. The present invention constructs an agent attention encoder ATE based on agent attention, and performs hierarchical feature extraction through L-layer cascaded agent attention encoder ATE to realize the initial feature encoding matrix The layer-by-layer fusion and enhancement of spatial-spectral information significantly reduces the computational complexity while maintaining strong representation capabilities, achieving a good balance between computational efficiency and feature expression capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is the overall flow chart of the present invention.
[0018] Figure 2 Schematic diagram of the structure of the proxy attention encoder of the present invention.
[0019] Figure 3 Schematic diagram of the agent attention of the present invention.
[0020] Figure 4 Schematic diagram of the multi-head agent attention mechanism of the present invention.
[0021] Figure 5 Schematic diagram of the information aggregation module of the present invention. DETAILED DESCRIPTION
[0022] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0023] like Figure 1 As shown, a hyperspectral image classification method based on agent attention and VIT architecture includes the following steps: S1, data preprocessing: preprocess the hyperspectral image data and divide it into training set and test set.
[0024] This invention uses the Pavia University dataset, which is collected by the Reflection Optical System Imaging Spectrometer (ROSIS-3) sensor over the city of Pavia, Italy, and contains 610×340 pixels and 103 spectral bands. The images are divided into 9 categories, with a total of 42,776 labeled samples, including asphalt, grass, gravel, trees, metal sheets, bare soil, asphalt, bricks, and shadows.
[0025] The specific process of step S1 is as follows: S11, data set division: The original hyperspectral data set D is randomly divided into training sets in a ratio of 1:9. and test set , where the training set accounts for , the test set accounts for ; S12, proxy pixel block sampling: in pixel coordinates Centered on the pixel, the processed data set is cropped pixel by pixel, with a window size of , get the sample set of proxy pixel blocks centered on the pixel , where each proxy pixel block , The size is , is the number of spectral channels, It is expressed as: ; in, represents the trimming operation function, Represented by coordinates The proxy pixel block is centered; S13, normalization and data enhancement: First, the sampled proxy pixel blocks are Normalize it and The value of is standardized to the interval [0,1], and the processing method is: ; in, is the normalized data, and are the minimum and maximum values in the proxy pixel block, respectively; Then perform data enhancement operations; randomly rotate and flip the Perform data enhancement to obtain a new proxy pixel block sample set , where each new proxy pixel block sample , The size is , is the number of spectral channels, It is expressed as: ; in, represents the data enhancement function; S14, construct a set of training sample pairs: New proxy pixel patch samples Match the corresponding true category label , Indicates the serial number of the new proxy pixel block sample, and establishes the training sample pair ,in , is the total number of categories, and finally the training sample pair set is obtained , It is expressed as: ; in for The total number of samples in , and the data preprocessing operation is completed.
[0026] S2, Model Building: Construct a hyperspectral image classification model based on agent attention and VIT architecture.
[0027] The specific process of step S2 is as follows: Step S21: Construct a sliding window group coding module based on the sliding window mechanism and the proxy pixel block to encode the original hyperspectral data into an initial feature coding matrix , Contains spatial-spectral joint characterization of hyperspectral data.
[0028] The specific process of step S21 is as follows: S211, sliding window grouping; First, Leveled, by becomes The size is vector; then the vector filling operation is performed on the flattened vector group to meet the requirement of grouping with each vector as the center when performing sliding window grouping; finally, the sliding window grouping operation is performed, and the window size is vectors, the sliding step size is 1, and we get groups, each containing indivual The whole grouping process is expressed as: ; in, Represents the result after grouping. Indicates the leveling operation. represents a filling operation, Represents a sliding window grouping operation; S212, group coding: Encode each group and finally obtain the initial feature coding matrix ; First, linear projection is used to indivual The grouping of vectors is projected onto a In the vector space of size, Represents the feature dimension and adds position encoding to get a size of The initial feature encoding matrix , the projection process is expressed as: ; in, represents a linear projection operation, Represents positional encoding; positional encoding The calculation method is: ; ; in, Represents the position index, Represents a dimension index.
[0029] Step S22: Construct an agent-trans encoder ATE (Agent-Trans Encoder) based on agent attention, perform hierarchical feature extraction through L layers of cascaded agent-trans encoder ATE, and realize the initial feature encoding matrix The spatial-spectral information is fused and enhanced layer by layer to obtain the feature map.
[0030] The specific process of step S22 is as follows: S221, such as Figure 5 As shown, we build an information aggregation module IAM to aggregate The information of size of Get a size of The proxy tag matrix , represents the degree of information aggregation of the proxy labeling matrix; The specific operation of step S221 is as follows: S2211, the initial feature encoding matrix Reshape into a three-dimensional tensor , , , Represents the dimensions of the matrix and performs preliminary downsampling through an adaptive average pooling operation: ; in, is the feature tensor after downsampling, , represents the adaptive average pooling operation; S2212, design a multi-scale convolution layer, which contains k two-dimensional convolution operations with different kernel sizes. The size of each convolution kernel is , is an adjustable positive integer, which is used to adjust the downsampled feature tensor Perform parallel convolution operations to obtain the output feature maps of k convolution branches: ; in, , br means the convolution branches; For the The output feature map of the convolution branch, ; For the convolution kernel parameters; Represents parallel convolution operation; S2213, introduce the channel attention mechanism and calculate the attention weights of each channel: ; in, represents the global average pooling operation; and is the fully connected layer parameter, , , is the dimensionality reduction factor; express Activation function; express Activation function; is the channel attention weight, ; S2214, apply the channel attention weights to the output feature maps of each convolution branch, and add the results to obtain the fused feature map G: ; in, Represents a broadcast multiplication operation on the channel dimension; ; S2215, fusion feature map Reshape into a proxy-labeled matrix , : ; in, Represents a reshape operation, which will Fusion feature map Convert to The proxy tag matrix .
[0031] S222, such as Figure 3 As shown, the agent attention is constructed; the calculation of agent attention is divided into two stages: First, aggregate information and calculate proxy features , the proxy labeling matrix As a query, from the key Sum Aggregate information in ; specifically, calculate the proxy label matrix With key , and perform weighted summation of the values to obtain the proxy feature , the whole process is expressed as: ; ; ; ; in, is the query matrix, is the key matrix, is the value matrix, represents matrix multiplication, , , are the weight matrices for query, key, and value, respectively. represents the dimension of the key vector, represents the matrix transposition operation, and softmax is the activation function; Then, broadcast the information and get the final output of the agent attention ; Aggregate the proxy features As new values, the proxy label matrix As a key, As a query, a second self-attention calculation is performed; this process broadcasts the global information in the proxy features back to each query, and obtains the final output of the proxy attention : - ; in, - represents the agent attention calculation operation, Represents a matrix transpose operation.
[0032] S223, such as Figure 4 As shown in Figure 1, a multi-head agent attention mechanism is constructed; the multi-head agent attention mechanism Multi-Head Agent Attention is used, that is, the outputs of multiple agent self-attention heads are calculated in parallel, and then the outputs of multiple agent self-attention heads are spliced and linearly mapped; assuming that there are h agent attention heads, the calculation of the mth agent attention head is , the multi-head output is - , the calculation process is expressed as: - ; - ; in, and are all learnable parameter matrices, Represents a matrix concatenation operation.
[0033] S224, constructing a feed-forward neural network FFN; The feedforward neural network consists of two linear transformation layers and an activation function. Assuming the input is Z, Z is output through the feedforward neural network. for:
[0034] in, is the activation function, and are all weight matrices of FFN, , All are bias terms.
[0035] S225, constructing an L-layer cascaded agent attention encoder ATE; ATE includes an information aggregation module IAM, a multi-head agent attention mechanism and a feed-forward neural network FFN; The structure of each layer of ATE is exactly the same, the initial feature encoding matrix The proxy label matrix is obtained by aggregation through the information aggregation module IAM ; The four matrices are input into the multi-head agent attention mechanism together; the output of the multi-head agent attention mechanism is then passed to the feed-forward neural network FFN after residual connection and normalization; the output of FFN is residually connected again to obtain the first Feature map of the layer , ,in is the number of cascaded layers of ATE.
[0036] Step S23: construct a classification head consisting of an average pooling layer and a multi-layer perceptron MLP layer, interpret the feature map, predict the category to which the original pixel belongs, and obtain a classification prediction map.
[0037] The specific process of step S23 is as follows: First of all, Apply the global average pooling operation to obtain the global feature vector : ; Among them, the global eigenvector ; Then, feature transformation and classification are performed through two layers of MLP to obtain the classification probability vector : ; Among them, the classification probability vector , and is the weight matrix, , is the bias term; Finally, through The function will Convert to category probability distribution and generate classification prediction graph; ; in, is the predicted category probability distribution, the dimension is Class, and the predicted category is the category index with the highest probability.
[0038] S3, model training: Use the training set to train the hyperspectral image classification model constructed in step S2, calculate the loss value between the predicted result and the actual label during the training process, optimize the model parameters in an iterative manner until the model converges, and record and store the training optimized model parameters.
[0039] In the model training stage, the training set data is first input into the constructed hyperspectral image classification model for forward propagation calculation to obtain the predicted category probability distribution; the cross entropy loss function is used to calculate the loss value between the predicted result and the true label. The loss function Loss is: ; in, is the number of training samples, is the total number of categories, Indicates The samples belong to The true label value of the class, The model predicts The samples belong to The probability value of the class; The Adam optimizer is used for parameter optimization, the initial value of the learning rate is set to a constant, and the cosine annealing strategy is used to adjust the learning rate: ; in is the basic learning rate, is the adjusted learning rate, is the current training round, is the maximum number of training rounds; After all rounds of training are completed, save the global optimal model parameters.
[0040] S4, model reasoning: The hyperspectral data in the test set is input into the trained hyperspectral image classification model, the predicted category is obtained by reasoning, and the obtained predicted category is evaluated by the evaluation index.
[0041] The specific process of inference to obtain the predicted category is: The hyperspectral data to be predicted in the test set are input into the trained hyperspectral image classification model batch by batch. The spatial-spectral joint representation of the hyperspectral data to be predicted is extracted through the sliding window group encoding module. The feature map is generated after layer-by-layer feature fusion by the agent attention encoder ATE. The feature map is globally averaged pooled to obtain the global feature vector, and then the global feature vector is transformed and classified using a two-layer MLP to generate a classification probability vector. ,pass The function will It is converted into a category probability distribution, and finally the largest category index in the category probability distribution is used as the predicted category of the pixel to generate a classification prediction map.
[0042] Evaluation indicators include overall accuracy , average accuracy and coefficient, The coefficient is the evaluation index of the hyperspectral classification model. and The calculation formula is as follows: ; ; in represents the total number of categories, represents the number of correctly classified samples, represents the total number of samples, Indicates The number of correctly classified samples in the class, Indicates The total number of samples in the class.
[0043]
[0044] Table 1 shows the accuracy of the method of the present invention in the Pavia University dataset. Transformer (VIT) is the baseline model of this model, SpectralFormer is a model improved based on the VIT architecture, and Ours is a hyperspectral image classification model provided by the present invention.
[0045] Using overall accuracy , average accuracy and The coefficient is used to evaluate the accuracy of hyperspectral image classification. It can be seen from Table 1 that the classification result of the model of the present invention is significantly improved compared with the Baseline.
Claims
1. A hyperspectral image classification method based on proxy attention and VIT architecture, characterized in that: The following steps are involved: S1, data preprocessing: preprocess the hyperspectral image data and divide it into training set and test set; S2, model building: building a hyperspectral image classification model based on agent attention and VIT architecture; The specific process of step S2 is as follows: Step S21: Construct a sliding window group coding module based on the sliding window mechanism and the proxy pixel block to encode the original hyperspectral data into an initial feature coding matrix , Contains spatial-spectral joint characterization of hyperspectral data; Step S22: Construct an agent attention encoder ATE based on agent attention, perform hierarchical feature extraction through L layers of cascaded agent attention encoder ATE, and realize the initial feature encoding matrix The spatial-spectral information is fused and enhanced layer by layer to obtain the feature map; Step S23: construct a classification head consisting of an average pooling layer and a multi-layer perceptron MLP layer, interpret the feature map, predict the category to which the original pixel belongs, and obtain a classification prediction map; S3, model training: using the training set to train the hyperspectral image classification model constructed in step S2, calculating the loss value between the predicted result and the actual label during the training process, optimizing the model parameters in an iterative manner until the model converges, and recording and storing the optimal model parameters for the training; S4, model reasoning: The hyperspectral data in the test set is input into the trained hyperspectral image classification model, the predicted category is obtained by reasoning, and the obtained predicted category is evaluated by the evaluation index.
2. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 1 is characterized in that: The specific process of step S1 is as follows: S11, data set division: randomly divide the original hyperspectral data set D into training sets according to the proportion and test set , where the training set accounts for , the test set accounts for ; S12, proxy pixel block sampling: in pixel coordinates Centered on the pixel, the processed data set is cropped pixel by pixel, with a window size of , get the sample set of proxy pixel blocks centered on the pixel , where each proxy pixel block , The size is , is the number of spectral channels, It is expressed as: ; in, represents the trimming operation function, Represented by coordinates The proxy pixel block is centered; S13, normalization and data enhancement: First, the sampled proxy pixel blocks are Normalize it and The value of is standardized to the interval [0,1], and the processing method is: ; in, is the normalized data, and are the minimum and maximum values in the proxy pixel block, respectively; Then perform data enhancement operations; randomly rotate and flip the Perform data enhancement to obtain a new proxy pixel block sample set , where each new proxy pixel block sample , The size is , is the number of spectral channels, It is expressed as: ; in, represents the data enhancement function; S14, construct a set of training sample pairs: New proxy pixel patch samples Match the corresponding true category label , Indicates the serial number of the new proxy pixel block sample, and establishes the training sample pair ,in , is the total number of categories, and finally the training sample pair set is obtained , It is expressed as: ; in for The total number of samples in .
3. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 2 is characterized in that: The specific process of step S21 is as follows: S211, sliding window grouping; First, Leveled, by becomes The size is vector; then the vector filling operation is performed on the flattened vector group to meet the requirement of grouping with each vector as the center when performing sliding window grouping; finally, the sliding window grouping operation is performed, and the window size is vectors, the sliding step size is 1, and we get groups, each containing indivual The whole grouping process is expressed as: ; in, Represents the result after grouping. Indicates the leveling operation. represents the filling operation, Represents a sliding window grouping operation; S212, group coding: Encode each group and finally obtain the initial feature coding matrix ; First, linear projection is used to indivual The grouping of vectors is projected onto a In the vector space of size, Represents the feature dimension and adds position encoding to get a size of The initial feature encoding matrix , the projection process is expressed as: ; in, represents a linear projection operation, Represents positional encoding; positional encoding The calculation method is: ; ; in, Represents the position index, Represents a dimension index.
4. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 3 is characterized in that: The specific process of step S22 is as follows: S221, build information aggregation module IAM, aggregation The information of size of Get a size of The proxy tag matrix , represents the degree of information aggregation of the proxy labeling matrix; S222, construct agent attention; The calculation of agent attention is divided into two stages: First, aggregate information and calculate proxy features , the proxy labeling matrix As a query, from the key Sum Aggregate information in ; specifically, calculate the proxy label matrix With key , and perform weighted summation of the values to obtain the proxy feature , the whole process is expressed as: ; ; ; ; in, is the query matrix, is the key matrix, is the value matrix, represents matrix multiplication, , , are the weight matrices for query, key, and value, respectively. represents the dimension of the key vector, represents the matrix transposition operation, and softmax is the activation function; Then, broadcast the information and get the final output of the agent attention ; Aggregate the proxy features As new values, the proxy label matrix As a key, As a query, a second self-attention calculation is performed; this process broadcasts the global information in the proxy features back to each query, and obtains the final output of the proxy attention : - ; in, - represents the agent attention calculation operation, Represents a matrix transpose operation; S223, construct a multi-head proxy attention mechanism; use the multi-head proxy attention mechanism, that is, parallelly calculate the outputs of multiple proxy self-attention heads, and then concatenate and linearly map the outputs of multiple proxy self-attention heads; assuming there are h proxy attention heads, the calculation of the mth proxy attention head is , the multi-head output is - , the calculation process is expressed as: - ; - ; in, and are all learnable parameter matrices, Represents a matrix concatenation operation; S224, constructing a feed-forward neural network FFN; The feedforward neural network consists of two linear transformation layers and an activation function. Assuming the input is Z, Z is output through the feedforward neural network. for: ; in, is the activation function, and are all weight matrices of FFN, , All are bias terms; S225, constructing an L-layer cascaded agent attention encoder ATE; ATE includes an information aggregation module IAM, a multi-head agent attention mechanism and a feed-forward neural network FFN; The structure of each layer of ATE is exactly the same, the initial feature encoding matrix The proxy label matrix is obtained by aggregation through the information aggregation module IAM ; The four matrices are input into the multi-head agent attention mechanism together; the output of the multi-head agent attention mechanism is then passed to the feed-forward neural network FFN after residual connection and normalization; the output of FFN is residually connected again to obtain the first Feature map of the layer , ,in is the number of cascaded layers of ATE.
5. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 4 is characterized in that: The specific operation of step S221 is as follows: S2211, the initial feature encoding matrix Reshape into a three-dimensional tensor , , , Represents the dimensions of the matrix and performs preliminary downsampling through an adaptive average pooling operation: ; in, is the feature tensor after downsampling, , represents the adaptive average pooling operation; S2212, design a multi-scale convolution layer, which contains k two-dimensional convolution operations with different kernel sizes. The size of each convolution kernel is , is an adjustable positive integer, which is used to adjust the feature tensor after downsampling Perform parallel convolution operations to obtain the output feature maps of k convolution branches: ; in, , br means the convolution branches; For the The output feature map of the convolution branch, ; For the convolution kernel parameters; Represents parallel convolution operation; S2213, introduce the channel attention mechanism and calculate the attention weight of each channel: ; in, represents the global average pooling operation; and is the fully connected layer parameter, , , is the dimensionality reduction factor; express Activation function; express Activation function; is the channel attention weight, ; S2214, apply the channel attention weights to the output feature maps of each convolution branch, and add the results to obtain the fused feature map G: ; in, Represents a broadcast multiplication operation on the channel dimension; ; S2215, fusion feature map Reshape into a proxy-labeled matrix , : ; in, Represents a reshape operation, which will Fusion feature map Convert to The proxy tag matrix .
6. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 5 is characterized in that: The specific process of step S23 is as follows: First of all, Apply the global average pooling operation to obtain the global feature vector : ; Among them, the global eigenvector ; Then, feature transformation and classification are performed through two layers of MLP to obtain the classification probability vector : ; Among them, the classification probability vector , and is the weight matrix, , is the bias term; Finally, through The function will Convert to category probability distribution and generate classification prediction graph; ; in, is the predicted category probability distribution, the dimension is Class, and the predicted category is the category index with the highest probability.
7. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 6 is characterized in that: In step S3, in the model training stage, the training set data is first input into the constructed hyperspectral image classification model for forward propagation calculation to obtain the predicted category probability distribution; the cross entropy loss function is used to calculate the loss value between the predicted result and the true label, and the loss function Loss is: ; in, is the number of training samples, is the total number of categories, Indicates The samples belong to The true label value of the class, The model predicts The samples belong to The probability value of the class; The Adam optimizer is used for parameter optimization, the initial value of the learning rate is set to a constant, and the cosine annealing strategy is used to adjust the learning rate: ; in is the basic learning rate, is the adjusted learning rate, is the current training round, is the maximum number of training rounds; After all rounds of training are completed, save the global optimal model parameters.
8. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 6 is characterized in that: In step S4, the specific process of inferring the predicted category is: The hyperspectral data to be predicted in the test set are input into the trained hyperspectral image classification model batch by batch. The spatial-spectral joint representation of the hyperspectral data to be predicted is extracted through the sliding window group encoding module. The feature map is generated after layer-by-layer feature fusion by the agent attention encoder ATE. The feature map is globally averaged pooled to obtain the global feature vector, and then the global feature vector is transformed and classified using a two-layer MLP to generate a classification probability vector. ,pass The function will It is converted into a category probability distribution, and finally the largest category index in the category probability distribution is used as the predicted category of the pixel to generate a classification prediction map.
9. The hyperspectral image classification method based on proxy attention and VIT architecture according to claim 6, characterized in that: In step S4, the evaluation index includes the overall accuracy , average accuracy and coefficient, The coefficient is the evaluation index of the hyperspectral classification model. and The calculation formula is as follows: ; ; in represents the total number of categories, represents the number of correctly classified samples, represents the total number of samples, Indicates The number of correctly classified samples in the class, Indicates The total number of samples in the class.
Citation Information
Patent Citations
Coding and decoding structure semantic segmentation model based on position attention mechanism
CN115908793A
Hyperspectral image classification method based on high-order interactive convolutional network
CN118506112A
Medical image segmentation method based on hierarchical agency attention
CN118799569A
Cited By
Hyperspectral image classification method based on dynamic scale and image optimal transmission
CN120375100A
Original mass spectrum data classification method based on multi-channel embedded representation
CN120524271A
Medical image classification method and device based on Top-K strategy
CN120783188A
Equipment residual life prediction method based on broadcast attention mechanism
CN120804498A
Method for identifying spectral and spatial features of hyperspectral image
CN120877130A