Hyperspectral image classification method based on attention mechanism capsule network

Through the capsule network and EM routing mechanism based on attention mechanism, the problems of translation invariance and data dependence in convolutional neural networks in hyperspectral image classification are solved, and efficient and accurate classification effects are achieved.

CN120259711APending Publication Date: 2025-07-04HOHAI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411930429.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing convolutional neural networks have problems such as translation invariance and strong dependence on training data in hyperspectral image classification, resulting in insufficient classification accuracy and stability, especially when the sample size is insufficient.

Method used

A capsule network based on attention mechanism is adopted, combined with the spectral attention module and EM routing mechanism, features are represented through vectors and information transmission is optimized, reducing computational complexity and improving feature extraction efficiency.

Benefits of technology

It improves the accuracy and efficiency of hyperspectral image classification, reduces dependence on large-scale sample data sets, enhances the robustness of the network and sensitivity to spatial position information, and can still be effectively classified especially when sample data is insufficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259711A_ABST
    Figure CN120259711A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral image classification method based on an attention mechanism capsule network. According to the method, a spectral attention module is constructed, spectral information is fully extracted, more prominent features are selected for an input sample, and calculation complexity is reduced; eM routing is adopted in the capsule network, the information transmission process can be effectively optimized, information loss is reduced, the feature transmission efficiency is improved, the robustness of the network is enhanced, and dependence on a large number of samples is reduced. After a contrast experiment with a 3D-CNN and a capsule network is carried out, it is proved that the method is superior to other classification methods in classification precision. According to the method, rich spatial spectrum information is fully utilized, and compared with traditional routing, it is ensured that the low-level capsules accurately transmit the features to the high-level capsules, and attitude information is better utilized. According to the experimental result, the classification accuracy is effectively improved, and the classification effect is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a hyperspectral image classification method based on an attention mechanism capsule network, belonging to the technical field of image processing. Background Art

[0002] With the continuous progress of remote sensing technology, modern technology has been able to easily obtain various high-resolution remote sensing images, which provide us with rich spatial information and can reveal important information in multiple fields such as urban environment, transportation, meteorology, hydrology, and military. Therefore, how to efficiently interpret the content of remote sensing images has become a hot topic in current research, and the requirements for intelligent recognition and classification technologies are also constantly increasing. Especially in hyperspectral image processing, it can not only obtain the spatial information of the target, but also provide detailed data in the spectral dimension, realizing the organic integration of spatial information and spectral information. Therefore, hyperspectral images can be regarded as three-dimensional data. However, there are also some challenges in processing these images. Due to the large number of bands in hyperspectral images and the high correlation between bands, the number of training samples required for the classification task has increased significantly. Insufficient samples may affect the stability and accuracy of the model, resulting in the so-called "curse of dimensionality". In addition, there may be a large amount of redundant information in the images, which also makes data processing more complex.

[0003] In recent years, more and more researchers have applied machine learning and deep learning methods to hyperspectral image classification. Among them, the convolutional neural network (CNN) is one of the most commonly used network structures. CNN can automatically learn to extract robust features from data and has strong feature learning ability, especially suitable for image processing tasks. However, in hyperspectral image classification, CNN also exposes some limitations. First, CNN has translational invariance, which means it is insensitive to the spatial position changes of objects in the image. Although this feature is beneficial in many applications, for hyperspectral image classification, the precise positioning of spatial information is often crucial. Therefore, when processing hyperspectral data with spatial position information sensitivity, the effect of CNN may be limited. Second, CNN usually requires a large amount of training data to improve the generalization ability of the network. However, the data acquisition cost of hyperspectral images is high, and the annotation of training data is often difficult, which limits the application of CNN in hyperspectral image classification.

[0004] On this basis, the proposed capsule network overcomes the limitations of CNN to a certain extent. Different from the scalar neurons in CNN, the capsule network uses vectors to represent the instantiation parameters of objects. The length of the vector reflects the possibility of the existence of the entity, while the direction carries more spatial and orientation information. This representation method makes the capsule network more robust in processing images with spatial and pose transformations. To further improve the information transmission efficiency, the capsule network introduces a dynamic routing mechanism, which can intelligently transfer information from low levels to high levels in the network, thus enhancing the network's ability to model complex data. In addition, the attention mechanism is widely used in neural networks, aiming to extract more effective features by weighting the regions of interest. The attention mechanism can help the network focus on the most discriminative parts in the massive data, thereby improving the efficiency and accuracy of feature extraction. Based on this, this study proposes a method combining the attention mechanism capsule network for hyperspectral image classification. This method can not only effectively retain the key information of hyperspectral images, but also optimize the network training process by reducing the dependence on large-scale sample data sets, thereby improving the accuracy and efficiency of hyperspectral image classification. Summary of the Invention

[0005] To solve the above problems existing in the prior art, the present invention proposes a classification method based on an attention mechanism capsule network, which has good classification effects.

[0006] The present invention provides the following solutions to solve the above problems:

[0007] A hyperspectral image classification method based on an attention mechanism capsule network (capsule network) includes the following steps:

[0008] S1. Obtain a hyperspectral image: Download a hyperspectral image set;

[0009] S2. Use the principal component analysis method (PCA) to reduce the dimension of the normalized data set, and use n to represent the number of principal components retained after dimension reduction;

[0010] S3. Construct a spectral attention module, and obtain the prominent features of the image by subtracting feature maps. Take the mean value of all pixel points on each band as the weight of each band, which reduces the computational complexity compared with directly adopting a self-attention structure;

[0011] S4. Construct a capsule network model based on the attention mechanism. The operation steps of the capsule network are as follows:

[0012] A matrix capsule can capture activations (probabilities) like a neuron, but also captures a pose matrix. A pose matrix defines the translation and rotation of an object, which is equivalent to the change in the perspective of an object. The present invention employs EM routing. The purpose of EM (Expectation Maximization) routing is to group capsules to form a part-whole relationship by using clustering techniques (EM). A higher-level feature is detected by finding the consensus of votes from capsules in the lower layer. A vote v from capsule i to parent capsule j ij can be calculated by multiplying the pose matrix M of capsule i i by a perspective-invariant transformation matrix W ij . The calculation formula is as follows: v ij = M i W ij

[0013] The probability that a capsule i is grouped as a part-whole relationship into capsule j is based on the vote v ij and the proximity to the votes of other capsules (vo 1j ...vo kj ). W ij is learned through a cost function and backpropagation. Even when the perspective changes, the pose matrix and the vote change in a coordinated manner. Therefore, the transformation matrix is the same for any perspective of the object: perspective invariance. For different orientations of the object, we need a set of transformation matrices and a parent capsule.

[0014] When capsules are assigned, EM routing groups capsules to form a higher-level capsule at runtime. It also calculates the assignment probability r ij to quantify the runtime connection between a capsule and its parent capsule. The calculation of the capsule output is different from that of the neurons in a deep network. In EM clustering, we represent data points by a Gaussian distribution. In EM routing, we still model the pose matrix of the parent capsule using a Gaussian model.

[0015] The pose matrix and the activation value of the output capsule are iteratively calculated using EM routing. The EM method alternately calls step E and step M to fit the data points to a mixture Gaussian model. Step E determines the probability r ij for each data point to be assigned to the parent capsule. Step M recalculates the values of the Gaussian model based on r ij . After repeating the iteration multiple times, the final a j is the output of the parent capsule.

[0016]

[0017] The above a and V are the activation value and the vote of the sub-capsule respectively. We initialize the assignment probability r with a uniform distribution ijThat is, initially the child capsules have the same associations as any parent capsule. We call step M to calculate the updated Gaussian model (μ, σ) and the parent activation a j , based on a, V, and the current r ij . Then we call step E to recalculate the assignment probability r based on the new Gaussian model and the new a j ij .

[0018] Details of step M (pseudocode):

[0019]

[0020] In step M, we calculate μ and σ based on the activation a of the child capsules i , the current r ij and the votes V. Step M also recalculates the costs and the activation a of the parent capsules j . β v and β α are trained separately. In our implementation, λ (the reciprocal of the temperature parameter) is increased by 1 after each routing iteration.

[0021] Details of step E (pseudocode):

[0022]

[0023] In step E, we recalculate the assignment probability r based on the new μ, σ, and aj ij . If the vote is closer to μ of the updated Gaussian model, the assignment increases.

[0024] The output of a capsule, including the activation value and the pose matrix, is calculated by EM routing. We use EM routing to calculate the output of the parent capsule based on the transformation matrix W and the activation value and pose matrix of the child capsules. However, in the absence of errors, Matrix Capsules still largely depend on the transformation matrix W trained by backpropagation ij and the parameter β v and β α .

[0025] In EM routing, we quantify the connection between the child capsules and the parent capsules by calculating the assignment probability r ij . This value is important but has a short lifespan. We re-initialize it using a uniform distribution for each data point before EM routing calculation. In any case, whether for training or testing, we use EM routing to calculate the output of the capsules.

[0026] S5. Input the test set, test the proposed network model, output the prediction results, and generate a classification map

[0027] ​Loss function: The calculation formula is as follows: L = ((max(0, m - (a t - a i ))) 2

[0028] Among them, m starts from 0.2 and increases linearly by 0.05 each time until it reaches 0.5.

[0029] The capsule network finally returns the probability value that a pixel belongs to a certain category. The v in S4 ij is the output prediction label result.

[0030] The hyperspectral image classification method based on the attention mechanism capsule network adopted by the present invention has the following advantages compared with the prior art methods.

[0031] The present invention selects a spectral attention module, which can fully extract spectral information and obtain prominent features of the image. Compared with other self-attention structures, it reduces the computational complexity.

[0032] The present invention adopts an innovative method of capsule network in the field of deep learning. In terms of feature extraction, by using vectors to represent features, compared with the way of using scalars to represent features in traditional methods, the capsule network can better capture and utilize the rich spectral and spatial information in hyperspectral images. Vectors can not only express the existence and possibility of features, but also effectively retain spatial relationships and direction information, so that when the network processes hyperspectral data, it can more accurately identify and classify different objects and scenes. The capsule network can learn with less training data and still achieve good classification results. This enables the capsule network to play a greater advantage in hyperspectral image classification, especially when the sample data is relatively insufficient.

[0033] The present invention improves the dynamic routing in the capsule network. The EM routing is selected. Compared with the traditional routing, the EM routing algorithm can effectively optimize the information transmission process, ensure that the low-level capsules accurately transmit features to the high-level capsules, reduce information loss and improve the feature transmission efficiency. In addition, by retaining the spatial position information, the EM routing avoids the loss of spatial structure in the traditional pooling layer, enabling the network to improve the classification accuracy when processing data with rich spatial information such as hyperspectral images. At the same time, the EM routing enhances the robustness of the network, especially more adaptable when dealing with image translation, rotation and other transformations, and can be effectively trained on a smaller dataset, reducing the dependence on a large number of samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a flowchart of an embodiment of the present invention.

[0035] Figure 2It is a three-dimensional structure diagram of a hyperspectral image.

[0036] Figure 3 It is a diagram of the spectral attention module.

[0037] Figure 4 It is the comparative experimental results of 3D Convolution (3D-CNN), Capsule Network, and the method of the present invention on Indian Pines and University of Pavia datasets respectively.

[0038] Figure 5(a) is the false-color map of the Indian Pines dataset and the classification result maps after applying 3D-CNN, Capsule Network, and the method of the present invention respectively; Figure 5(b) is the false-color map of the University of Pavia dataset and the classification result maps after applying 3D-CNN, Capsule Network, and the method of the present invention respectively. Detailed implementation manners

[0039] 1. Simulation environment:

[0040] For the platform of the embodiment of the present invention, the hardware: the processor is Intel Core i5-12400F, the graphics card is 3060Ti, the memory is 32G, and the software: Windows10 system, PyCharm2022.

[0041] 2. Simulation implementation

[0042] The first dataset selected by the present invention is the Indian Pines dataset, which is acquired by the AVIRIS sensor at the Indian Pine test site in northwestern Indiana, including 224 bands, with a wavelength range of 400 - 2500nm. After removing the bands covering the water absorption area, 200 bands remain, and the ground reference objects include 16 categories. All hyperspectral images are divided into a training set and a test set. The number of samples in the training set is 300, and the remaining samples are used for testing, and the selection method is random. The principal component analysis method is used for the input hyperspectral images, and 30 components are selected, with a single pixel size of 25×25×1. More prominent image information is obtained through the spectral attention module. The first layer is a convolutional layer with 64 1×1×1 kernels with a stride of 1, the second layer is the primary capsule layer with 32 3×3×1 kernels with a stride of 2 and 32 4×4 pose matrices, the third layer is the digital capsule layer with 16 3×3×1 kernels with a stride of 2 and 16 4×4 pose matrices, and the last layer is a fully connected layer. Through the above process, the test dataset is classified to obtain the classification result.

[0043] The second dataset selected in this invention is the University of Pavia dataset, which is provided by the University of Pavia in Italy. It includes 103 bands with a wavelength range from 0.43μm to 0.86μm, and the ground reference objects include 9 categories. All hyperspectral images are divided into a training set and a test set. The number of samples in the training set is 300, and the remaining samples are used for testing, and the selection method is random. The principal component analysis method is used for the input hyperspectral images, and 30 components are selected, with the size of a single pixel being 19×19×1. More prominent image information is obtained through the spectral attention module. The first layer is a convolutional layer with 64 1×1×1 kernels with a stride of 1, the second layer is a primary capsule layer with 32 3×3×1 kernels with a stride of 2 and 32 4×4 pose matrices, the third layer is a digital capsule layer with 16 3×3×1 kernels with a stride of 2 and 16 4×4 pose matrices, and the last layer is a fully connected layer. Through the above process, the test dataset is classified to obtain the classification results.

Claims

1. Hyperspectral image classification method based on attention mechanism capsule network, and the features of this method include the following steps: S1: Obtain hyperspectral images, select two datasets, Indian Pines and University of Pavia, divide them into training sets and test sets, with the number of training samples both being 300, and the selection method being random. S2: Use the principal component analysis method to preprocess the normalized datasets. After performing dimensionality reduction on the datasets through principal component analysis, 30 components are selected from each. The size of a single pixel in the former dataset is 25×25×1, and the size of a single pixel in the latter dataset is 19×19×1. S3: Construct a spectral attention module, obtain the prominent features of the image by subtracting feature maps, take the mean value of all pixel points on each band as the weight of each band, and compared with directly adopting the self-attention structure, the computational complexity is reduced. S4: Construct a capsule network model based on the attention mechanism, input the training set to learn and obtain more advanced features of the image. S5: Input the test set, test the proposed network model, output the prediction results, and generate a classification map.

2. The spectral attention module according to claim 1, wherein It first uses three convolutions as feature extractors for the input data respectively to output feature maps. The three feature maps will all pass through the ReLU activation function and BN normalization, and then calculate the difference between the first two feature maps, subtract the corresponding elements of the feature maps. As the network is trained, the difference between the two feature maps will become more obvious, thus highlighting the important information of the feature maps. Then, a convolution with a scale of 1 (1x1) is used to increase the non-linear features of the difference feature map, and the ReLU activation function is used to convert the values in the difference feature map into positive values. Calculate the mean value of the pixel points on each band of the difference feature map as the weight of the corresponding band, and use the Softmax activation function for normalization to obtain the spectral weight matrix. Finally, multiply the corresponding elements of the spectral weight matrix by the third feature map to obtain the spectral weight feature map, and in order to make full use of the extracted shallow spectral information, add the spectral weight feature map to the input data.

3. The capsule network model according to claim 1, wherein The selected dynamic routing is EM routing, which can effectively optimize the information transmission process, ensure that low-level capsules accurately transmit features to high-level capsules, reduce information loss and improve the feature transmission efficiency. In addition, EM routing retains spatial position information and avoids the loss of spatial structure by traditional pooling layers, enabling the network to improve the classification accuracy when processing data with rich spatial information such as hyperspectral images. At the same time, EM routing enhances the robustness of the network, especially being more adaptable when dealing with image translations, rotations and other transformations, and can be effectively trained on smaller datasets, reducing the dependence on a large number of samples.

Citation Information

Cited By

  • Attention mechanism and capsule layer fused multi-class crop cultivation area classification method

    CN120932113A