Hyperspectral image classification model pre-training method and classification device

By mapping hyperspectral image samples into uniaxial feature vectors and performing long-range feature analysis, the problems of insufficient data volume and difficulty in labeling in hyperspectral image classification are solved, the classification ability and generalization performance of the model are improved, and the computational complexity is reduced.

CN120298733APending Publication Date: 2025-07-11SHANDONG WOMENS UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311470896.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing hyperspectral image classification model has high requirements for training data volume, high labeling cost, and difficult training of small sample data. The complexity of the self-attention calculation of the Transformer architecture limits its generalization performance in hyperspectral image classification.

Method used

Using the pre-training method of shifted window mask spectroscopy, the hyperspectral image samples are mapped into a uniaxial feature vector, and long-range feature analysis is performed through self-attention calculation in the logic window, and the weight and bias are updated using a linear reconstructor. An efficient uniaxial window hierarchical calculation strategy is designed to avoid the secondary complexity of self-attention calculation.

Benefits of technology

It effectively solves the problems of insufficient data and difficulty in labeling in hyperspectral image classification, improves the classification ability and generalization performance of the model, reduces the computational complexity, and improves the computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298733A_ABST
    Figure CN120298733A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral image classification model pre-training method and a classification device. The pre-training method comprises the following steps of: mapping a hyperspectral image sample into a uniaxial feature vector and then randomly covering the uniaxial feature vector; performing long-range feature analysis on the covered uniaxial feature vector to obtain a global feature; reconstructing a hyperspectral image according to the global features, calculating the loss of the global features and the reconstructed hyperspectral image, and updating the weight and bias in long-range feature analysis according to the loss; and migrating the updated weight and bias to the hyperspectral image classification model. According to the method, the problems of limited sample data and difficulty in labeling can be solved, the limitation of a private data set can be broken through, and the classification capability and generalization performance of the hyperspectral image are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral image classification, and more particularly to a pre-training method and a classification device for a hyperspectral image classification model. Background Art

[0002] Hyperspectral images are a further extension of multispectral remote sensing technology. Compared with traditional color or infrared remote sensing images, hyperspectral images have higher spectral resolution and richer spectral information, and their rich information is helpful for object classification, material detection, environmental analysis, etc. They are applied in the fields of environmental monitoring, agriculture, geological exploration, medical diagnosis, etc.

[0003] The goal of the hyperspectral image classification task is to classify each pixel in the image and determine the type of ground object it contains. Common ground object categories include vegetation, water bodies, urban buildings, etc. Each pixel in the image has a corresponding high-dimensional spectral vector.

[0004] In recent years, a large number of attempts have been made in hyperspectral image classification using deep learning; among them, the vision transformer (ViT) model with self-attention mechanism as the core has been favored by researchers. Compared with the traditional deep learning network CNN, ViT can better model the global context information of hyperspectral images, thus obtaining better classification results.

[0005] However, ViT has high requirements for the amount of training data and is restricted by the quadratic complexity of the self-attention calculation of the transformer architecture. The current difficulty is to solve the problems of high hyperspectral data annotation cost and difficult training of small sample data, so as to improve the classification ability and generalization performance of the ViT model in the case of limited hyperspectral image samples. Summary of the Invention

[0006] In view of this, to at least partially solve the above problems, the present invention provides a pre-training method and a classification device for hyperspectral image classification based on shifted window masked spectroscopy, aiming to obtain excellent performance in hyperspectral image classification through pre-training on a small hyperspectral sample dataset.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] On the one hand, the present invention discloses a pre-training method for a hyperspectral image classification model, including the following steps:

[0009] Map the hyperspectral image samples into single-axis feature vectors and then perform random masking;

[0010] Perform long-range feature analysis on the covered unidirectional feature vector, including dividing the unidirectional feature vector into logical windows and obtaining global features through self-attention calculation within the logical windows;

[0011] According to the global features, reconstruct the hyperspectral image, calculate the loss between the global features and the reconstructed hyperspectral image, and update the weights and biases in the long-range feature analysis according to the loss;

[0012] Transfer the updated weights and biases to the hyperspectral image classification model.

[0013] Optionally, the hyperspectral image sample is data collected together for c spectral channels in each patch.

[0014] Optionally, use a unidirectional continuous cross-correlation layer to map the sample of the hyperspectral image to a unidirectional feature vector according to the following formula;

[0015] l = GELU(LN(x spa +x spe ))

[0016] In the formula, x spa is the learning result of the spatial branch of the unidirectional continuous cross-correlation layer, and the expression is: x spa =(conv spa (x')); where the kernel stride of conv spa is (p,p), and p represents patch_size;

[0017] x spe is the learning result of the spatial-spectral branch of the unidirectional continuous cross-correlation layer, and the expression is: x spe =conv spe (x')+conv spe (x' T ), where the kernel strides of conv spe are both (1,p 2 ), and p represents patch_size.

[0018] Optionally, the height of the logical window is always 1, and the length of the logical window depends on the number of divided patches.

[0019] Optionally, use a linear reconstructor to reconstruct the hyperspectral image, including:

[0020] Map the features to the original features through a linear rollback layer and map the size to the original size through a linear decoding layer.

[0021] Optionally, the hyperspectral image classification model has the same long-range feature analysis process as the training method, and the input data is the uniaxial feature vector corresponding to the hyperspectral image sample, and the output is the classification result of the hyperspectral image sample.

[0022] On the other hand, the present invention discloses a hyperspectral image classification device, wherein a hyperspectral image classification model is built in the device.

[0023] The hyperspectral classification model has a pre-training branch.

[0024] The pre-training branch sequentially includes

[0025] a uniaxial continuous cross-correlation layer for mapping the patch samples divided according to the hyperspectral image into uniaxial feature vectors;

[0026] a window marking unit for marking the uniaxial feature vectors to achieve random masking;

[0027] a band shift transformer for performing long-range feature analysis on the masked uniaxial feature vectors to obtain global features; and

[0028] a linear reconstructor for reconstructing the hyperspectral image according to the global features; and calculating the loss between the global features and the reconstructed hyperspectral image, and updating the weights and biases of the band shift transformer according to the loss.

[0029] Optionally, the hyperspectral image classification model shares the uniaxial continuous cross-correlation layer and the band shift transformer with the pre-training branch, wherein the uniaxial continuous cross-correlation layer is connected to the band shift transformer, and the other end of the band shift transformer is connected to a linear classifier.

[0030] The linear classifier is used to directly output the classification result of the hyperspectral image sample.

[0031] Optionally, the band shift transformer includes 4 stages, and in each stage, 2 SwinTransformer blocks and 1 patch merging are sequentially included in the data transmission direction.

[0032] Through the above technical solutions, it can be seen that the pre-training method and classification device of the hyperspectral image classification model disclosed by the present invention can effectively solve the problems of limited sample data and difficult annotation by fully learning the potential features within the spectrum and the connection information between spectra through the pre-training process, and can break through the limitations of private data sets, thereby improving the classification ability and generalization performance of hyperspectral images.

[0033] In addition, mapping the image samples to a single-axis vector can avoid the quadratic complexity inherent in standard self-attention, thereby improving the computational efficiency.

[0034] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structure particularly pointed out in the written description, claims, and drawings. Brief Description of the Drawings

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0036] Figure 1 It is a flowchart of the pre-training method for the hyperspectral image classification model of the present invention;

[0037] Figure 2 It is a schematic diagram of the mapping process of the hyperspectral image samples of the present invention;

[0038] Figure 3 It is a comparison chart of the feature learning through windows of the method of the present invention and other algorithms; where (a) is the ViT algorithm, (b) is the Swin-Transformer algorithm, and (c) is the method of the present invention;

[0039] Figure 4 It is a schematic diagram of the network architecture in the classification device of the present invention;

[0040] Figure 5 It is the spectral curves of all land classes in the Indian Pines dataset of the present invention, and one sample is selected for each class;

[0041] Figure 6 It is the spectral curves of all land classes in the Pavia University dataset of the present invention, and one sample is selected for each class;

[0042] Figure 7 It is the spectral curves of all land classes in the Houston2013 dataset of the present invention, and one sample is selected for each class;

[0043] Figure 8 It is a comparison chart of the classification performance on the baseline network and the fine-tuning classification performance of the Swin-MSP network for different datasets of the present invention;

[0044] Figure 9: It is the vector diagram of the data set IP of the present invention, where (a) is the original vector diagram, (b) is the masked vector diagram, and (c) is the vector diagram after reconstructing the spectrum. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] In order to solve the problems of high cost of hyperspectral data annotation and difficulty in training small sample data, the embodiment of the present invention discloses a hyperspectral image classification model pre-training method. At the same time, during the pre-training and hyperspectral image classification process, the image samples are mapped into single-axis feature vectors, which can effectively avoid the quadratic complexity of self-attention calculation of the traditional transformer architecture.

[0047] Embodiment 1

[0048] This application first provides a hyperspectral image classification model pre-training method, such as Figure 1 As shown, the following steps are included:

[0049] The hyperspectral image samples are mapped into single-axis feature vectors and then randomly masked;

[0050] Performing long-range feature analysis on the masked single-axis feature vector, including dividing the strip feature vector into logical windows, and obtaining global features through self-attention calculation within the logical windows;

[0051] Reconstructing a hyperspectral image according to the global features, calculating the loss of the global features and the reconstructed hyperspectral image, and updating the weights and biases in the long-range feature analysis according to the loss;

[0052] The updated weights and biases are transferred to the hyperspectral image classification model.

[0053] For hyperspectral image data, each pixel corresponds to a land category, and there is a potential relationship between each pixel and its surrounding pixels. In order to explore this potential relationship, this application treats the pixel to be classified and its surrounding pixels with a radius of r as a patch, that is,

[0054] patch_size=h patch =w patch =2r+1

[0055] For edge pixels, the filling method can be used to meet the requirements.

[0056] If the input image is given as The c spectral channels of each patch are collected together as a sample, so the size of each 3D sample is x'∈(p×p×c) (p represents patch_size).

[0057] Subsequently, the samples were flattened and arranged sequentially, e.g. Figure 2 As shown, we get x'∈p×(p·c), and use the single-axis continuous cross-correlation layer spatial branch to perform a two-dimensional convolution on the sample with a kernel step size of (p, p), and we get:

[0058] x spa =(conv spa (x'))

[0059] where x spa ∈(1×p×dim).

[0060] On the other hand, the single-axis continuous cross-correlation layer spatial spectrum branch performs the following operations:

[0061] x spe =conv spe (x')+conv spe (x' T )

[0062] in, conv spe The kernel step size is (1,p 2 ): Further, x spe Flatten to achieve multi-spectral feature fusion, that is, x spe ∈(1×p×dim), the results of the last two branches are added:

[0063] l = GELU(LN(x spa +x spe ))

[0064] In this embodiment, l is the single-axis feature vector obtained by mapping, which is randomly masked and enters the next long-range feature analysis process, including window division of the single-axis feature vector, and obtaining the global feature through self-attention calculation within the window.

[0065] In this process, this embodiment randomly masks the spectrum. In the pre-training process, the classic Swin-MAE increases the minimum unit of window division from patch to a logical window composed of multiple patches. This application redesigns the masking strategy for hyperspectral images. In order not to destroy the continuity of information between spectra, this embodiment chooses to use patch-level random masking instead of window masking to encourage the model to learn the potential features in the spectrum and the connection information between spectra.

[0066] In addition, this method fits the data into a single axis, enabling efficient utilization of the sliding window mechanism. Specifically: First, the spectral dimension is flattened, and window partitioning is performed on this single-axis vector. The height of the logical window will always be 1, and the length of the logical window depends on the number of patches divided (e.g., Figure 3 c is 4 patches). Subsequently, the feature range is expanded through block fusion in the spectral dimension, which reflects long-range feature analysis in the spectral dimension on the hyperspectral image, and thus achieves feature reading from local to global. To clearly demonstrate the advantages of this application, a comparison is made with the feature reading methods provided in ViT( Figure 3 a) and Swin-T ransformer( Figure 3 b). Through the comparison, it can be seen that

[0067] the method provided by this application has the following advantages:

[0068] 1) The limited attention within each window can improve the modeling ability of local spectral features;

[0069] 2) The shifted window allows self-attention to span adjacent spectral bands, comprehensively learning inter-band correlations;

[0070] 3) Block fusion gradually expands the capture range, gradually grasping local to global features;

[0071] 4) The single-axis vector avoids the quadratic complexity inherent in standard self-attention, thereby improving computational efficiency.

[0072] Furthermore, based on the global features, the hyperspectral image is reconstructed, the loss between the global features and the reconstructed hyperspectral image is calculated, and the weights and biases in the long-range feature analysis are updated according to the loss;

[0073] In this application, after the masked feature vector undergoes long-range feature analysis, its size and dimension will change. Therefore, a linear reconstructor composed of two linear layers is used to remap the features back to the original dimension. Specifically, the features are mapped back to the original features through a linear rollback layer, and the size is mapped back to the original size through a linear decoding layer,

[0074] and further, the mean absolute error is calculated with the original input to evaluate the quality of the reconstructed pixels, obtaining the loss between the global features and the reconstructed hyperspectral image.

[0075] The present invention can force the encoder to focus its capabilities on modeling the masked tokens through a simple linear reconstructor, rather than leaving this task to the decoder.

[0076] When the predetermined number of iterations is reached or the loss threshold is satisfied, the update of the parameters is stopped. At this time, the finally updated weights and biases are migrated to the hyperspectral image classification model.

[0077] It should be noted that the hyperspectral image classification model of the present invention has the same long-range feature analysis process as the training method, and the input data is the single-axis feature vector corresponding to the hyperspectral image sample, and the output is the classification result of the hyperspectral image sample.

[0078] Generally speaking, the present invention can solve the deficiencies in the prior art, including:

[0079] 1) Insufficient pre-training data

[0080] Previous supervised pre-training models such as ImageNet pre-training rely heavily on a large amount of manually labeled data, which limits the transfer ability of the model. The present invention proposes self-supervised pre-training on unlabeled data through sample random occlusion, and uses a large amount of unlabeled image data without manual annotation, thus solving the problem of insufficient data volume;

[0081] 2) Weak semantic consistency

[0082] The image representations learned by early self-supervised methods are different from the semantic information mastered by the human visual system, and the semantic consistency is insufficient. This application draws on psychological research and trains by occluding spectral semantic regions to make the model learn representations that are more sensitive to semantic information;

[0083] 3) Insufficient utilization of Transformer encoding ability

[0084] Vision Transformer introduces stronger global modeling capabilities, but the early Transformer pre-training methods have poor effects. Therefore, based on Swin-MAE, the present invention fully exploits its powerful encoding-decoding ability through occlusion training to make its fusion of local and global information better;

[0085] 4) Computational optimization

[0086] Due to the inherent quadratic nature of attention calculation in previous methods, it is difficult to achieve a qualitative improvement in computational efficiency. The efficient single-axis window hierarchical calculation strategy designed by the present invention greatly reduces the complexity of attention calculation and gets rid of the computational efficiency dilemma of the attention-based model.

[0087] Embodiment 2

[0088] The present invention further discloses a hyperspectral image classification device with a built-in hyperspectral image classification model, wherein the hyperspectral classification model has a pre-training branch; as Figure 4 shown;

[0089] The pre-training branch sequentially includes:

[0090] A uniaxial continuous cross-correlation layer for mapping patch samples divided from the hyperspectral image into uniaxial feature vectors;

[0091] A window marking unit for marking the uniaxial feature vectors to achieve random masking;

[0092] A spectral band shifting transformer for performing long-range feature analysis on the masked uniaxial feature vectors to obtain global features; and

[0093] A linear reconstructor for reconstructing the hyperspectral image according to the global features; and calculating the loss between the global features and the reconstructed hyperspectral image, and updating the weights and biases of the spectral band shifting transformer according to the loss.

[0094] In the present invention, the Swin-MAE architecture is improved to be applicable to hyperspectral images. The overall training network architecture mainly consists of a uniaxial continuous cross-correlation layer (UC3L), a spectral band shifting transformer (SFBT), and a decoder.

[0095] The uniaxial continuous cross-correlation layer is used to map the hyperspectral image into the input of the SFBT. The SFBT is improved from the Swin-Transformer encoder for hyperspectral images, and the decoder uses different schemes in pre-training and fine-tuning.

[0096] Swin Transformer divides the picture into non-overlapping blocks, and then combines the blocks into logically windows for window self-attention calculation. That is, the original Swin-Transformer limits the self-attention calculation to a constant local two-dimensional window, and expands the feature capture range through continuous block fusion (such as Figure 3 b). In the spectral band shifting transformer in the present invention, it is improved to be applicable to hyperspectral images. Specifically, a scheme for calculating uniaxial shifted window self-attention in the spectral dimension is designed, including: dividing the uniaxial feature vectors into windows, and obtaining global features through self-attention calculation within the windows.

[0097] In this network, the same SFBT is used for pre-training, fine-tuning, and image classification. Specifically, in the pre-training process, the input data of the SFBT is the uniaxial feature vectors with part of the spectrum masked by a window marking unit (Window masking), and the decoder is a linear reconstructor for reconstructing the masked spectral features.

[0098] Pre-training calculates the L1 loss (mean absolute error, MAE) between the output of the linear reconstructor and the original input to evaluate the gap between the reconstructed pixels and the original pixels;

[0099] During the fine-tuning process, SFBT uses the weights and biases transferred from the pre-training stage. The decoder is replaced by a linear classifier, which is responsible for outputting the land category corresponding to the pixel. The present invention uses a simple linear classifier to map the output of the encoder to the corresponding land category for classification.

[0100] This design in the present invention will force the encoder to learn a large number of latent representations during the pre-training stage, and perform a good initialization for the downstream classification task.

[0101] Furthermore, the hyperspectral image classification model shares a single-axis continuous cross-correlation layer and a band-shift transformer with the pre-training branch. Among them, the single-axis continuous cross-correlation layer is connected to the band-shift transformer, and the other end of the band-shift transformer is connected to a linear classifier;

[0102] The linear classifier is used to directly output the classification result of the hyperspectral image sample.

[0103] At the same time, the band-shift transformer includes 4 stages. In order to adapt to the improved architecture process, in each stage, 2 SwinTransformer blocks and 1 patch merging are included in sequence according to the data transmission direction.

[0104] In order to verify the effectiveness of the Swin-MSP pre-training model proposed by the present invention, three groups of ablation experiments were carried out on three public datasets. In the first group of experiments, grid search was performed to determine the optimal pre-training learning rate scaling value for each individual dataset. In the second group of experiments, the influence of different masking ratios on the reconstructed spectral features during pre-training was systematically explored. In the third group of experiments, the influence of different network depths and the number of multi-head self-attention heads on the pre-training of masked spectral features was evaluated.

[0105] At the same time, classical deep learning networks such as 1D-CNN, 2D-CNN, RNN, miniGCN, the original VisionTransformer (ViT) network, and two networks based on (ViT) were used for comparison. The benefits can be directly attributed to the effectiveness of the pre-training method proposed in this application.

[0106] The networks based on (ViT) include:

[0107] MAEST: A novel Masked Auto-Encoding Spectral Transformer (MAEST) that incorporates two distinct collaborative branches: 1) a reconstruction path that dynamically discovers the most robust encoding features based on a masked auto-encoding strategy; 2) a classification path that embeds these features into a transformer network to classify the data, with a focus on features that can better reconstruct the input. Different from other existing models, given the complexity of the HS remote sensing image field mentioned above, this novel design aims to learn fine-grained transformer features.

[0108] SpectralFormer (SF): A novel backbone network. In addition to the banded representation in classical transformers, SF can also learn spectral local sequence information from adjacent bands of HS images, thereby generating grouped spectral embeddings. To reduce the possibility of losing valuable information during hierarchical propagation, a cross-layer skip connection is designed to adaptively learn and fuse the "soft" residuals across layers, transferring memory-like components from the shallow layers to the deep layers.

[0109] The datasets include:

[0110] A: Indian pines (IP)

[0111] This scene was collected by the AVIRIS sensor at the Indian Pine test site in northwestern Indiana. The pixel resolution is 145×145, containing 224 spectral bands. After removing background pixels, 10,249 pixels are used, and 200 bands are used after removing the strips covering the water absorption area. The fine-tuning data ratio of this dataset is 10%. Among them, the spectral curves of all land classes, such as Figure 5 shown;

[0112] B: Pavia University (PU)

[0113] This is a scene obtained during the flight activities of the ROSIS sensor over Pavia in northern Italy. The pixel resolution is 610×610, containing 103 spectral bands. After removing unclassified background pixels, 42,776 pixels are used, and we select the first 100 bands. The fine-tuning data ratio of this dataset is 1%.

[0114] The spectral curves are as Figure 6 shown;

[0115] C: Houston2013 (HOU)

[0116] Figure 7Acquired by the Innovation, Technology, Research, Excellence and Service (ITRES) Compact Airborne Spectral Imager (CASI)-1500 sensor, this image contains the campus of the University of Houston and its surrounding area in Texas, USA. The pixel resolution is 349x1905 and contains 144 bands with wavelengths ranging from 364 nm to 1046 nm. 15029 pixels are used after removing unclassified background pixels. We select the last 140 bands.

[0117] Experimental setup

[0118] In the present invention, the parameter setting follows the SimMIM network, and L1 loss is used to perform the spectral feature reconstruction task. In addition, the optimizer (AdamW) and learning rate scheduling strategy of Swin Transformer are inherited. For pre-training, all samples in the data set are used for training 500 times. For fine-tuning, the data samples used in the above table are used, initialized with pre-trained weights and biases, and fine-tuned for 100 epochs. Through this controlled experimental setting, the performance of the pre-training method proposed in this application can be directly evaluated.

[0119] It is worth noting that in this work, the focus is on the reconstruction and analysis of spectral features, so the patchsize is set to 5 and the feature mapping dimension is set to 96.

[0120] Experimental Results

[0121] The performance results of the Swin-MSP network proposed in this application are as follows Figure 8 As shown. The performance of training from scratch (baseline performance) and the performance of fine-tuning after pre-training are compared (10% of the data samples are taken for each). The overall accuracy (OA) of the baseline performance stabilizes after 300 epochs of training. In contrast, fine-tuning using pre-trained weights and biases not only exceeds the limitations of traditional training in terms of accuracy, but also requires fewer training cycles. This shows that the model in the present invention is able to learn sufficient potential representations during pre-training, which greatly promotes the performance of downstream classification tasks.

[0122] At the same time, the training sample settings and comparison results of the three data sets are shown in Tables 1-6. By using the same training sample settings as the comparison network, the accuracy of OA: 86.60%, AA: 93.28% (IP data set), OA: 98.49%, AA: 98.34% (PU data set), OA: 94.40%, AA: 95.67% (HOU data set) was achieved on the three data sets. They are 2.45% and 2.31% higher (IP data set), 7.43% and 8.34% (PU data set), and 5.85% and 6.78% (HOU data set) than the highest results of the comparison network, respectively. This is enough to prove the advanced nature of the network proposed in this application.

[0123] Table 1 Land Categories and Data Division of IP Dataset

[0124] Class No. Class Name Total Fine-tune 1 Alfalfa 46 15 2 Corn-notill 1428 50 3 Corn-mintill 830 50 4 Corn 237 50 5 Grass-pasture 483 50 6 Grass-trees 730 50 7 Grass-pasture-mowed 28 15 8 Hay-windrowed 478 50 9 Oats 20 15 10 Soybean-notill 972 50 11 Soybean-mintill 2455 50 12 Soybean-clean 593 50 13 Wheat 205 50 14 Woods 1265 50 15 Buildings-Grass-Trees-Drives 386 50 16 Stone-Steel-Towers 93 50 Total 10249 695

[0125] Table 2 Comparison Results of IP Dataset

[0126]

[0127] Table 3 Land Cover Categories and Data Division of PU Dataset

[0128] Class No. Class Name Total Fine-tune 1 Asphalt 6631 548 2 Meadows 18649 540 3 Gravel 2099 392 4 Trees 3064 524 5 Painted metal sheets 1345 265 6 Bare Soil 5029 532 7 Bitumen 1330 375 8 Self-Blocking Bricks 3682 514 9 Shadows 947 231 Total 42776 3921

[0129] Table 4 Comparison Results of PU Dataset

[0130]

[0131] Table 5 Land Cover Categories and Data Division of HOU Dataset

[0132]

[0133]

[0134] Table 6 Land Cover Categories and Data Division of HOU Dataset

[0135]

[0136] In addition, taking the first dataset as an example, Figure 9 the original vector diagram (a), the covered vector diagram (b), and the vector diagram after reconstructed spectrum (c) of its samples are respectively shown; where the horizontal axis represents the spectral band, and the vertical axis represents the pixels arranged in a flattened manner when the patch size is 5. The coherence of the vertical axis represents the degree of mixing of pixel points, while the coherence of the horizontal axis represents the correlation degree between spectra. It can be seen that the IP samples show complex mixing and spectral correlation, and from the vector diagram of the spectral reconstruction result, it can be seen that the proposed model has excellent reconstruction effects both globally and locally.

[0137] This experiment systematically evaluated the pre-training of the improved Swin-MAE architecture on small-scale hyperspectral datasets, revealing the potential benefits of spectral feature reconstruction pre-training for downstream classification tasks. The results show that the proposed model is very suitable for publicly available small-scale hyperspectral data, which is proved by the strong performance on multiple benchmark datasets. The present invention will provide a new perspective for pre-training strategies and model architectures of computer vision tasks involving small unlabeled hyperspectral image collections with limited data availability.

[0138] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.

[0139] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A pre-training method for a hyperspectral image classification model, characterized in that, Including the following steps: Mapping the hyperspectral image sample into a uniaxial feature vector and then performing random masking; Performing long-range feature analysis on the masked uniaxial feature vector, including dividing the uniaxial feature vector into logical windows and obtaining global features through self-attention calculation within the logical windows; Reconstructing the hyperspectral image according to the global features, calculating the loss between the global features and the reconstructed hyperspectral image, and updating the weights and biases in the long-range feature analysis according to the loss; Transferring the updated weights and biases to the hyperspectral image classification model.

2. A pre-training method for a hyperspectral image classification model according to claim 1, wherein The hyperspectral image sample is the data collected together for c spectral channels in each patch.

3. A pre-training method for a hyperspectral image classification model according to claim 1, characterized in that, Using a uniaxial continuous cross-correlation layer to map the sample of the hyperspectral image into a uniaxial feature vector according to the following formula; l = GELU(LN(x spa + x spe )) where x spa is the learning result of the single-axis continuous cross-correlation layer spatial branch, and the expression is: x spa =(conv spa (x')); where the kernel stride of conv spa is (p, p), and p represents patch_size; x spe is the learning result of the single - axial continuous cross - correlation layer spatial - spectral branch, and the expression is: x spe = conv spe (x') + conv spe (x' T ), where the kernel strides of conv spe are both (1, p 2 ), and p represents patch_size.

4. A pre-training method for a hyperspectral image classification model according to claim 1, wherein The height of the logical window is always 1, and the length of the logical window depends on the number of divided patches.

5. A pre-training method for a hyperspectral image classification model according to claim 1, characterized in that Reconstructing the hyperspectral image using a linear reconstructor, including: Mapping the features into the original features through a linear rollback layer and mapping the size into the original size through a linear decoding layer.

6. A pre-training method for a hyperspectral image classification model according to claim 1, characterized in that, The hyperspectral image classification model has the same long-range feature analysis process as the training method, and the input data is the uniaxial feature vector corresponding to the hyperspectral image sample, and the output is the classification result of the hyperspectral image sample.

7. A hyperspectral image classification device, characterized in that, Built-in hyperspectral image classification model, The hyperspectral classification model has a pre-training branch; The pre-training branch sequentially includes A uniaxial continuous cross-correlation layer for mapping the patch samples divided from the hyperspectral image into a uniaxial feature vector; A window marking unit for annotating the uniaxial feature vector to achieve random masking; A band shift transformer for performing long-range feature analysis on the masked uniaxial feature vector to obtain global features; and A linear reconstructor for reconstructing the hyperspectral image according to the global features; and calculating the loss between the global features and the reconstructed hyperspectral image, and updating the weights and biases of the band shift transformer according to the loss.

8. An apparatus for hyperspectral image classification according to claim 7, characterized in that, The hyperspectral image classification model shares the uniaxial continuous cross-correlation layer and the band shift transformer with the pre-training branch, wherein the uniaxial continuous cross-correlation layer is connected to the band shift transformer, and the other end of the band shift transformer is connected to a linear classifier; The linear classifier is used to directly output the classification result of the hyperspectral image sample.

9. A hyperspectral image classification device according to claim 7 or 8, characterized in that The band shift transformer includes 4 stages, and in each stage, in the data transmission direction, it sequentially includes 2 SwinTransformer blocks and 1 patch merging.

Citation Information

Cited By

  • A hyperspectral image airplane target recognition method based on self-supervised learning

    CN122530998A

  • A hyperspectral image airplane target recognition method based on self-supervised learning

    CN122530998B