Small sample hyperspectral image open set identification method and system based on task adaptation
By constructing a task-adaptive small sample hyperspectral image open set recognition method, using the spectral spatial selection transformer model and multi-head attention mechanism to generate unknown prototypes, the problem of low recognition accuracy of hyperspectral image open set under small sample conditions is solved, and effective recognition of unknown categories of samples is achieved.
Patent Information
- Application Number
- CN202510447850.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art hyperspectral image open-set recognition method has low classification accuracy under small sample conditions, making it difficult to effectively identify samples of unknown categories.
A small sample hyperspectral image open-set recognition method based on task adaptation is constructed, and the characteristics are extracted using the spectral space selection transformer model, and unknown prototypes are generated through multi-head attention mechanism and multi-layer perceptron, and classified in combination with meta-learning tasks.
The accuracy of open set recognition of small sample hyperspectral images is improved, the robustness of the model is enhanced, and the samples of unknown categories can be effectively identified.
Smart Images

Figure CN120355995A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly relates to a small-sample hyperspectral image open-set recognition method and system based on task adaptation. Background Art
[0002] With the rapid development of remote sensing technology, hyperspectral images have been widely used in many fields such as agricultural monitoring, environmental assessment, geological exploration, military reconnaissance, etc. because they can provide rich spectral information. However, in these diverse practical application scenarios, hyperspectral image data presents an open-set characteristic. Specifically, when performing target recognition on hyperspectral images, the test set often contains class samples that have never been encountered in the training stage, which is the open-set recognition problem. In addition, there are many difficulties and challenges in the acquisition and annotation processes of hyperspectral image data, which makes the number of labeled samples available for model training extremely limited in many practical tasks.
[0003] Methods for open-set recognition of hyperspectral images have been widely studied in recent years. In the article 'Few - Shot Hyperspectral Image Classification With Unknown Classes Using Multitask Deep Learning' published by Liu et al. in 《IEEE Transactions on Geoscience and Remote Sensing》 in 2020, a reconstruction-based open-set recognition method for hyperspectral images was proposed. Based on the assumption that the reconstruction error of known samples is less than that of unknown samples, this method uses extreme value theory to fit the reconstruction error of samples in the test session and sets a threshold to identify unknown classes. In this way, the model can effectively identify samples that do not belong to the known classes, thus achieving relatively excellent classification performance. However, this method requires sufficient training samples to achieve good classification performance and it is difficult to obtain satisfactory classification results in the case of small samples. Summary of the Invention
[0004] The purpose of the present invention is to provide a small-sample hyperspectral image open-set recognition method and system based on task adaptation for the deficiencies of the existing technology, so as to solve the problem of low classification accuracy caused by insufficient training samples in open-set recognition of hyperspectral images.
[0005] The purpose of the present invention is achieved by the following technical solutions: The first aspect of the present invention provides a small-sample hyperspectral image open-set recognition method based on task adaptation, including the following steps: Input the hyperspectral image to be classified into a pre-constructed open-set recognition model for hyperspectral images to obtain the hyperspectral image classification result; The construction method of the pre-constructed open-set recognition model for hyperspectral images is as follows: Input the hyperspectral image to be classified into a feature extraction model to extract features; Input the extracted features into an unknown class prototype generation model to generate unknown class prototypes; Classify the hyperspectral image based on the distance between the extracted features and the unknown class prototypes; Among them, the feature extraction model uses a spectral-spatial selective transformer model; the unknown class prototype generation model consists of a multi-head attention mechanism and a multi-layer perceptron; the input end of the unknown class prototype generation model is connected to the output end of the feature extraction model, calculates the known class prototypes based on the output of the feature extraction model, and generates unknown class prototypes after processing the known class prototypes through the multi-layer attention mechanism and the multi-layer perceptron.
[0006] Furthermore, the spectral-spatial selective transformer model includes: A central pixel position embedding module, which is used to fuse the spatial information of the input hyperspectral image to obtain embedded features; A spatial selection transformer module, which is used to extract global features and denoise the embedded features; A central pixel mixing module, which is used to fuse the central pixel features of the spatial selection transformer module to achieve feature extraction.
[0007] Furthermore, the central pixel position embedding module obtains the corresponding position encoding according to the relative distance of each pixel to the central pixel, and the expression of the position encoding is:
[0008] Among them, and are variables representing any indexes of the pixel columns and rows in the input image; represents the column index corresponding to the central pixel; represents the row index corresponding to the central pixel; represents the set of position embeddings; represents the size of the image; The expression of the embedded features is:
[0009] Among them, represents the input image.
[0010] Furthermore, the spatial selection transformer module includes: a spatial selection attention module, a multi-layer perceptron layer, and two normalization layers; a residual connection is adopted before the spatial selection attention module and the multi-layer perceptron layer; the embedded features enter the spatial selection attention module to extract global features after passing through the first normalization layer, and the global features are then input into the second normalization layer and enter the multi-layer perceptron layer to output the final global features.
[0011] Furthermore, the specific process of the spatial selection attention module extracting global features is as follows: The spatial selection attention module converts the input three-dimensional data into a two-dimensional matrix, and then generates a query matrix, a key matrix, and a value matrix through linear projection, specifically as follows:
[0012] Among them, 、 and are linear projection matrices; is the query matrix; is the key matrix; is the value matrix; Calculate the pixel correlation matrix by multiplying the query matrix with the transpose of the key matrix, and only retain the most relevant pixel for each pixel to obtain the most relevant index matrix. The expression of the most relevant index matrix is as follows:
[0013] Among them, represents obtaining the index of the most relevant pixel in the correlation matrix; Use the most relevant index matrix to calculate the most relevant matrices of the key matrix and the value matrix respectively. The calculation expressions are:
[0014] Among them, gather means obtaining the most relevant values in the key matrix and the value matrix according to the index; Obtain the output of the spatial selection attention module by calculating the self-attention of the query matrix, the key matrix, and the value matrix and local context enhancement. The calculation expression of the output matrix is:
[0015] Among them, is the local context enhancement term, which obtains enhanced features through depthwise separable convolution.
[0016] Furthermore, the central pixel mixing module fuses the central pixel features of each spatial selection transformer module. The calculation expression in the fusion process is:
[0017] Among them, is a one-dimensional vector, represents the The central pixel feature of the output feature of a spatial selective attention module.
[0018] Furthermore, input the extracted features into an unknown class prototype generation model to generate unknown class prototypes; specifically: Calculate the known class prototypes based on the output of the feature extraction model. The calculation expression of the known class prototypes is:
[0019] Where, Represents the feature extraction process of the feature extraction model, Represents the Number of features of the Represents the th known class prototype; Based on each known class prototype, calculate the self-attention weights between the known class prototypes; the calculation expression of the self-attention weights is:
[0020] Where, And Are linear projection matrices, Is the feature length of the known class prototypes; Is the known class prototype matrix; Generate class negative prototypes based on the known class prototypes and the self-attention weights. The calculation expression of the class negative prototypes is:
[0021] Where, Represents the softmax function, Is a linear projection matrix; Calculate the average value of the class negative prototypes and input it into a multi-layer perceptron to obtain the unknown class prototypes. The calculation expression of the unknown class prototypes is:
[0022] Where, Represents the class negative prototype generated by the th class.
[0023] Furthermore, the loss function of the pre-constructed hyperspectral image open-set recognition model is:
[0024] Where, Is the cross-entropy loss function, Is the class negative prototype loss function; Furthermore, the calculation expression of the class negative prototype loss function Is:
[0025] Where, Denote the set of known class query set samples of the n-th class, and denote the set of unknown class query set samples.
[0026] A second aspect of the present invention provides a small sample hyperspectral image open set recognition system based on task adaptation, comprising: A feature extraction module that inputs the hyperspectral image to be classified into a feature extraction model to extract features; An unknown class prototype generation module that inputs the extracted features into an unknown class prototype generation model to generate unknown class prototypes; A classification module that classifies the hyperspectral image based on the distance between the extracted features and the unknown class prototypes; Wherein, the feature extraction model adopts a spectral space selection transformer model; the unknown class prototype generation model consists of a multi-head attention mechanism and a multi-layer perceptron; the input end of the unknown class prototype generation model is connected to the output end of the feature extraction model, and the average value of the data features of each category output by the feature extraction model is obtained to obtain known class prototypes.
[0027] The beneficial effects of the present invention are as follows: The present invention discloses a small sample hyperspectral image open set recognition method based on task adaptation. In this method, the constructed hyperspectral image feature extraction model and unknown class prototype generation model use a small amount of data in the source domain dataset and the target domain dataset to construct a meta-learning task, realize the training of the model and learn to generate unknown class prototypes to recognize unknown classes. The unknown class prototypes are generated using the prototypes of each category in the target domain, and classification is performed based on the distance between the features of the hyperspectral image to be classified and each prototype. The feature extraction model adopts a spectral space selection transformer model, which strengthens the utilization of the central pixel information of the input hyperspectral image by using the central pixel position embedding module and the central pixel mixing module, and extracts the global features of the image using the spatial selection transformer module and reduces the influence of irrelevant spatial information. The unknown class generation model consists of a multi-head attention mechanism and a multi-layer perceptron, and uses the known class prototypes to generate unknown class prototypes. By this method, the accuracy of small sample hyperspectral image open set recognition is improved, the robustness of the model is enhanced, and it can be widely applied to fields such as smart agriculture, mineral exploration, environmental detection, and medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is a schematic structural diagram of the small-sample hyperspectral image open-set recognition method in the embodiment of the present invention; Figure 2 It is a schematic block diagram of the central pixel position embedding module in the embodiment of the present invention; Figure 3 It is a schematic structural diagram of the spatial selection attention module in the embodiment of the present invention; Figure 4 It is a schematic structural diagram of the unknown class prototype generation module in the embodiment of the present invention; Figure 5-1 It is a pseudo-color composite map of the University of Pavia dataset in the embodiment of the present invention; Figure 5-2 It is the ground truth map of the University of Pavia dataset in the embodiment of the present invention; Figure 5-3 It is a simulation result map of classifying the University of Pavia using the existing MDL4OW method in the embodiment of the present invention; Figure 5-4 It is a simulation result map of classifying the University of Pavia in the embodiment of the present invention. Detailed implementation manners
[0030] In order to make the objectives and technical solutions of the present invention clearer and easier to understand, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] The concept of the present invention is to provide a task-adaptive small-sample hyperspectral image open-set recognition method and system. The method includes: constructing a hyperspectral image feature extraction model and an unknown class prototype generation model, and using a small amount of data in the source domain dataset and the target domain dataset to construct a meta-learning task to train the model and learn to generate unknown class prototypes to identify unknown class prototypes.
[0032] In the test stage, the hyperspectral image to be classified and a small amount of training data in the target domain dataset are input into the trained feature extraction model. The features of the training data are used to calculate various prototypes and generate unknown class prototypes, and classification is performed based on the distances between the features of the hyperspectral image to be classified and each prototype.
[0033] The feature extraction model adopts a spectral-spatial selective transformer model, which enhances the utilization of the central pixel information of the input hyperspectral image by using the central pixel position embedding module and the central pixel mixing module, and extracts the global features of the image and reduces the influence of irrelevant spatial information by using the spatial selective transformer module. The unknown class generation model consists of a multi-head attention mechanism and a multi-layer perceptron, and uses the known class prototypes to generate unknown class prototypes.
[0034] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Among them, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0035] Embodiment of a task-adaptive few-shot hyperspectral image open-set recognition method: As Figure 1 shown, a task-adaptive few-shot hyperspectral image open-set recognition method specifically includes the following steps: A hyperspectral image feature extraction model and an unknown class prototype generation model are constructed. In the training stage, a meta-learning task is constructed using a small amount of data in the source domain dataset and the target domain dataset to train the model and learn to generate unknown class prototypes to recognize unknown classes.
[0036] In this embodiment, the Indian Pines hyperspectral image data is taken as an example for classification and specific elaboration.
[0037] In this embodiment, Houston University 2013, Chikusei, Kennedy Space Center and Botswana are selected as the source domain datasets, and the University of Pavia is selected as the target domain dataset. The pseudo-color composite map corresponding to the data in the University of Pavia dataset is as Figure 5-1 shown, and the corresponding true value map is as Figure 5-2 shown. Among them, the University of Pavia dataset contains 9 types of ground objects, a total of 42,776 labeled samples. The first eight categories are selected as known classes, and the last category is selected as the unknown class. Since the number of spectral bands of the source domain data and the target domain data is different, the optimal clustering band selection method is used to unify the number of bands. In this example, the number of bands is unified to 100. In the training stage, 55 classes are selected from the source domain dataset, with 200 samples in each class. Among the 8 known classes in the target domain dataset, 5 samples are selected from each class to form all the training samples in the training stage. Gaussian noise is applied to the small number of known class samples in the target domain to generate additional samples, thereby alleviating the situation of insufficient training data in the target domain.
[0038] Secondly, a meta-learning task is constructed. Sixteen classes are sampled from the training set samples, and 20 samples are sampled from each class to form the training set of the meta-learning task. Arbitrarily select 8 of these classes as known classes, and the remaining 8 classes as the simulated unknown class query set. From the known classes, one sample is selected from each class to construct the support set, and the remaining samples form the known class query set, thus forming the training set of the meta-learning task.
[0039] The hyperspectral image feature extraction model uses a spectral-spatial selection transformer model, which is Figure 1 composed of three parts as shown. The first part uses a two-dimensional convolution and central pixel position embedding module to fuse the spatial information of the input data, while highlighting the importance of the central pixel. The second part consists of a series of spatial selection transformer blocks, which can reduce the interference of irrelevant spectral-spatial noise in the data while extracting global features through the spatial selection attention mechanism. The last part uses the central pixel mixing module to fuse the central pixel features of each spatial selection transformer block, and then realizes the final extraction of features.
[0040] The encoding schematic diagram of the central pixel position embedding module described in the first part of the spectral-spatial selection transformer model is as Figure 2 shown. It obtains the corresponding position encoding according to the relative distance of each pixel to the central pixel. The expression of its position encoding is:
[0041] where and are variables representing any indices of the pixel columns and rows in the input image. and represent the column and row indices corresponding to the central pixel. represents the set of position embeddings, represents the size of the image. After the central pixel embedding module encodes the input data, the spatial information of the input data is fused through the two-dimensional convolution module to obtain the embedded features. The expression of the embedded features is:
[0042] where represents the input image. Specifically, the input matrix size, convolution kernel size, the size of each convolution kernel, and the stride of the two-dimensional convolution module in this embodiment can all be set. For example, in this embodiment, the size of the input matrix is (13, 13, 100), the number of convolution kernels is 64, the size of each convolution kernel is (3, 3), the padding is (1, 1), and the stride is (1, 1). According to this setting, 64 matrices with a size of (13, 13) are output, and these matrices are integrated into a matrix with a size of (13, 13, 64).
[0043] The spatial selection transformer block described in the second part of the spectral spatial selection transformer model consists of three main modules: a spatial selection attention module, a multi-layer perceptron layer, and two layers of normalization layers. Residual connections are employed before the spatial selection attention module and the multi-layer perceptron layer.
[0044] The structure of the spatial selection attention module is as Figure 3 shown. The input three-dimensional data is converted into a two-dimensional matrix, and then three different matrices are generated through linear projection, namely the query matrix, the key matrix, and the value matrix. The matrix calculation expressions are as follows:
[0045] where, , and are linear projection matrices. The pixel correlation matrix is calculated by multiplying the query matrix with the transpose of the key matrix, and only the most relevant pixels of each pixel are retained to obtain the most relevant index matrix. The expression of the most relevant index matrix is as follows:
[0046] where, represents the index of obtaining the most relevant pixel in the correlation matrix. The most relevant matrices of the key matrix and the value matrix are calculated respectively using the most relevant index matrix. The calculation expressions are:
[0047] where gather means obtaining the most relevant values in the key matrix and the value matrix according to the index. Finally, the output of the spatial selection attention module is obtained by calculating the self-attention of the query matrix, the key matrix, and the value matrix, as well as local context enhancement. The calculation expression of the output matrix is:
[0048] where, is the local context enhancement term, which obtains enhanced features through depthwise separable convolution.
[0049] Specifically, in this embodiment, the matrix size of the input to the spatial selection attention module is (13, 13, 64), which is converted into a matrix of (169, 64). The matrix sizes generated by linear projection remain unchanged. The size of the pixel correlation matrix obtained by multiplying the query matrix with the transpose of the key matrix is (169, 169). The size of the most relevant index matrix obtained by obtaining the top 24 most relevant pixels is (169, 64), and the sizes of the most relevant matrices of the key matrix and the value matrix are (169, 64). The size of the output obtained by calculating the self-attention of the query matrix, the key matrix, and the value matrix, as well as local context enhancement of the spatial selection attention module is (169, 64).
[0050] The central pixel fusion mechanism in the third part fuses the central pixel features of each spatial selection transformer block. The calculation expression of the fusion process is:
[0051] where is a one-dimensional vector, represents the central pixel feature of the output feature of the -th spatial selective attention module. Specifically, there are five layers of spatial selective attention modules in this embodiment. The size of the output feature of the spatial selective attention module is (169, 64). The central pixels of multiple layers of features are spliced to obtain a central pixel splicing matrix with a size of (5, 64), and the size of the output feature obtained by fusion is (64).
[0052] The structure of the unknown class prototype generation model is as Figure 4 shown. The input end is connected to the output end of the feature extraction model. The average value of the feature data of each category output by the feature extraction model is obtained to obtain the known class prototype. The calculation expression of the known class prototype is:
[0053] where, represents the feature extraction process of the feature extraction model, represents the -th category feature quantity. The self-attention weight between the known class prototypes is calculated using the known class prototypes. The weight calculation expression is:
[0054] where, and are linear projection matrices, is the set of known class prototypes, is the feature length of the known class prototype. The weight matrix is normalized and the category negative prototype is generated. The calculation expression of the category negative prototype is:
[0055] where, represents the softmax function, is a trainable linear projection matrix. The average value of the category negative prototypes is calculated and input into a multi-layer perceptron to obtain the unknown class prototype. The calculation expression of the unknown class prototype is:
[0056] where, represents the category negative prototype generated by the -th class.
[0057] Specifically, there are 8 known classes in this embodiment. The size of the obtained set of known class prototypes is (8, 64). The size of the self-attention weight between the calculated known class prototypes is (8, 8), and a set of category negative prototypes with a size of (8, 64) is calculated. The unknown class prototype with a size of (64) is obtained by taking the average value and passing it through a multi-layer perceptron.
[0058] Train the hyperspectral image open-set recognition model using the above-mentioned source domain training set and target domain training set. The training process includes: Construct meta-learning tasks using a small amount of data from the source domain dataset and target domain dataset; input the support set and query set into the constructed hyperspectral feature extraction network to obtain support set features and query set features; use the support set features to obtain known class prototypes and then generate unknown class prototypes; Calculate the cross-entropy loss function and negative prototype loss function using the query set features and various prototypes, update the gradient of the loss function through backpropagation, and perform forward propagation to update the network training parameters according to the direction of the loss function decrease until the maximum number of training sets.
[0059] Specifically, the loss function is:
[0060] Among them, is the cross-entropy loss function, is the class negative prototype loss function, and the calculation expression of the class negative prototype loss function is:
[0061] Among them, represents the set of samples of the th class, represents the set of unknown class samples.
[0062] This embodiment also includes inputting the test data and a small amount of data from the target domain dataset into the trained feature extraction model, taking the average value of the features of each class-labeled sample as the class prototype and generating unknown class prototypes, and classifying based on the distances between the features of the hyperspectral image to be classified and each unknown class prototype, and obtaining the hyperspectral image classification result according to the classification result.
[0063] The classification result can be evaluated by the classification accuracy (PA), average classification accuracy (AA), overall classification accuracy (OA), and KAPPA coefficient of each output class, and the final target domain hyperspectral image classification result map is output. The calculation formulas of the average classification accuracy (AA), overall classification accuracy (OA), and KAPPA coefficient in the evaluation indicators are respectively:
[0064]
[0065]
[0066] Among them, is the number of correctly classified into this class, is the number of samples of this class that are misclassified into other classes, is the number of non - samples of this class that are classified into other classes, is the number of other samples that are misclassified into samples of this class, where p c 、p 0 respectively satisfy:
[0067]
[0068] In the formula, represents the sum of all elements in the i row of the confusion matrix, that is, the sum of the actual quantities of this class, represents the sum of all elements in the i column of the confusion matrix, that is, the sum of the quantities predicted as this class, N represents the sum of all elements, and c is the number of classes. represents the elements on the diagonal of the confusion matrix, that is, the number of correct predictions for each class.
[0069] In order to better verify the classification method in this embodiment, this embodiment uses a simulation system to classify the University of Pavia data by using the classification method in this embodiment and the classification method in the prior art respectively, and repeats the test 5 times with 5 random seeds respectively, calculates the confusion matrix, the classification accuracy PA of each class, the average classification accuracy AA, the overall classification accuracy OA, and the KAPPA coefficient, and obtains the classification results as shown in Figure 5-3 and Figure 5-4 , and obtains the test results as shown in Table 1. Among them Figure 5-3 is the classification result obtained by the classification method in the prior art, Figure 5-4 is the classification result obtained by the classification method in this embodiment. By Figure 5-3 and Figure 5-4 comparison, it can be found that the ground - object distribution of the classification result map obtained by the method proposed in the present invention is closer to the true - value map, and the classification effect obtained for most classes is better, realizing better classification performance.
[0070] Table 1: Simulation experiment results of the present invention and the comparative method
[0071] It can be seen from Table 1 that compared with the current relatively advanced MDL4OW method, the present invention has improved the classification accuracy for classes numbered 2, 5, 8, and 9, and at the same time, the average classification accuracy, the overall classification accuracy, and the KAPPA coefficient have also been significantly improved.
[0072] An embodiment of the present invention provides a small-sample hyperspectral image open-set recognition system based on task adaptation, including: A feature extraction module that inputs the hyperspectral image to be classified into a feature extraction model to extract features; An unknown class prototype generation module that inputs the extracted features into an unknown class prototype generation model to generate unknown class prototypes; A classification module that classifies the hyperspectral image based on the distance between the extracted features and the unknown class prototypes; wherein, the feature extraction model uses a spectral space selection transformer model; the unknown class prototype generation model consists of a multi-head attention mechanism and a multi-layer perceptron; the input end of the unknown class prototype generation model is connected to the output end of the feature extraction model, and the average value of the data features of each category output by the feature extraction model is obtained to obtain known class prototypes.
Claims
1. A small-sample hyperspectral image open-set recognition method based on task adaptation, characterized in that, It includes the following steps: Input the hyperspectral image to be classified into a pre-constructed open-set recognition model of hyperspectral images to obtain the hyperspectral image classification result; The construction method of the pre-constructed open-set recognition model of hyperspectral images is as follows: Input the hyperspectral image to be classified into a feature extraction model to extract features; Input the extracted features into an unknown class prototype generation model to generate unknown class prototypes; Classify the hyperspectral image based on the distance between the extracted features and the unknown class prototypes; Among them, the feature extraction model uses a spectral-spatial selective transformer model; the unknown class prototype generation model consists of a multi-head attention mechanism and a multi-layer perceptron; the input end of the unknown class prototype generation model is connected to the output end of the feature extraction model, calculates the known class prototypes based on the output of the feature extraction model, and generates unknown class prototypes after processing the known class prototypes through the multi-layer attention mechanism and the multi-layer perceptron.
2. The open-set recognition method for small-sample hyperspectral images based on task adaptation according to claim 1, characterized in that, The spectral-spatial selective transformer model includes: A central pixel position embedding module, which is used to fuse the spatial information of the input hyperspectral image to obtain embedded features; A spatial selective transformer module, which is used to extract global features from the embedded features and denoise them; A central pixel mixing module, which is used to fuse the central pixel features of the spatial selective transformer module to achieve feature extraction.
3. The small-sample hyperspectral image open-set recognition method based on task adaptation according to claim 2, wherein The central pixel position embedding module obtains the corresponding position encoding according to the relative distance of each pixel to the central pixel, and the expression of the position encoding is: Among them, and are variables representing arbitrary indices of pixel columns and rows in the input image; represents the column index corresponding to the central pixel; represents the row index corresponding to the central pixel; represents a set of positional embeddings; represents the size of the image; The expression of the embedded features is: Among them, represents the input image.
4. The small-sample hyperspectral image open-set recognition method based on task adaptation according to claim 2, wherein The spatial selective transformer module includes: a spatial selective attention module, a multi-layer perceptron layer, and two normalization layers; a residual connection is used before the spatial selective attention module and the multi-layer perceptron layer; the embedded features enter the spatial selective attention module to extract global features after passing through the first normalization layer, and the global features enter the multi-layer perceptron layer after passing through the second normalization layer to output the final global features.
5. The open-set recognition method for small-sample hyperspectral images based on task adaptation according to claim 4, wherein The specific process of the spatial selective attention module extracting global features is: The spatial selective attention module converts the input three-dimensional data into a two-dimensional matrix, and then generates a query matrix, a key matrix, and a value matrix through linear projection, as follows: Among them, , and are linear projection matrices; is the query matrix; is the key matrix; is the value matrix; Calculate the pixel correlation matrix by multiplying the query matrix with the transpose of the key matrix, and obtain the most relevant index matrix by keeping only the most relevant pixel for each pixel. The expression of the most relevant index matrix is as follows: Among them, represents obtaining the index of the most relevant pixel in the correlation matrix; Calculate the most relevant matrices of the key matrix and the value matrix respectively using the most relevant index matrix, and the calculation expression is: Among them, gather represents obtaining the most relevant values in the key matrix and the value matrix according to the index; The output of the spatial selection attention module is obtained by calculating the self-attention of the query matrix, key matrix, and value matrix and enhancing the local context. The calculation expression of the output matrix is as follows: Among them, is a local context enhancement item, which obtains enhanced features through depthwise separable convolution.
6. The method for open-set recognition of small-sample hyperspectral images based on task adaptation according to claim 2, wherein The central pixel mixing module fuses the central pixel features of each spatial selection transformer module, and the calculation expression for the fusion process is as follows: Among them, is a one-dimensional vector, indicating the central pixel feature of the output feature of the th spatial selective attention module.
7. The small-sample hyperspectral image open-set recognition method based on task adaptation according to claim 1, wherein The process of inputting the extracted features into the unknown class prototype generation model to generate unknown class prototypes is specifically: Calculate the known class prototypes based on the output of the feature extraction model, and the calculation expression of the known class prototypes is: Among them, represents the feature extraction process of the feature extraction model, represents the number of the features of the th known class prototype; Based on each known class prototype, calculate the self-attention weights between the known class prototypes; the calculation expression of the self-attention weights is as follows: Among them, and are linear projection matrices, is the feature length of the known class prototype; is the known class prototype matrix; Generate a class negative prototype based on a known class prototype and self-attention weights. The calculation expression for the class negative prototype is as follows: Among them, represents the softmax function, is the linear projection matrix; Calculate the average value of the class negative prototypes and input it into a multi-layer perceptron to obtain the unknown class prototype. The calculation expression for the unknown class prototype is as follows: Among them, represents the class negative prototype generated by the th class.
8. The small-sample hyperspectral image open-set recognition method based on task adaptation according to claim 1, characterized in that The loss function of the pre-constructed open-set recognition model of hyperspectral images is: Among them, is the cross-entropy loss function, is the class negative prototype loss function.
9. The method for open-set recognition of small-sample hyperspectral images based on task adaptation according to claim 8, wherein The negative prototype loss function of the category has the following calculation expression: Among them, represents the set of known class query set samples of the class, and represents the set of unknown class query set samples.
10. A small-sample hyperspectral image open-set recognition system based on task adaptation, characterized in that It includes: A feature extraction module, which inputs the hyperspectral image to be classified into a feature extraction model to extract features; An unknown class prototype generation module, which inputs the extracted features into an unknown class prototype generation model to generate unknown class prototypes; A classification module, which classifies the hyperspectral image based on the distance between the extracted features and the unknown class prototypes; Among them, the feature extraction model uses a spectral-spatial selective transformer model; the unknown class prototype generation model consists of a multi-head attention mechanism and a multi-layer perceptron; the input end of the unknown class prototype generation model is connected to the output end of the feature extraction model, and the average value of the data features of each category output by the feature extraction model is obtained to obtain the known class prototypes.
Citation Information
Cited By
Small sample open set identification method and system based on multi-modal negative prototype, and medium
CN120804600A
Small sample open set recognition method and system based on multi-modal negative prototype, and medium
CN120804600B
Planetary surface interested target spectrum identification system and method
CN121582695A