Text-driven hyperspectral image terrain classification method
Through text-driven hyperspectral image geographic classification method, the integration of spectral, spatial and semantic information is solved, and the problem of insufficient characterization of geographic categories in the prior art is achieved, and higher classification accuracy and consistency are achieved, which is suitable for agricultural monitoring and environmental assessment.
Patent Information
- Application Number
- CN202510348570.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing hyperspectral image classification methods fail to effectively utilize the hierarchical relationship, semantic correlation and topological structure between land objects, resulting in low classification accuracy.
The text-driven hyperspectral image geographic classification method is used to extract the semantic features of geographic categories through text encoder, and image features are extracted in combination with the HSI encoder, spectral, spatial and semantic information are fused, and feature representation capabilities are improved using multi-head self-attention mechanism and feedforward neural network.
It significantly improves the accuracy and accuracy of land objects classification, can more accurately reflect the essential properties of land objects, and is suitable for practical application scenarios such as land use type classification and ecological environment assessment.
Smart Images

Figure CN120298885A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image classification, and particularly relates to a method for classifying ground objects in hyperspectral images driven by text. Background Art
[0002] With the rapid development of remote sensing technology, hyperspectral images have shown great application potential in the fields of precise ground object classification, environmental monitoring, resource exploration, etc. due to their rich spectral information and spatial information. Hyperspectral images (HSIs) can provide fine reflection information of ground objects in multiple spectral bands, which is of great significance for distinguishing ground object types with subtle spectral differences.
[0003] However, existing methods for classifying ground objects in hyperspectral images face some challenges and limitations. On the one hand, traditional classification methods based on spectral features, such as Spectral Angle Mapping (SAM), parallelepiped classification method, etc., mainly rely on spectral features and ignore the correlation of ground objects at the spatial structure and semantic levels. This leads to a significant reduction in classification accuracy when dealing with ground objects with similar spectral features but different spatial distributions. For example, in urban areas, the rooftops of buildings and road surfaces may have similar spectral reflection characteristics, but their spatial layouts and functional uses are very different, and it is difficult to effectively distinguish them only relying on spectral features.
[0004] On the other hand, in recent years, hyperspectral image classification methods based on deep learning have been widely used. This method can automatically learn and extract complex features in hyperspectral data, thereby improving classification performance. However, existing ground object classification methods generally use one-hot encoding to represent ground object categories, and this method has the following technical defects: One-hot encoding regards ground object categories as independent discrete labels and fails to effectively represent important information such as hierarchical relationships, semantic associations, and topological structures between categories. For example, the vegetation category and the crop category are closely related in actual application scenarios, but in the one-hot encoding representation, the two are treated as completely independent labels. This limits the model's ability to learn deep features of ground object categories, resulting in limited classification accuracy and generalization performance. In addition, one-hot encoding cannot effectively fuse multi-source information such as semantic descriptions, environmental backgrounds, and spatial associations, and these information are of great value for improving the accuracy of ground object classification. For example, the environmental background and spatial position relationship of ground objects can provide auxiliary information to help distinguish similar ground object categories, but one-hot encoding cannot incorporate this information into the model, further exacerbating the limitation of classification accuracy.
[0005] In summary, existing methods for classifying ground objects in hyperspectral images have problems in making full use of spectral, spatial, and semantic information. There is an urgent need for a new classification method to overcome the above problems, improve the accuracy and efficiency of classifying ground objects in hyperspectral images, and better meet the actual application requirements. Summary of the invention
[0006] The purpose of the present invention is to solve the problem that the existing land object classification methods generally use one-hot encoding to represent the land object categories, but the one-hot encoding fails to effectively represent the important information such as the hierarchical relationship, semantic association, topological structure between categories, and the one-hot encoding cannot effectively integrate multi-source information such as semantic description, environmental background and spatial association, resulting in low classification accuracy. A text-driven hyperspectral image land object classification method is proposed.
[0007] A text-driven hyperspectral image object classification method. The specific process is as follows:
[0008] Step 1: obtain a hyperspectral image HSI of a marked object type, preprocess the hyperspectral image HSI of the marked object type to obtain a preprocessed hyperspectral image HSI, and use the preprocessed hyperspectral image HSI as a training set;
[0009] Step 2: The hyperspectral image HSI in each patch preprocessed in step 1 is sequentially subjected to normalization enhancement, noise enhancement, multi-scale transformation enhancement, and horizontal mirror transformation enhancement to obtain an enhanced hyperspectral image HSI in each patch;
[0010] Step 3: Obtain text description information of the ground object category in the hyperspectral image HSI in each patch preprocessed in step 1;
[0011] Step 4: Input the text description information of the ground feature category corresponding to the hyperspectral image HSI in each patch obtained in step 3 into the text encoder for encoding, and the text encoder outputs the semantic features of the text description information;
[0012] Step 5: Input the enhanced hyperspectral image HSI in each patch obtained in step 2 into the HSI encoder, and the HSI encoder outputs the image features in each patch. The image features in each patch output by the HSI encoder are sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result;
[0013] Step 6: Train the text encoder and the HSI encoder to obtain the trained text encoder and the HSI encoder;
[0014] Step 7: Collect the hyperspectral image data to be tested and generate classification results based on the trained HSI encoder.
[0015] The beneficial effects of the present invention are:
[0016] The present invention provides a text-driven method for classifying objects in hyperspectral images, which overcomes the problems of insufficient fusion of spectral and spatial information and lack of utilization of semantic information in the prior art by effectively fusing spectral, spatial, and semantic information, thereby significantly improving the accuracy of object classification.
[0017] The text-driven method for classifying objects in hyperspectral images of the present invention has the following advantages and characteristics:
[0018] Full utilization of multi-source information:
[0019] By fusing the spectral and spatial information of hyperspectral images and the semantic information contained in text descriptions, the present invention can comprehensively characterize the features of objects, significantly improving the accuracy and precision of object classification. Compared with traditional classification methods that rely only on spectral or spatial features, the classification accuracy of the present invention has significant advantages.
[0020] Enhanced semantic consistency of classification results:
[0021] The present invention effectively improves the semantic consistency and rationality of object classification results by introducing text descriptions, enabling them to more accurately reflect the essential attributes of objects. This technical advantage significantly enhances the value of classification results in practical application scenarios such as land use type classification and ecological environment assessment, and can provide classification information that better meets the semantic understanding needs of users.
[0022] Improved classification accuracy:
[0023] The present invention uses the MUUFL Gulfport dataset for testing. This dataset was collected by the ROSIS sensor on the campus of the University of Southern Mississippi in November 2010 and contains a hyperspectral image (HSI) with a size of 325×220 pixels and 64 effective spectral bands. The dataset covers 11 urban land cover classes and a total of 53,687 ground truth pixels are labeled. In the experiment, 100 pixels are randomly selected from each class for training, and the remaining 52,587 pixels are used for testing.
[0024] The present invention achieves classification accuracies of OA: 91.95%, AA: 93.17%, and Kappa: 89.39% on the MUUFL dataset through a text-driven method and optimizing the structure of the HSI encoder. It is suitable for fast classification processing of large-scale hyperspectral image data and meets the requirements of real-time and efficiency in practical applications. Description of the Drawings
[0025] Figure 1 It is a flowchart of the method for classifying objects in hyperspectral images;
[0026] Figure 2 It is a structural diagram of the HSI encoder;
[0027] Figure 3 This is the structural diagram of the spatial extractor in the HSI encoder. DETAILED DESCRIPTION
[0028] Specific implementation method 1: The specific process of a text-driven hyperspectral image object classification method in this implementation method is as follows:
[0029] Step 1: obtain a hyperspectral image HSI of a marked object type, preprocess the hyperspectral image HSI of the marked object type to obtain a preprocessed hyperspectral image HSI, and use the preprocessed hyperspectral image HSI as a training set;
[0030] The hyperspectral image data of a certain area is selected as the data set. The data set contains various types of objects, such as vegetation, buildings, water bodies, roads, etc., and has corresponding object annotation information for model training and classification effect verification;
[0031] Step 2: The hyperspectral image HSI in each patch preprocessed in step 1 is sequentially subjected to normalization enhancement, noise enhancement, multi-scale transformation enhancement, and horizontal mirror transformation enhancement to obtain an enhanced hyperspectral image HSI in each patch;
[0032] Step 3: Obtain text description information of the ground object category in the hyperspectral image HSI in each patch preprocessed in step 1;
[0033] Step 4: Input the text description information (answer) of the ground object category corresponding to the hyperspectral image HSI in each patch obtained in step 3 into the text encoder for encoding, and the text encoder outputs the semantic features of the text description information;
[0034] Step 5: Input the enhanced hyperspectral image HSI in each patch obtained in step 2 into the HSI encoder, and the HSI encoder outputs the image features in each patch. The image features in each patch output by the HSI encoder are sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result;
[0035] Step 6: Train the text encoder and the HSI encoder to obtain the trained text encoder and the HSI encoder;
[0036] Step 7: Collect the hyperspectral image data to be tested and generate classification results based on the trained HSI encoder.
[0037] Each pixel is assigned to a specific feature category. The results are further optimized through visual verification to make them more consistent with the distribution patterns and semantic logic of actual features, thus providing reliable data support for applications such as agricultural monitoring, resource management, and environmental assessment.
[0038] Meanwhile, standard evaluation metrics (Overall Accuracy OA, Average Accuracy AA, Kappa coefficient) are used to evaluate the classification results to verify the performance of the model.
[0039] Specific Embodiment 2: Different from Specific Embodiment 1, in Step 1, a hyperspectral image HSI of the marked ground object type is obtained, and the hyperspectral image HSI of the marked ground object type is preprocessed to obtain a preprocessed hyperspectral image HSI, and the preprocessed hyperspectral image HSI is used as the training set;
[0040] The hyperspectral image data of a certain area is selected as the data set. The data set contains various ground object types, such as vegetation, buildings, water bodies, roads, etc., and has corresponding ground object annotation information for model training and classification effect verification;
[0041] The specific process is as follows:
[0042] Obtain a hyperspectral image HSI of the marked ground object type;
[0043] Patch extraction is performed on the hyperspectral image HSI of the marked ground object type, and the size of each patch is h×w;
[0044] Among them, h is the length of each patch, and w is the width of each patch;
[0045] h×w = 9×9;
[0046] The hyperspectral image HSI after patch extraction is used as the training set.
[0047] Other steps and parameters are the same as those in Specific Embodiment 1.
[0048] Specific Embodiment 3: Different from Specific Embodiment 1 or 2, in Step 2, the hyperspectral image HSI in each patch preprocessed in Step 1 is sequentially subjected to normalization enhancement, noise addition enhancement, multi-scale transformation enhancement, and horizontal mirror transformation enhancement to obtain the hyperspectral image HSI in each enhanced patch;
[0049] The specific process is as follows:
[0050] Step 2-1. Perform normalization enhancement on the hyperspectral image HSI in each patch preprocessed in Step 1 to obtain the hyperspectral image HSI in each patch after normalization enhancement;
[0051] The normalization enhancement formula is:
[0052]
[0053] Among them,
[0054] x i The \(i\)-th sample in the hyperspectral image HSI of each preprocessed patch;
[0055] x′ i The \(i\)-th sample in the hyperspectral image HSI of each patch after normalization enhancement. After min-max normalization, the sample data is mapped to (0, 1);
[0056] min(x i ) is the minimum value of \(x\) in each patch (the maximum pixel value in each patch), and max(x i ) is the maximum value of \(x\) in each patch (the minimum pixel value in each patch); i ) is the maximum value of \(x\) in each patch; i The minimum value of \(x\) in each patch (the minimum pixel value in each patch);
[0057] Step 2: Perform noise addition enhancement on the hyperspectral image HSI of each patch after normalization enhancement to obtain the hyperspectral image HSI of each patch after noise addition enhancement; the expression is:
[0058]
[0059] where,
[0060] \(\mu\) is the mean value and \(\delta\) is the variance;
[0061] x″ i The \(i\)-th sample after noise addition enhancement of the \(i\)-th sample in the hyperspectral image HSI of each patch after normalization enhancement;
[0062] Step 3: Perform multi-scale transformation enhancement on the hyperspectral image HSI of each patch after noise addition enhancement to obtain the hyperspectral image HSI of each patch after multi-scale transformation enhancement; the specific process is as follows:
[0063] For the hyperspectral image HSI of each patch after noise addition enhancement, take a window of size 7×7 or 5×5 with the patch center as the origin, and then add each 7×7 or 5×5 window to a 9×9 window size to obtain the hyperspectral image HSI after multi-scale transformation enhancement;
[0064] Step 4: Perform horizontal mirror transformation enhancement on the hyperspectral image HSI of each patch after multi-scale transformation enhancement to obtain the hyperspectral image HSI of each patch after horizontal mirror transformation enhancement; the expression is:
[0065]
[0066] Among them,
[0067] (x″′ i , y″′ i ) is the coordinate (coordinate in the pixel coordinate system) of the i-th sample in the hyperspectral image HSI in each patch after multi-scale transformation enhancement;
[0068] is the coordinate of the i-th sample after horizontal mirror transformation enhancement for the i-th sample in the hyperspectral image HSI in each patch after multi-scale transformation enhancement;
[0069] w is the width of the i-th sample in the hyperspectral image HSI in each patch after multi-scale transformation enhancement (the width of the i-th sample is 9).
[0070] Other steps and parameters are the same as those in the first or second specific implementation manner.
[0071] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that in step three, the text description information of the ground object categories in the hyperspectral image HSI in each patch after the preprocessing in step one is obtained;
[0072] The specific process is as follows:
[0073] Using a large language model, through instruction input, generate the text description information of the ground object categories in the hyperspectral image HSI in each patch after the preprocessing in step one;
[0074] Instructions and answer examples:
[0075] Instruction: We have 15 categories, 1 - healthy grass, 2 - oppressed grass, 3 - synthetic grass, 4 - trees, 5 - soil, 6 - water body, 7 - residential area, 8 - commercial area, 9 - road, 10 - highway, 11 - railway, 12 - parking lot 1, 13 - parking lot 2, 14 - tennis court, 15 - runway; On the hyperspectral image dataset, what are the useful visual features for distinguishing the 15 categories? Represent with the most significant short description for each category;
[0076] Answer:
[0077] 1. Healthy grass: Represents high NIR reflectance and strong chlorophyll absorption;
[0078] 2. Oppressed grass: Reduced near-infrared reflectance and shifted red edge;
[0079] 3. Synthetic grass: The spectral features are the spectral peaks of artificial materials and uniform spectral patterns;
[0080] 4. Trees: The spectral features are strong vegetation features and shadow changes;
[0081] 5. Soil: The spectral feature is a strong mineral absorption feature;
[0082] 6. Water body: The spectral feature is extremely low near-infrared reflectance and absorption in the short-wave infrared;
[0083] 7. Residential area: The spectral feature is a mixed material feature (diverse roof types);
[0084] 8. Commercial area: The spectral feature is that the spectral of the roof material is uniform;
[0085] 9. Road: The spectral feature is an asphalt feature and a continuous pattern;
[0086] 10. Highway: The spectral feature is a broad asphalt feature with clear boundaries;
[0087] 11. Railway: The spectral feature is a metal and gravel feature;
[0088] 12. Parking lot 1: The spectral feature is an asphalt feature with markings;
[0089] 13. Parking lot 2: The spectral feature is a different asphalt / concrete feature;
[0090] 14. Tennis court: The spectral feature is a unique court surface material;
[0091] 15. Runway: The spectral feature is an artificial material feature, different from the lawn.
[0092] Other steps and parameters are the same as those in any one of the specific embodiments 1 to 3.
[0093] Specific embodiment 5: The difference between this embodiment and any one of the specific embodiments 1 to 4 is that in step 4, the text description information (answer) of the ground object category corresponding to the hyperspectral image HSI in each patch obtained in step 3 is input into the text encoder for encoding, and the text encoder outputs the semantic features of the text description information; the specific process is as follows:
[0094] Input the 15 pieces of text description information (answer) of the ground object category corresponding to the hyperspectral image HSI in each patch into the RemoteCLIP text encoder for encoding, and the RemoteCLIP text encoder outputs the semantic features of the text description information;
[0095] Establish the association between remote sensing images and professional text information to achieve the precise mapping of remote sensing visual content and professional semantic narration. RemoteCLIP embeds remote sensing professional knowledge into the model by integrating remote sensing images and professional text information, significantly enhancing the cognitive ability for complex ground objects and phenomena, such as accurately distinguishing vegetation, analyzing land use patterns, and identifying water characteristics. Compared with the general CLIP model, RemoteCLIP shows higher accuracy and efficiency in remote sensing specific tasks, attributed to its specialized learning and understanding of the remote sensing environment.
[0096] Other steps and parameters are the same as those in any one of the first to fourth specific embodiments.
[0097] Specific embodiment six: The difference between this embodiment and any one of the first to fifth specific embodiments is that in step five, the hyperspectral image HSI in each enhanced patch obtained in step two is input into the HSI encoder, and the HSI encoder outputs the image features in each patch. The image features in each patch output by the HSI encoder are sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result; the specific process is as follows:
[0098] The HSI encoder sequentially includes a first 1×1 convolutional layer, a BN layer, a spatial feature extractor, and a spectral feature extractor;
[0099] The hyperspectral image HSI in each enhanced patch obtained in step two is sequentially input into the first 1×1 convolutional layer and the BN layer, and the BN layer outputs feature A;
[0100] The feature A output by the BN layer is input into the spatial feature extractor, and the spatial feature extractor outputs feature B;
[0101] The feature B output by the spatial feature extractor is input into the spectral feature extractor, and the spectral feature extractor outputs feature C;
[0102] Feature C is sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result.
[0103] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.
[0104] Specific embodiment seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that the spatial feature extractor includes:
[0105] The first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, the third 1×1 convolution layer, the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, the fifth 1×1 convolution layer, the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, the seventh 1×1 convolution layer;
[0106] The working process of the spatial feature extractor is as follows:
[0107] The output feature A of the BN layer is sequentially input into the first 7×7 depth convolution layer, the first LN layer, the second 1×1 convolution layer, the first GELU, the first GRN, and the third 1×1 convolution layer, and the third 1×1 convolution layer outputs the feature A';
[0108] The output feature A' of the third 1×1 convolution layer and the output feature A of the BN layer are element-wise added to obtain the feature A";
[0109] The feature A" is sequentially input into the second 7×7 depth convolution layer, the second LN layer, the fourth 1×1 convolution layer, the second GELU, the second GRN, and the fifth 1×1 convolution layer, and the fifth 1×1 convolution layer outputs the feature A''';
[0110] The output feature A''' of the fifth 1×1 convolution layer and the feature A" are element-wise added to obtain the feature
[0111] Feature Is sequentially input into the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, and the seventh 1×1 convolution layer, and the seventh 1×1 convolution layer outputs the feature
[0112] The output feature of the seventh 1×1 convolution layer And the feature Are element-wise added to obtain the feature B;
[0113] The feature B is the output feature of the spatial feature extractor.
[0114] Other steps and parameters are the same as those in any one of the specific embodiments one to six.
[0115] Specific embodiment eight: The difference between this embodiment and any one of the specific embodiments one to seven is that the spectral feature extractor includes:
[0116] The fourth LN layer, the multi-head self-attention mechanism MHSA, the fifth LN layer, the feed-forward neural network (FFN);
[0117] The working process of the spectral feature extractor is as follows:
[0118] The output feature B of the spatial feature extractor is input into the fourth LN layer, and the fourth LN layer outputs the feature B';
[0119] The output feature B' of the fourth LN layer is input into the multi-head self-attention mechanism MHSA, and the multi-head self-attention mechanism MHSA outputs the feature B";
[0120] The output feature B" of the multi-head self-attention mechanism MHSA and the output feature B of the spatial feature extractor are input into the fifth LN layer, and the fifth LN layer outputs the feature B''';
[0121] The output feature B''' of the fifth LN layer is input into the feed-forward neural network (FFN), and the feed-forward neural network (FFN) outputs the feature
[0122] The output feature of the multi-layer perceptron MLP The output feature B of the spatial feature extractor and the output feature B" of the multi-head self-attention mechanism MHSA are element-wise added to obtain the feature C;
[0123] The output feature C of the spectral feature extractor is used as the output feature of the spectral feature extractor.
[0124] Other steps and parameters are the same as those in any one of the specific embodiments one to seven.
[0125] The HSI data after data augmentation is input through the HSI encoder to extract HSI features, including spectral features and spatial features.
[0126] As Figure 2 、 Figure 3 shown in the HSI encoder, the HSI encoder is responsible for extracting spatial and spectral features from hyperspectral data. The HSI encoder consists of two main parts: a spatial feature extractor and a spectral feature extractor. Among them, the spatial feature extractor is based on a convolutional neural network (CNN), and the spectral feature extractor is based on the Transformer architecture.
[0127] Among them, the spatial feature extractor aims to efficiently extract local spatial features from hyperspectral images. This extractor consists of 3 identical modules, which use large-kernel depthwise separable convolutions (such as 7×7) to aggregate global features in order to capture the spatial distribution information of ground objects. The depthwise separable convolution decomposes the standard convolution, extracts spatial features through depthwise convolution, and fuses channel information through pointwise convolution. Layer Normalization stabilizes the feature distribution and improves the training efficiency. Subsequently, the number of feature channels is adjusted through a 1×1 convolution layer, non-linearity is introduced by combining with the GELU activation function, and the common information is passed through the Global Response Normalization (GRN) layer. Finally, the number of channels is reduced through a 1×1 convolution layer to reduce the computational complexity. Residual connections enhance the network's learning ability for spatial features, which is crucial for understanding the spatial features of ground objects. Through the spatial feature extractor, the model can capture the spatial distribution patterns and texture features of ground objects, such as the geometry of farmland, the outline of buildings, and the coverage of vegetation. Through operations such as large-kernel convolution, the spatial feature extractor can generate multi-level spatial feature maps, which not only contain local details but also retain global spatial relationships.
[0128] The spatial features extracted by the spatial feature extractor then enter the spectral feature extractor, which aims to extract global features from the spectral dimension of hyperspectral images. This extractor uses the Transformer architecture, and its core components include the multi-head self-attention mechanism (MHSA) and the feed-forward neural network (FFN). The MHSA can capture the long-range dependencies between spectral bands and dynamically aggregate spectral information; the MLP enhances the non-linear expression ability of features. Through the Transformer block, the extractor gradually constructs global spectral features. Through the attention mechanism, detailed spectral information is extracted from the hyperspectral data. The spectral feature extractor can identify the reflectance differences of different ground objects in each band, thereby accurately distinguishing different ground object types such as vegetation, water bodies, and soil.
[0129] The finally generated feature representation combines high-resolution spatial information and detailed spectral information, and can more comprehensively and accurately describe the ground object attributes of each pixel point. This rich feature representation provides strong support for subsequent classification tasks, enabling the classification model to better handle complex ground object distributions and changes, improving the classification accuracy and robustness, and being applicable to various application scenarios such as agricultural monitoring, resource management, and environmental assessment.
[0130] Specific Embodiment Nine: What is different from one of Specific Embodiments One to Eight is that in Step Six, the text encoder and the HSI encoder are trained to obtain the trained text encoder and HSI encoder;
[0131] The specific process is as follows:
[0132] Calculate the integrated HSI and text contrast loss function based on the semantic features of the text description information output by the text encoder in step four and the image features (feature C output by the spectral feature extractor) output by the HSI encoder in step five.
[0133]
[0134] s(F HSI ,F text ) = norm(F HSI ) T norm(F text )
[0135] Where N represents the total number of hyperspectral images HSI, i represents the i-th one; k represents the k-th one; i = 1, 2, …, N; k = 1, 2, …, N;
[0136] represents an intermediate variable;
[0137] represents an intermediate variable;
[0138] represents the image features output by the HSI encoder when the i-th hyperspectral image HSI is input to the HSI encoder;
[0139] represents the image features output by the HSI encoder when the k-th hyperspectral image HSI is input to the HSI encoder;
[0140] represents the semantic features of the text description information output by the text encoder when the 15 text description information (answers) corresponding to the i-th hyperspectral image HSI are input to the text encoder;
[0141] represents the semantic features of the text description information output by the text encoder when the 15 text description information (answers) corresponding to the k-th hyperspectral image HSI are input to the text encoder;
[0142] represents the similarity between the image features corresponding to the i-th hyperspectral image HSI and the semantic features corresponding to the i-th hyperspectral image HSI;
[0143] represents the similarity between the image features corresponding to the i-th hyperspectral image HSI and the semantic features corresponding to the k-th hyperspectral image HSI;
[0144] Represents the similarity between the semantic features corresponding to the \(i\)-th hyperspectral image HSI and the image features corresponding to the \(i\)-th hyperspectral image HSI;
[0145] Represents the similarity between the semantic features corresponding to the \(i\)-th hyperspectral image HSI and the image features corresponding to the \(k\)-th hyperspectral image HSI;
[0146] s(F HSI ,F text ) represents the similarity between the image features corresponding to the hyperspectral image HSI and the semantic features corresponding to the hyperspectral image HSI;
[0147] F HSI Represents the image features output by the HSI encoder in step five; F text Represents the semantic features of the text description information output by the text encoder in step four;
[0148] norm(F HSI ) represents normalizing the image features output by the HSI encoder in step five; norm(F text ) represents normalizing the semantic features of the text description information output by the text encoder in step four;
[0149] The superscript \(T\) represents taking the transpose;
[0150] \(\tau\) represents a temperature hyperparameter;
[0151] Based on the classification result \(p(x\) ic ) output in step five, calculate the HSI classification loss function
[0152]
[0153] where
[0154] \(C\) represents the total number of classes;
[0155] y ic represents the ground truth of whether the \(i\)-th sample belongs to the \(c\)-th class of land cover; \(p(x\) ic ) represents the probability of predicting that the \(i\)-th sample belongs to the \(c\)-th class of land cover;
[0156] Based on the integrated HSI and text contrast loss function and the HSI classification loss function calculate the total loss function
[0157] Until the total loss function converges, obtain the trained text encoder and HSI encoder.
[0158] Other steps and parameters are the same as those in any one of the first to eighth specific embodiments.
[0159] Specific Embodiment Ten: The difference between this embodiment and any one of the first to ninth specific embodiments is that the total loss function is calculated based on the integrated HSI and text contrast loss function and the HSI classification loss function and is expressed as: That is:
[0160]
[0161] The present invention adopts a strategy of combining contrast loss and classification loss to improve the accuracy of hyperspectral image land cover classification. Among them, the contrast loss enhances the discrimination between classes by exploring the potential correlation between hyperspectral images (HSI) and text modalities, pulling closer the sample features of the same class with different modalities and pushing away the features of different classes, so as to achieve accurate identification of land cover differences. The classification loss guides the model optimization by reducing the error between the model prediction result and the true label to achieve accurate land cover classification.
[0162] Finally, through training, the final land cover classification model is obtained.
[0163] Other steps and parameters are the same as those in any one of the first to ninth specific embodiments.
[0164] The present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A text-driven method for classifying features in hyperspectral images, characterized in that: The specific process of the method is as follows: Step 1: Obtain the hyperspectral image HSI of the marked ground object type, preprocess the hyperspectral image HSI of the marked ground object type to obtain the preprocessed hyperspectral image HSI, and use the preprocessed hyperspectral image HSI as the training set; Step 2: Successively perform normalization enhancement, noise addition enhancement, multi-scale transformation enhancement, and horizontal mirror transformation enhancement on the hyperspectral image HSI in each patch after the preprocessing in Step 1 to obtain the enhanced hyperspectral image HSI in each patch; Step 3: Obtain the text description information of the ground object category in the hyperspectral image HSI in each patch after the preprocessing in Step 1; Step 4: Input the text description information of the ground object category corresponding to the hyperspectral image HSI in each patch obtained in Step 3 into the text encoder for encoding, and the text encoder outputs the semantic features of the text description information; Step 5: Input the enhanced hyperspectral image HSI in each patch obtained in Step 2 into the HSI encoder, the HSI encoder outputs the image features in each patch, the image features in each patch output by the HSI encoder are successively input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result; Step 6: Train the text encoder and the HSI encoder to obtain the trained text encoder and HSI encoder; Step 7: Collect the hyperspectral image data to be measured, and generate the classification result based on the trained HSI encoder.
2. A text-driven hyperspectral image ground object classification method according to claim 1, characterized in that: In Step 1, obtain the hyperspectral image HSI of the marked ground object type, preprocess the hyperspectral image HSI of the marked ground object type to obtain the preprocessed hyperspectral image HSI, and use the preprocessed hyperspectral image HSI as the training set; The specific process is as follows: Obtain the hyperspectral image HSI of the marked ground object type; Perform patch extraction on the hyperspectral image HSI of the marked ground object type, and the size of each patch is h×w; Where h is the length of each patch and w is the width of each patch; h×w = 9×9; Use the hyperspectral image HSI after patch extraction as the training set.
3. A text-driven hyperspectral image ground object classification method according to claim 2, characterized in that: In Step 2, successively perform normalization enhancement, noise addition enhancement, multi-scale transformation enhancement, and horizontal mirror transformation enhancement on the hyperspectral image HSI in each patch after the preprocessing in Step 1 to obtain the enhanced hyperspectral image HSI in each patch; The specific process is as follows: Step 2-1: Perform normalization enhancement on the hyperspectral image HSI in each patch after the preprocessing in Step 1 to obtain the normalized enhanced hyperspectral image HSI in each patch; The normalization enhancement formula is: Where xi is the i-th sample in the hyperspectral image HSI in each patch after preprocessing; x′ i is the i-th sample in the hyperspectral image HSI of each patch after normalization enhancement; min(x i ) is the minimum value of x i in each patch; max(x i ) is the maximum value of x i in each patch; Step 2-2: Perform noise addition enhancement on the normalized enhanced hyperspectral image HSI in each patch to obtain the noise addition enhanced hyperspectral image HSI in each patch; The expression is: Where μ is the mean value and δ is the variance; x″ i It is the i-th sample after noise addition enhancement for the i-th sample in the hyperspectral image HSI in each patch after normalization enhancement; Step 23: Perform multi-scale transformation enhancement on the hyperspectral image HSI in each patch after noise addition and enhancement to obtain the hyperspectral image HSI in each patch after multi-scale transformation enhancement. The specific process is as follows: For the hyperspectral image HSI in each patch after noise addition and enhancement, take a window of size 7×7 or 5×5 with the patch center as the origin, and then add each 7×7 or 5×5 window to a 9×9 window size with 0 to obtain the hyperspectral image HSI after multi-scale transformation enhancement. Step 24: Perform horizontal mirror transformation enhancement on the hyperspectral image HSI in each patch after multi-scale transformation enhancement to obtain the hyperspectral image HSI in each patch after horizontal mirror transformation enhancement. The expression is: where (x″′ i ,y″′ i ) is the coordinate of the i-th sample in the hyperspectral image HSI of each patch after multi-scale transform enhancement; It is the coordinate of the $i$-th sample after horizontal mirror transformation enhancement for the $i$-th sample in the hyperspectral image (HSI) of each patch after multi-scale transformation enhancement; w is the width of the i-th sample in the hyperspectral image HSI in each patch after multi-scale transformation enhancement.
4. A text-driven hyperspectral image ground object classification method according to claim 3, characterized in that: In step 3, obtain the text description information of the ground object category in the hyperspectral image HSI in each patch after preprocessing in step 1. The specific process is as follows: Use a large language model to generate the text description information of the ground object category in the hyperspectral image HSI in each patch after preprocessing in step 1 through instruction input. Instructions and answer examples: Instruction: We have 15 categories, 1 - healthy grass, 2 - stressed grass, 3 - synthetic grass, 4 - trees, 5 - soil, 6 - water body, 7 - residential area, 8 - commercial area, 9 - road, 10 - highway, 11 - railway, 12 - parking lot 1, 13 - parking lot 2, 14 - tennis court, 15 - runway; on the hyperspectral image dataset, what are the useful visual features for distinguishing 15 categories? Represent with the most significant short description for each category. Answer:
1. Healthy grass: Represents high NIR reflectance and strong chlorophyll absorption.
2. Stressed grass: Reduced near-infrared reflectance and shifted red edge.
3. Synthetic grass: The spectral characteristics are the spectral peaks of artificial materials and uniform spectral patterns.
4. Trees: The spectral characteristics are strong vegetation characteristics and shadow changes.
5. Soil: The spectral characteristics are strong mineral absorption characteristics.
6. Water body: The spectral characteristics are extremely low near-infrared reflectance and absorption in the short-wave infrared.
7. Residential area: The spectral characteristics are mixed material characteristics.
8. Commercial area: The spectral characteristics are uniform roof material spectra.
9. Road: The spectral characteristics are asphalt characteristics and continuous patterns.
10. Highway: The spectral characteristics are broad asphalt characteristics and clear boundaries.
11. Railway: The spectral characteristics are metal and gravel characteristics.
12. Parking lot 1: The spectral characteristics are asphalt characteristics with markings.
13. Parking lot 2: The spectral characteristics are different asphalt / concrete characteristics.
14. Tennis court: The spectral characteristics are unique court surface materials.
15. Runway: The spectral characteristics are artificial material characteristics, different from lawns.
5. A text-driven hyperspectral image feature classification method according to claim 4, characterized in that: In step 4, input the text description information of the ground object category corresponding to the hyperspectral image HSI in each patch obtained in step 3 into a text encoder for encoding, and the text encoder outputs the semantic features of the text description information. The specific process is as follows: The 15 text description information of the ground object categories corresponding to the hyperspectral image HSI in each patch is input into the RemoteCLIP text encoder for encoding, and the RemoteCLIP text encoder outputs the semantic features of the text description information.
6. A text-driven hyperspectral image feature classification method according to claim 5, characterized in that: In step 5, the enhanced hyperspectral image HSI in each patch obtained in step 2 is input into the HSI encoder, the HSI encoder outputs the image features in each patch, and the image features in each patch output by the HSI encoder are sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result; the specific process is as follows: The HSI encoder sequentially includes a first 1×1 convolutional layer, a BN layer, a spatial feature extractor, and a spectral feature extractor; The enhanced hyperspectral image HSI in each patch obtained in step 2 is sequentially input into the first 1×1 convolutional layer and the BN layer, and the BN layer outputs feature A; The feature A output by the BN layer is input into the spatial feature extractor, and the spatial feature extractor outputs feature B; The feature B output by the spatial feature extractor is input into the spectral feature extractor, and the spectral feature extractor outputs feature C; The feature C is sequentially input into the fully connected layer and the softmax layer, and the softmax layer outputs the classification result.
7. A text-driven hyperspectral image ground object classification method according to claim 6, characterized in that: The spatial feature extractor includes: A first 7×7 depth convolutional layer, a first LN layer, a second 1×1 convolutional layer, a first GELU, a first GRN, a third 1×1 convolutional layer, a second 7×7 depth convolutional layer, a second LN layer, a fourth 1×1 convolutional layer, a second GELU, a second GRN, a fifth 1×1 convolutional layer, a third 7×7 depth convolutional layer, a third LN layer, a sixth 1×1 convolutional layer, a third GELU, a third GRN, and a seventh 1×1 convolutional layer; The working process of the spatial feature extractor is as follows: The feature A output by the BN layer is sequentially input into the first 7×7 depth convolutional layer, the first LN layer, the second 1×1 convolutional layer, the first GELU, the first GRN, and the third 1×1 convolutional layer, and the third 1×1 convolutional layer outputs feature A'; The feature A' output by the third 1×1 convolutional layer is element-wise added to the feature A output by the BN layer to obtain feature A"; The feature A" is sequentially input into the second 7×7 depth convolutional layer, the second LN layer, the fourth 1×1 convolutional layer, the second GELU, the second GRN, and the fifth 1×1 convolutional layer, and the fifth 1×1 convolutional layer outputs feature A'''; The output feature A″′ of the fifth 1×1 convolutional layer is element-wise added to the feature A″ to obtain the feature Feature Input the third 7×7 depth convolution layer, the third LN layer, the sixth 1×1 convolution layer, the third GELU, the third GRN, and the seventh 1×1 convolution layer in sequence. The seventh 1×1 convolution layer outputs features Output features of the seventh 1×1 convolutional layer and the features are element-wise added to obtain Feature B; Feature B is the output feature of the spatial feature extractor.
8. A text-driven hyperspectral image ground object classification method according to claim 7, characterized in that: The spectral feature extractor includes: A fourth LN layer, a multi-head self-attention mechanism MHSA, a fifth LN layer, and a feed-forward neural network; The working process of the spectral feature extractor is as follows: The feature B output by the spatial feature extractor is input into the fourth LN layer, and the fourth LN layer outputs feature B'; The feature B' output by the fourth LN layer is input into the multi-head self-attention mechanism MHSA, and the multi-head self-attention mechanism MHSA outputs feature B"; The feature B" output by the multi-head self-attention mechanism MHSA and the feature B output by the spatial feature extractor are input into the fifth LN layer, and the fifth LN layer outputs feature B'''; The output features B″′ of the fifth LN layer are input into a feed-forward neural network (FFN), and the feed-forward neural network (FFN) outputs features The output features of the multi-layer perceptron MLP The output features B of the spatial feature extractor and the output features B″ of the multi-head self-attention mechanism MHSA are element-wise added to obtain feature C; The output feature C of the spectral feature extractor is used as the output feature of the spectral feature extractor.
9. A text-driven hyperspectral image ground object classification method according to claim 8, characterized in that: In step six, the text encoder and the HSI encoder are trained to obtain the trained text encoder and HSI encoder. The specific process is as follows: Calculate the integrated HSI and text contrast loss function based on the semantic features of the text description information output by the text encoder in step four and the image features output by the HSI encoder in step five s(F HSI ,F text ) = norm(F HSI ) T norm(F text ) Among them, N represents the total number of hyperspectral images HSI, i represents the i-th one; k represents the k-th one; i = 1, 2, …, N; k = 1, 2, …, N; Denote intermediate variables; Denote intermediate variables; Indicates the input of the i-th hyperspectral image HSI to the HSI encoder and the image features output by the HSI encoder; Denote the k-th hyperspectral image HSI input to the HSI encoder and the image features output by the HSI encoder; Represent the semantic features of the text description information output by the text encoder for the 15 pieces of text description information of the ground object categories corresponding to the i-th hyperspectral image HSI input to the text encoder. Represent the semantic features of the text description information output by the text encoder for the 15 pieces of text description information of the ground object categories corresponding to the k-th hyperspectral image HSI input to the text encoder. s(F i hsi ,F i text ) represents the similarity between the image features corresponding to the i-th hyperspectral image HSI and the semantic features corresponding to the i-th hyperspectral image HSI; Indicates the similarity between the image features corresponding to the i-th hyperspectral image HSI and the semantic features corresponding to the k-th hyperspectral image HSI; s(F i text ,F i hsi ) represents the similarity between the semantic features corresponding to the i-th hyperspectral image HSI and the image features corresponding to the i-th hyperspectral image HSI; Indicates the similarity between the semantic features corresponding to the i-th hyperspectral image HSI and the image features corresponding to the k-th hyperspectral image HSI; s(F HSI ,F text ) represents the similarity between the image features corresponding to the hyperspectral image HSI and the semantic features corresponding to the hyperspectral image HSI; F HSI Represents the image features output by the HSI encoder in step five; F text Represents the semantic features of the text description information output by the text encoder in step four; norm(F HSI ) represents normalizing the image features output by the HSI encoder in step five; norm(F text ) represents normalizing the semantic features of the text description information output by the text encoder in step four; The superscript T represents the transpose; τ represents a temperature hyperparameter; Based on the classification result p(x ic ) output in Step Five, calculate the HSI classification loss function Among them, C represents the total number of categories; y ic Indicates the true value of whether the i-th sample belongs to the c-th ground object category; p(x ic ) represents the probability of predicting that the i-th sample belongs to the c-th ground object category; Based on the integrated HSI and text contrast loss function and the HSI classification loss function Calculate the total loss function Until the total loss function converges to obtain the trained text encoder and HSI encoder.
10. A text-driven hyperspectral image ground object classification method according to claim 9, characterized in that: The above-mentioned integrated HSI and text contrast loss function and HSI classification loss function to calculate the total loss function is expressed as: