A multi-modal hyperspectral image classification method and system

By extracting 3D cubes and center pixels from hyperspectral images and other modal images, and combining convolution and Transformer modules, the problem of information fusion in multimodal remote sensing image classification is solved, achieving efficient feature extraction and classification while reducing computational complexity.

CN119888333BActive Publication Date: 2025-12-05XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411955490.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-28
Publication Date
2025-12-05
Estimated Expiration
2044-12-28

AI Technical Summary

Technical Problem

In existing multimodal remote sensing image classification technologies, the data information structures of each modality are complex, making it difficult to fully extract effective information for information fusion. Furthermore, the Transformer network architecture is complex and computationally intensive, hindering its deployment in practical applications.

Method used

The algorithm extracts 3D cubes and center pixels from hyperspectral images and other modal images centered on pixels. It extracts spatial and pixel-level features through image block convolution and pixel convolution modules. It then combines residual cross-attention tokenization module and Transformer hybrid feature module to perform feature fusion, generating single-modality and fused modality tokens, reducing information loss and performing classification.

Benefits of technology

It effectively extracts and fuses spatial, spectral, and geometric features of multimodal remote sensing images, reduces information loss, improves classification performance, and reduces computational complexity, making it suitable for practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888333B_ABST
    Figure CN119888333B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal hyperspectral image classification method and system, based on 3D cube and original center pixel, extract the spatial feature and pixel level feature of hyperspectral image and other modal image, based on image block and center pixel, fully exploit the spatial spectral feature and unique feature of HSI and other modal image, and on this basis, single modal token and mixed modal token are constructed, to the number of row of input feature matrix linearly related with the calculation cost of adaptive fusion multi-modal feature, reduce the loss of information, improve the classification effect, at the same time, the token generated not only includes each single modal token, also includes multi-modal fusion token, comprehensively considers various feature fusion conditions, maximum possibly reduces the information loss brought in feature fusion process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image classification, and relates to a multi-modal hyperspectral image classification method and system. BACKGROUND

[0002] In response to global environmental challenges such as climate change and desertification, detailed observation of the Earth becomes crucial. With the growing demand, remote sensing technology has become an important tool for scientific research and environmental monitoring due to its unique advantages. In recent years, with the development of remote sensing technology, remote sensing image classification has gradually become one of the hotspots in the field of remote sensing image research. The goal of remote sensing image classification is to assign land cover classes to each pixel in a remote sensing image, where each class represents a specific land object, such as water, vegetation, urban buildings, etc. Remote sensing image classification has been widely used in forestry, mineral exploration and mapping, target detection, environmental monitoring, urban planning, biodiversity protection, and disaster response and management.

[0003] In order to further improve the accuracy of remote sensing image classification, different sensor data is introduced for classification based on hyperspectral image classification, which is called multi-modal remote sensing image classification. Specifically, for the same geographical area, different sensors are used for data capture, and the data obtained can provide different feature information for the same observed land cover area. Each modality can provide unique and complementary information to HSI: LiDAR collects depth and intensity information, measures the height of objects on the Earth's surface, SAR data provides amplitude and phase geometry information, and DSM can provide three-dimensional morphology of the Earth's surface. Through deep learning, multi-modal remote sensing image classification needs to extract deep learning features of HSI and other modal images, and then input the features into the classifier for prediction. The main challenges of this task are: 1) LiDAR, SAR and DSM and other modal remote sensing images contain unique information. How to extract reliable information from them and use it for subsequent information fusion is a big problem. 2) Due to its unique network architecture, Transformer can flexibly integrate different types of data, which makes it very suitable for performing multi-modal tasks. How to use this feature to assist the method in the fusion of multi-modal remote sensing image features while ensuring that information is not lost too much is a problem worth considering.

[0004] Traditional methods fuse the morphological features (such as morphological features and attribute features) of HSI and LiDAR through subspace learning and other ways to enhance the classification effect. In addition, global fusion techniques such as extinction features (EPs) are also used to improve the extraction ability of multi-modal features. Traditional classifiers such as support vector machine (SVM) show good potential in combination with multiple features and multi-source data environment. The emergence of deep learning models has greatly improved the classification ability of multi-modal remote sensing data, especially the performance of convolutional neural network (CNN) in feature extraction and fusion. Through structures such as 1-D, 2-D and 3-D CNN, spatial and spectral features can be effectively extracted from HSI, while elevation information can be obtained from LiDAR data, and multi-modal fusion at the feature level and decision level can be further optimized by using parameter sharing and other strategies to further optimize the classification performance.

[0005] In recent years, the Transformer model has gradually become a core technology in multi-modal remote sensing image classification due to its long-distance dependency modeling capability. The Transformer network processes HSI and LiDAR data through multi-head self-attention mechanism (MSA) to achieve feature extraction and deep fusion. To improve the complementarity and compatibility between different modalities, multi-scale self-attention and cross-attention modules are used to fuse modal features. In addition, the Transformer encoder further improves the fusion effect between HSI and other modal data by combining the modal comprehensive features generated by the modal attention module. In the prior art, although the introduction of data from different sensors can provide more rich feature information for the same geographic area, each modality data (such as HSI, LiDAR, SAR, and DSM) contains unique and complementary information, and the information structure is complex, making it difficult to fully extract effective information for subsequent information fusion. Secondly, how to effectively extract and fuse spatial, spectral and geometric features in multi-modal data is still a difficult problem in the classification task. In addition, although the Transformer has the advantage of flexible integration of multi-modal data, its network architecture is complex and the computational load is large, which hinders their deployment in practical applications. SUMMARY

[0006] The present application aims to solve the problems in the prior art that the information structure of each modality data is complex, it is difficult to fully extract effective information for subsequent information fusion, it is difficult to effectively extract and fuse spatial, spectral and geometric features, and the network architecture is complex and the computational load is large, which hinders their deployment in practical applications, and provides a multi-modal hyperspectral image classification method and system.

[0007] To achieve the above-mentioned purpose, the technical scheme is adopted as follows:

[0008] A multi-modal hyperspectral image classification method comprises the following steps:

[0009] A 3D cube and a raw central pixel of the original hyperspectral image data are extracted pixel by pixel, and a 3D cube and a raw central pixel of other modal image data are extracted;

[0010] Based on the 3D cube and the raw central pixel, spatial features and pixel-level features of the hyperspectral image and the other modal image are extracted;

[0011] The spatial features and the pixel-level features are converted into a token set containing a single modal and a fusion modal, and the token set is fused;

[0012] The fused token set is classified to obtain a classification result.

[0013] Further improvements of the present application are:

[0014] The 3D cube of the original hyperspectral image data and the other modal image data is extracted pixel by pixel, comprising:

[0015]

[0016] The raw central pixel of the original hyperspectral image data and the other modal image data comprises:

[0017]

[0018] In the formula, PT M1 represents a hyperspectral image block; PT M2 represents an other modal image block; PX M1 the central pixel of the hyperspectral image; PX M2 the central pixel of the other modal image; M1 represents HSI, M2 represents other modal remote sensing data, P is the width and height of the image data, and C is the channel number of the hyperspectral image.

[0019] Based on PT M1 and PT M2 spatial spectral features of the HSI and spatial features of the other modal image are extracted through an image block convolution module;

[0020] Based on PX M1 and PX M2 detailed spectral features of the HSI and pixel-level unique features of the other modal image are extracted through a pixel convolution module.

[0021] The extraction of the spatial features of the hyperspectral image and the other modal image comprises the following steps:

[0022]

[0023] wherein PT M1 denotes a hyperspectral image patch; PT M2 denotes other modality image patch; Reshape(·) denotes a reshaping operation; Conv2D(·) denotes a two-dimensional convolution layer, a batch normalization layer and a ReLU activation layer; Conv3D(·) denotes a three-dimensional convolution layer, a batch normalization layer and a ReLU activation layer.

[0024] The step of extracting pixel-level features of the hyperspectral image and the other modality image comprises the following steps:

[0025]

[0026] wherein PX Mi denotes a center pixel, wherein Mi denotes the hyperspectral image M1 or the other modality image M2; Linear(·) denotes a linear layer with an output dimension d; Conv1D(·) denotes a one-dimensional convolution layer, a batch normalization layer and a ReLU activation layer.

[0027] The step of converting the spatial features and the pixel-level features into a token set comprising single modality and fusion modality comprises:

[0028] The single modality token T 11 generated by combining the hyperspectral image and the hyperspectral image group;

[0029] The fusion modality token T 12 generated by combining the hyperspectral image and the other modality image;

[0030] The fusion modality token T 21 generated by combining the other modality image and the hyperspectral image;

[0031] The single modality token T 22 generated by combining the other modality image and the other modality image;

[0032] The single modality token and the multi-modal fusion token are further fused by a Transformer mixing feature module, and the result is subjected to layer normalization and global average pooling to obtain a fusion result.

[0033] The token set after fusion is classified using a multi-layer perceptron.

[0034] A multi-modal hyperspectral image classification system comprises:

[0035] A data processing module is configured to extract a 3D cube of original hyperspectral image data and an original center pixel, and extract a 3D cube of other modality image data and an original center pixel, both on a pixel basis;

[0036] a feature extraction module configured to extract spatial features and pixel-level features of the hyperspectral image and other modal images based on the 3D cube and the original center pixel;

[0037] a fusion module configured to convert the spatial features and the pixel-level features into a token set containing single modal and fusion modal, and fuse the token set;

[0038] a classification module configured to classify the fused token set to obtain a classification result.

[0039] A terminal device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of the embodiments of the application when executing the computer program.

[0040] A computer readable storage medium stores a computer program, and the computer program is executable by a processor to implement the steps of the method according to any one of the embodiments of the application.

[0041] Compared with the prior art, the present application has the following beneficial effects:

[0042] The application discloses a multi-modal hyperspectral image classification method, which extracts spatial features and pixel-level features of a hyperspectral image and other modal images based on a 3D cube and an original center pixel, fully mines spatial spectral features and unique features of the HSI and other modal images based on an image block and the center pixel, and constructs single modal tokens and mixed modal tokens on the basis, so as to adaptively fuse multi-modal features with a linear correlation with the number of rows of an input feature matrix at a calculation cost, reduce information loss, improve classification effect, and meanwhile, the generated tokens contain not only single modal tokens but also multi-modal fusion tokens, comprehensively consider various feature fusion conditions, and maximally reduce information loss in the feature fusion process. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the application, and therefore should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0044] Figure 1 It is a model structure diagram of the application;

[0045] Figure 2 It is a residual cross-attention tokenization module structure diagram of the application;

[0046] Figure 3 The flow chart of the present application. DETAILED DESCRIPTION

[0047] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.

[0048] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.

[0049] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0050] In the description of the embodiments of the present application, it should be noted that, if the terms "upper", "lower", "horizontal", "inner" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship when the product of the present application is used, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second" and the like are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0051] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly inclined. For example, "horizontal" only means that its direction is relatively more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.

[0052] In the description of the embodiments of the present application, it should also be noted that, unless otherwise explicitly specified and limited, if the terms "arrangement", "installation", "connection", "connection" appear, they should be understood in a broad sense, for example, can be fixedly connected, can be detachably connected, or integrally connected; can be mechanically connected, can be electrically connected; can be directly connected, can be indirectly connected through an intermediate medium, or can be connected inside two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0053] The application will be described in further detail below with reference to the drawings:

[0054] Referring to Figures 1 to 3 The embodiment of the application discloses a multi-modal hyperspectral image classification method, which fully mines the features of HSI and other modal images, and meanwhile, well fuses the features, reduces the loss of information, and makes the hyperspectral image classification achieve good results.

[0055] The method introduces four effective modules, including an image block convolution module, a pixel convolution module, a residual cross-attention tokenization module and a Transformer mixed feature module, wherein:

[0056] The image block convolution module focuses on image blocks, extracts spatial spectral features of the hyperspectral image and spatial features of other modal images.

[0057] The pixel convolution module focuses on the center pixel, extracts detailed spectral features of the hyperspectral image and unique pixel-level features of other modal images.

[0058] The residual cross-attention tokenization module is used for fusing and converting the above features into single modal tokens and fusion modal tokens, which not only ensures the fusion of features, but also retains the features of each modal alone, avoiding the loss of information.

[0059] The Transformer mixed feature module is used for establishing long-distance dependence between tokens.

[0060] The above four modules are sequentially applied to predict the label of each pixel, further improving the classification performance.

[0061] Specifically, the following steps are included:

[0062] Step 1: pre-processing the input hyperspectral image and other modal images, extracting a 3D cube centered on a pixel and an original center pixel, obtaining processed training set and test set data, specifically including:

[0063] For the input hyperspectral image and other modal images Normalizing each channel, and cutting out image blocks corresponding to the pixels to be classified from the images of the two modalities respectively:

[0064]

[0065] Original center pixel:

[0066]

[0067] where M1 represents HSI, M2 represents other modal remote sensing data, P is the width and height of the image data, and C is the number of channels of the hyperspectral image.

[0068] Further, the other modal data includes LiDAR, SAR or DSM image data.

[0069] Step 2: The pre-processing module output is transmitted into the image block convolution module and the pixel convolution module for extracting spatial features and pixel-level features of the hyperspectral image and the other modal image, specifically including:

[0070] The image block convolution module uses different CNNs to extract spatial spectral features of the HSI (hyperspectral image) and spatial features of the other modal image, respectively.

[0071] The target of the pixel convolution module is to focus on mining the pixel-level features of the multi-modal remote sensing image.

[0072] Specifically, the network schematic diagram of the image block convolution module is as shown in Figure 1 For the HSI block PT M1 , it is first reshaped to the size of 1×C M1 ×p×p, and then input into a three-dimensional convolution layer with an output channel of 8, a convolution kernel size of 3×3×3, a step size of 1 and a pooling size of 0×1×1. The output feature with the size of 8×(CM1-2)×p×p is obtained. Then, the first dimension and the second dimension are merged to obtain a feature matrix with the size of (8CM1-16)×p×p. The matrix is input into a two-dimensional convolution with an output channel of d, a convolution kernel size of 3×3, a step size of 1 and a pooling size of 0, and a feature matrix with the size of d×p′×p′ is obtained, where p′=p-2. The matrix is further reshaped to (p′×p′)×d, which is the spatial spectral feature of the HSI block For the other modal image block PT M2 , it is first reshaped to a cube with the size of C M2 ×p×p, and then input into a two-dimensional convolution with an output channel of d, a convolution kernel size of 3×3, a step size of 1 and a pooling size of 0, and a feature matrix with the size of d×p′×p′ is obtained. Finally, the matrix is reshaped to a feature matrix with the size of (p′×p′)×d, which represents the spatial feature of the other modal.

[0073] The above process can be expressed as:

[0074]

[0075] ​In the formula: Reshape(·) – reshaping operation; Conv2D(·) – contains a 2D convolutional layer, a batch normalization layer, and a ReLU activation layer; Conv3D(·) – contains a 3D convolutional layer, a batch normalization layer, and a ReLU activation layer. Ultimately, the two image patch-level features, namely the spatial spectral features of HSI, are... Spatial features of other modal images This is the output of this module.

[0076] Furthermore, the network diagram of the pixel convolution module is as follows: Figure 1 As shown. Center pixels of HSI and other modalities. and As input to this module, it is first reshaped into a size of 1×C. M1 and 1×C M2 A one-dimensional vector. Then, it is input into a one-dimensional convolutional layer with 4 output channels, a kernel size of 3, a stride of 1, and a pooling size of 1, where convolutions are performed along the channel dimension, yielding 4×C. M1 and 4×C M2 The size feature is then fed into a batch normalization layer and a ReLU activation layer. Next, the features are concatenated in the second dimension, reshaping the data into a 1×4C size. M1 and 1×4C M2 The feature vectors are then fed into a linear layer with output dimension d to obtain the output. and The above process can be expressed as:

[0077]

[0078] In the formula: Mi—M1 or M2; Linear(·)—a linear layer with an output dimension of d; Conv1D(·)—containing a one-dimensional convolutional layer, a batch normalization layer, and a ReLU activation layer. Ultimately, HSI provides detailed spectral features. Unique pixel-level features compared to other modal images This is the output of this module.

[0079] Step 3: Based on spatial and pixel-level features, the residual cross-attention tokenization module converts multimodal features into tokens containing both single and fused modalities. The Transformer hybrid feature module further establishes long-range dependencies on the token set containing multimodal heterogeneous features, specifically including:

[0080] Specifically, the multi-modal features are converted into tokens containing single modal and fusion modal by using the residual cross-attention tokenization module. The long-distance dependencies of the token set containing multi-modal heterogeneous features are further established by the Transformer mixed feature module.

[0081] Further, the residual cross-attention tokenization module is as shown in Figure 2 It is used to further convert the heterogeneous features of each modal into a token set for input into the Transformer for classification. The residual cross-attention tokenization module receives the spatial-spectral features of HSI extracted on the image block the spatial features of other modal images the detailed spectral features of HSI and the unique features of other modal image pixels The corresponding key matrix and value matrix are generated by mapping through the weight matrix, and the residual vector is obtained by using global average pooling. For the pixel-level features and They are mapped into one-dimensional query matrices through weight matrices, and then each query matrix is combined with each pair of key matrix and value matrix one by one to perform cross-attention calculation. Finally, the residual vector is added to form a residual connection to enhance the classification effect, and a series of tokens T = [T 11 ,T 12 ,T 21 ,T 22 ] are obtained as the output of this module. The process can be summarized as:

[0082]

[0083] In the formula: 1D-RCA(·) is one-dimensional residual cross-attention, and the calculation method is as follows:

[0084]

[0085] In the formula: Mj and Mk are M1(HSI) or M2(other modal images); The i-th attention head and i ∈ [1, h]; Query matrix; Key matrix; Value matrix; and Learnable weight matrix, all attention heads are finally concatenated in the second dimension to obtain an output of size 1 × d (d h × h = d). The residual vector R Mj is used to increase the residual connection to overcome gradient explosion, and the generation method is as follows:

[0086]

[0087] wherein by combining Global Average Pooling (GAP) in the first dimension is obtained.

[0088] Since in 1D-RCA, the query matrix is a one-dimensional vector, 1D-RCA has lower computational complexity than ordinary MSA:

[0089] Ω(MSA) = 2p' 4 d + 4p' 2 d 2

[0090] Ω(1D-RCA) = p' 2 d 2 + 3p' 2 d + d 2 + d

[0091] wherein: p' 2 — the number of rows of the feature matrix and Thus, the global MSA is quadratically related to the number of rows of the input feature matrix, while the 1D-RCA is linearly related to the number of rows of the input feature matrix. Therefore, 1D-RCA is more suitable for processing input feature matrices with a large number of rows than MSA.

[0092] Further, the Transformer hybrid feature module is as shown in Figure 1 acts on the input token set to perform one-to-one information fusion, establish long-distance dependence, and better process serialized features. Two Transformer encoders are used to receive the token set T to globally establish long-distance dependence between the input token set, and further fuse single-modal tokens and multi-modal fusion tokens. Finally, the result is subjected to layer normalization and global average pooling to obtain T gap as the final output of the network.

[0093] In the present application, the residual cross-attention tokenization module is used to further convert the heterogeneous features of each modality into a token set for input into the Transformer for classification. The generated tokens not only contain single modality tokens, but also contain multi-modal fusion tokens. Various feature fusion situations are comprehensively considered to minimize the information loss in the feature fusion process. Meanwhile, residual connection is introduced to further retain the original spatial spectral features. In terms of computational complexity, 1D-RCA combined with PXConv makes the query matrix a one-dimensional vector, reducing the computational complexity of 1D-RCA itself. Moreover, the module reconstructs the spatial spectral features into only 4 tokens, greatly reducing the computational burden in the subsequent modules.

[0094] Further, in the present application, the Transformer mixed feature module is composed of multiple connected Transformer encoders, a normalization layer and a global average pooling layer. It can perform one-to-one information fusion on the input token set, establish long-distance dependence, and better handle the serialized features. In order to minimize the loss in multi-modal feature fusion, the Transformer mixed feature module performs global feature re-fusion on the 4 tokens generated by the residual cross-attention tokenization module, which respectively contain different modal combinations. It adaptively extracts decisive features for classification and establishes long-distance dependence between tokens.

[0095] Step 4: input the embedded features into a multilayer perceptron to classify the test samples and obtain the classification results.

[0096] Assuming there are N classes, the output of the Transformer mixed feature module is input into a multilayer perceptron to obtain the classification results The commonly used cross-entropy loss is used to train the proposed method. In the training phase, the cross-entropy loss between Y and the true label is calculated to update the network.

[0097] This embodiment fully excavates the spatial spectral features and unique features of HSI and other modal images based on image blocks and center pixels respectively, and constructs single modality tokens and mixed modality tokens. The multi-modal features are adaptively fused with a linear correlation between the number of input feature matrices and the computational cost, reducing information loss and improving classification effect.

[0098] A multi-modal hyperspectral image classification system, characterized in that it comprises:

[0099] A data processing module for extracting a 3D cube of original hyperspectral image data and an original center pixel, and extracting a 3D cube of other modal image data and an original center pixel based on the center pixel;

[0100] a feature extraction module configured to extract spatial features and pixel-level features of the hyperspectral image and the other modality image based on the 3D cube and the original center pixel;

[0101] a fusion module configured to convert the spatial features and the pixel-level features into a token set containing a single modality and a fusion modality, and fuse the token set;

[0102] a classification module configured to classify the fused token set to obtain a classification result

[0103] An embodiment of the terminal device provided by the present application provides a schematic diagram of the terminal device. The terminal device of the embodiment includes a processor, a memory, and a computer program stored in the memory and executable on the processor. The processor implements the steps in each of the method embodiments when executing the computer program. Alternatively, the processor implements the functions of each module / unit in each of the device embodiments when executing the computer program.

[0104] The computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application.

[0105] The terminal device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and other computing devices. The terminal device can include, but is not limited to, a processor and a memory.

[0106] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0107] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the terminal device by running or executing the computer program and / or modules stored in the memory, and calling the data stored in the memory.

[0108] The modules / units integrated in the terminal device, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0109] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multimodal hyperspectral image classification method, characterized in that, Includes the following steps: Using pixels as the center, extract the 3D cube and the original center pixel of the original hyperspectral image data, as well as the 3D cube and the original center pixel of other modal image data; Based on the 3D cube and the original center pixel, spatial and pixel-level features are extracted from hyperspectral images and other modal images; Spatial features and pixel-level features are transformed into a token set containing single-modality and fused-modality features. The token set is then fused. Multimodal features are transformed into tokens containing single-modality and fused-modality features using the residual cross-attention tokenization module. Long-distance dependencies are established on the token set containing multimodal heterogeneous features using the Transformer hybrid feature module. The residual cross-attention tokenization module is used to further transform the heterogeneous features of each modality into a token set for input into the Transformer for classification. The merged token set is classified, and the classification results are obtained.

2. The multimodal hyperspectral image classification method according to claim 1, characterized in that, The extraction of raw hyperspectral image data and other modal image data into a pixel-centered 3D cube includes: The extracted original center pixels of the original hyperspectral image data and other modal image data include: In the formula, Represents a hyperspectral image patch; Represents other modal image blocks; The center pixel of a hyperspectral image; The center pixel of other modal images; Indicates HSI, Represents other modal remote sensing data, For the width and height of the image data, This represents the number of channels in the hyperspectral image.

3. The multimodal hyperspectral image classification method according to claim 2, characterized in that, based on Spatial spectral features of HSI and spatial features of other modal images are extracted using an image block convolution module; based on The pixel convolution module extracts detailed spectral features of HSI and pixel-level unique features of other modal images.

4. The multimodal hyperspectral image classification method according to claim 1, characterized in that, The extraction of spatial features from hyperspectral images and other modal images includes the following steps: In the formula: Represents a hyperspectral image patch; Represents other modal image blocks; ( ) indicates a reshaping operation; 2 ( This indicates that the layer consists of a 2D convolutional layer, a batch normalization layer, and a ReLU activation layer. 3 ( This indicates that it contains a 3D convolutional layer, a batch normalization layer, and a ReLU activation layer.

5. The multimodal hyperspectral image classification method according to claim 1, characterized in that, The extraction of pixel-level features from hyperspectral images and other modal images includes the following steps: In the formula: Represents the center pixel, where, Represents hyperspectral images 1 or other modal images 2; ( ) represents an output dimension of Linear layers; 1 ( This indicates that it contains a one-dimensional convolutional layer, a batch normalization layer, and a ReLU activation layer.

6. The multimodal hyperspectral image classification method according to claim 1, characterized in that, The process of converting spatial features and pixel-level features into a token set containing single-modality and fused-modality features includes: Single-modal token generated by combining hyperspectral images ; Fusion mode tokens generated by combining hyperspectral images with other modal images ; Fusion modality tokens generated by combining other modal images with hyperspectral images ; Other modal images are combined with other modal images to generate a single modal token. The Transformer hybrid feature module further fuses single-modal tokens and multimodal fusion tokens, and then performs layer normalization and global average pooling on the results to obtain the fused result.

7. The multimodal hyperspectral image classification method according to claim 1, characterized in that, During classification, a multilayer perceptron is used to classify the fused token set.

8. A multimodal hyperspectral image classification system implementing the method of claim 1, characterized in that, include: The data processing module is used to extract the 3D cube and the original center pixel of the original hyperspectral image data, centered on the pixel, as well as to extract the 3D cube and the original center pixel of other modal image data. The feature extraction module is used to extract spatial and pixel-level features from hyperspectral images and other modal images based on a 3D cube and the original center pixel. The fusion module is used to transform spatial features and pixel-level features into a token set containing single modalities and fused modalities, and to fuse the token set; The classification module is used to classify the merged token set and obtain the classification results.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.