A Multimodal Remote Sensing Data Classification Method Based on a Selective State Space Model of Linear Time Series
Through cross-modal information interaction based on linear time series selective state space model, the problem of insufficient global feature extraction in multimodal remote sensing data classification is solved, and more efficient information capture and classification effects are achieved.
Patent Information
- Application Number
- CN202411183605.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-08-27
AI Technical Summary
The multimodal remote sensing data classification method in the prior art lacks the ability to extract global feature, especially in the extraction of long-range dependencies.
Using a linear time series selective state space model, a cross-modal space fusion module and a multi-layer perceptron are constructed, and information interaction is performed using one-dimensional convolution and SiLU activation function to realize cross-modal feature extraction and classification of hyperspectral and lidar data.
The global modeling capability of multimodal remote sensing data classification model is improved, the modeling constraints of convolutional neural networks are alleviated, and more efficient information capture and classification effects are achieved.
Smart Images

Figure CN119152366B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for classifying multi-modal remote sensing data, and in particular to a method for classifying multi-modal remote sensing data based on a linear time series selective state space model, belonging to the technical field of image processing. Background Art
[0002] In the field of remote sensing, multi-modal usually refers to the imaging results of scenes and targets obtained under different sensors, such as multi-spectral, hyperspectral, synthetic aperture radar, and lidar. Reasonably using multi-modal remote sensing data can provide more comprehensive descriptive information for ground objects in terms of spectrum, time, and space, thereby improving the interpretation ability of remote sensing data to meet the needs of practical applications such as military reconnaissance and smart agriculture.
[0003] To promote the in-depth and extensive application of multi-modal remote sensing data in the above applications, effective information processing means are required. Classification, as one of the important links in remote sensing data processing technology, has always been a hot issue in current research. In recent years, due to its powerful feature extraction ability, deep learning has become the mainstream method for classifying multi-modal remote sensing data. Among many deep learning methods, the method of multi-modal remote sensing data based on the CNN model has attracted much attention, but the extraction of long-range dependencies is lacking. Therefore, a method for classifying multi-modal remote sensing data with high-efficiency global feature extraction ability is needed. Summary of the Invention
[0004] A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present certain concepts in a simplified form as a prelude to the more detailed description to follow.
[0005] In view of this, to solve the deficiency of the existing method for classifying multi-modal remote sensing data in terms of global feature extraction ability, the present invention provides a method for classifying multi-modal remote sensing data based on a linear time series selective state space model.
[0006] The technical solution is as follows: A method for classifying multi-modal remote sensing data based on a linear time series selective state space model, comprising the following steps:
[0007] S1. Obtain multi-modal remote sensing data;
[0008] S2. Establish a mapping layer based on a multi-layer perceptron for different modalities, input the multi-modal remote sensing data, divide the output of the multi-modal remote sensing data into different sample blocks, add the position encoding to the sample blocks to obtain a sequential multi-modal remote sensing representation vector;
[0009] S3. Construct a cross-modal spatial fusion module. The multi-modal remote sensing representation vectors of the input sequence are subjected to spatial information interaction in different modalities to obtain the output of the cross-modal spatial fusion module;
[0010] S4. Iterate step S3 N times, and obtain the classification result by taking the average of the outputs of all cross-modal spatial fusion modules and using a multi-layer perceptron.
[0011] Furthermore, the multi-modal remote sensing data includes hyperspectral data D HSI and ground surface model D LiDAR , and the ground surface model D LiDAR is obtained by denoising and rasterizing the data collected by lidar, is a real number set with length H, width W, and spectral number L, is a real number set with length H and width W.
[0012] Furthermore, the output of the multi-modal remote sensing data includes the output O HSI of hyperspectral data D HSI and the output O LiDAR of ground surface model D LiDAR , and the multi-modal remote sensing representation vectors of the sequence include the final output O HSI of hyperspectral data D HSI' and the final output O LiDAR of ground surface model D LiDAR' .
[0013] Furthermore, construct a cross-modal spatial fusion module based on a linear time series selective state space model, and input the output O HSI of the final hyperspectral data D HSI' and the output O LiDAR of the final ground surface model D LiDAR' into the one-dimensional convolution and SiLU activation function of the cross-modal spatial fusion module respectively to obtain the first output O' HSI and the second output O' LiDAR ;
[0014] O' HSI =σ(Conv(MLP(O HSI' ))) (1)
[0015] O' LiDAR =σ(Conv(MLP(O LiDAR' ))) (2)
[0016] where Conv is a one-dimensional convolution operation, MLP is a multi-layer perceptron, σ represents the SiLU activation function, and z HSI represents the final hyperspectral data DHSI Output O HSI' Output, z LiDAR Indicates the final ground surface model D LiDAR Output O LiDAR' Output;
[0017] Input the first output O' HSI and the second output O' LiDAR into the linear time series selective state space block respectively. The first output O' HSI outputs the first parameter B HSI , the second parameter C HSI and the time scale parameter Δ of the hyperspectral modality HSI , and the second output O' LiDAR outputs the third parameter B LiDAR , the fourth parameter C LiDAR and the time scale parameter Δ of the lidar modality LiDAR , and perform spatial information interaction of different modalities on the outputs obtained from the first output O' HSI and the second output O' LiDAR after passing through the linear time series selective state space block;
[0018] The spatial information interaction process of different modalities is expressed as:
[0019] B HSI , C HSI , Δ HSI = MLP(O' HSI ) (3)
[0020] B LiDAR , C LiDAR , Δ LiDAR = MLP(O' LiDAR ) (4)
[0021]
[0022]
[0023]
[0024] y = (y HSI ⊙ z HSI ) ⊙ (y LiDAR ) ⊙ σ(MLP(O LiDAR )) (8)
[0025] where is the discrete parameter corresponding to the state equation for extracting HSI, is BHSI The corresponding discrete parameter, ZOH is the zero-order hold, A is a continuous variable, y LiDAR is the hybrid feature of LiDAR and HSI based on the output of the linear time series selective state space block, is used to extract the corresponding discrete parameter in the state equation of LiDAR, h t-1 is the system state at time t-1, is B LiDAR The corresponding discrete parameter, y HSI is the hybrid feature of HSI and LiDAR based on the output of the linear time series selective state space block, y is the output of the cross-modal space fusion module, ⊙ is the Hadamard product, HSI is the hyperspectral image, and LiDAR is the lidar data.
[0026] The beneficial effects of the present invention are as follows: Compared with the multi-modal remote sensing data method based on the CNN model, the present invention alleviates the modeling constraints of the convolutional neural network and improves the global modeling ability of the multi-modal remote sensing data classification model through the global receptive field and dynamic weighting; based on the state space model, the present invention uses the linear time series modeling method to capture relevant information more efficiently and effectively across long sequences through selective state space, and realizes efficient channel information interaction between modalities by learning channel dependencies with shared weights. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0028] Figure 1 is a schematic flow chart of a multi-modal remote sensing data classification method based on a linear time series selective state space model;
[0029] Figure 2 is a schematic flow chart of an embodiment of a multi-modal remote sensing data classification method based on a linear time series selective state space model. DETAILED DESCRIPTION OF THE INVENTION
[0030] In order to make the technical solutions and advantages in the embodiments of the present invention clearer, the following further describes the exemplary embodiments of the present invention in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0031] Refer to Figure 1 and Figure 2Specifically describe this embodiment. A multi-modal remote sensing data classification method based on a linear time series selective state space model specifically includes the following steps:
[0032] S1. Obtain multi-modal remote sensing data;
[0033] S2. Establish a mapping layer based on a multi-layer perceptron for different modalities. Input the multi-modal remote sensing data, divide the output of the multi-modal remote sensing data into different sample blocks, add the position encoding to the sample blocks to obtain a sequential multi-modal remote sensing representation vector;
[0034] S3. Construct a cross-modal spatial fusion module. Input the sequential multi-modal remote sensing representation vector, and through spatial information interaction of different modalities, obtain the output of the cross-modal spatial fusion module;
[0035] S4. Iterate step S3 N times, take the average of the outputs of all cross-modal spatial fusion modules, and use a multi-layer perceptron to output the classification result;
[0036] Specifically, the position encoding is obtained by adding a vector with random integer values between 0, 1, 2,..., N to each sample block, where N is the number of sample blocks, providing information about the position for the model.
[0037] Further, the multi-modal remote sensing data includes hyperspectral data D HSI and a ground surface model D LiDAR , and the ground surface model D LiDAR is obtained by denoising and rasterizing the data collected by lidar, is a set of real numbers with length H, width W, and spectral number L, is a set of real numbers with length H and width W.
[0038] Further, the output of the multi-modal remote sensing data includes the output O HSI of the hyperspectral data D HSI and the output O LiDAR of the ground surface model D LiDAR , and the sequential multi-modal remote sensing representation vector includes the final output O HSI of the hyperspectral data D HSI' and the final output O LiDAR of the ground surface model D LiDAR' .
[0039] Further, construct a cross-modal spatial fusion module based on a linear time series selective state space model, and combine the final output O HSI of the hyperspectral data D HSI' and the final output O LiDAR of the ground surface model D LiDAR'Input them into the one-dimensional convolution and SiLU activation function of the cross-modal spatial fusion module respectively to obtain the first output O'. HSI and the second output O'. LiDAR ;
[0040] O' HSI = σ(Conv(MLP(O HSI' ))) (1)
[0041] O' LiDAR = σ(Conv(MLP(O LiDAR' ))) (2)
[0042] where Conv is the one-dimensional convolution operation, MLP is the multi-layer perceptron, σ represents the SiLU activation function, z HSI represents the output O HSI of the final hyperspectral data D HSI' of the output, z LiDAR represents the output O LiDAR of the final ground surface model D LiDAR' of the output;
[0043] Input the first output O' HSI and the second output O' LiDAR into the linear time series selective state space block respectively. After passing through the linear time series selective state space block, the first output O' HSI outputs the first parameter B HSI , the second parameter C HSI and the time scale parameter Δ HSI of the hyperspectral modality. After passing through the linear time series selective state space block, the second output O' LiDAR outputs the third parameter B LiDAR , the fourth parameter C LiDAR and the time scale parameter Δ LiDAR of the lidar modality. Perform spatial information interaction of different modalities on the outputs obtained after passing the first output O' HSI and the second output O' LiDAR through the linear time series selective state space block;
[0044] The spatial information interaction process of different modalities is expressed as:
[0045] B HSI ,C HSI ,Δ HSI = MLP(O' HSI ) (3)
[0046] B LiDAR ,C LiDAR ,Δ LiDAR = MLP(O'LiDAR ) (4)
[0047]
[0048]
[0049]
[0050] y = (y HSI ⊙z HSI )⊙(y LiDAR ⊙σ(MLP(O LiDAR )) (8)
[0051] Among them, is the discrete parameter corresponding to the state equation for extracting HSI, is B HSI corresponding discrete parameter, ZOH is the zero-order hold, A is a continuous variable, y LiDAR is the hybrid feature of LiDAR and HSI based on the output of the linear time series selective state space block, is the discrete parameter corresponding to the state equation for extracting LiDAR, h t-1 is the system state at time t - 1, is B LiDAR corresponding discrete parameter, y HSI is the hybrid feature of HSI and LiDAR based on the output of the linear time series selective state space block, y is the output of the cross-modal space fusion module, ⊙ is the Hadamard product, HSI is the hyperspectral image, and LiDAR is the lidar data;
[0052] Specifically, referring to Figure 2 , the hyperspectral image HSI and lidar data LiDAR are respectively divided into sample blocks. After the sample blocks are superimposed with the corresponding position encodings, they are input into the cross-modal space fusion module. During the cross-modal space fusion process, Liner means maintaining linearity. The hyperspectral image HSI and lidar data LiDAR respectively undergo a one-dimensional convolution operation conv and an activation function SiLU to obtain the time scale parameter Δ HSI of the hyperspectral modality, the first parameter B HSI , the second parameter C HSI , the time scale parameter Δ LiDAR of the lidar modality, the third parameter B LiDAR , the fourth parameter C LiDAR . represents the outer product of matrices. The above process is repeated N times to output the classification result map.
[0053] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art, having the benefit of the foregoing description, will appreciate that other embodiments can be contemplated within the scope of the invention as thus described. Further, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and not to limit or circumscribe the inventive subject matter. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure herein is illustrative, and not restrictive, the scope of the invention being defined by the appended claims.
Claims
1. A multi-modal remote sensing data classification method based on a selective state space model of linear time series, characterized in that, It includes the following steps: S1. Obtain multi-modal remote sensing data; S2. Establish mapping layers based on multi-layer perceptrons for different modalities, input the multi-modal remote sensing data, divide the output of the multi-modal remote sensing data into different sample blocks, add the positional encoding to the sample blocks to obtain the sequence of multi-modal remote sensing representation vectors; S3. Construct a cross-modal spatial fusion module, input the sequence of multi-modal remote sensing representation vectors, and obtain the output of the cross-modal spatial fusion module through spatial information interaction of different modalities; S4. Iterate step S3 N times, take the average of the outputs of all cross-modal spatial fusion modules, and use a multi-layer perceptron to output the classification result; In the above S1, the multi-modal remote sensing data includes hyperspectral data D HSI and ground surface model D LiDAR , the ground surface model D LiDAR is obtained by denoising and rasterizing the data collected by lidar, is a real number set with length H, width W, and spectral number L, is a real number set with length H and width W; In the above S2, the output of the multimodal remote sensing data includes the output O of the hyperspectral data D HSI and the output O of the ground surface model D HSI ; the multimodal remote sensing characterization vector of the sequence includes the output O of the final hyperspectral data D LiDAR and the output O of the final ground surface model D LiDAR ; HSI HSI' LiDAR LiDAR' ; In S3, a cross-modal spatial fusion module based on a linear time series selective state space model is constructed, and the final hyperspectral data D HSI output O HSI' and the output O LiDAR of the final ground surface model D LiDAR' are respectively input into the one-dimensional convolution and SiLU activation function of the cross-modal spatial fusion module to obtain the first output O' HSI and the second output O' LiDAR ; O' HSI = σ(Conv(MLP(O HSI' ))) (1) O' LiDAR = σ(Conv(MLP(O LiDAR' ))) (2) Among them, Conv is a one-dimensional convolution operation, MLP is a multi-layer perceptron, σ represents the SiLU activation function, and z HSI represents the output O HSI of the final hyperspectral data D HSI' ; z LiDAR represents the output O LiDAR of the final ground surface model D LiDAR' ; Input the first output O' HSI and the second output O' LiDAR into the linear time series selective state space block respectively. After passing through the linear time series selective state space block, the first output O' HSI outputs the first parameter B HSI , the second parameter C HSI and the time scale parameter Δ of the hyperspectral modality HSI . After passing through the linear time series selective state space block, the second output O' LiDAR outputs the third parameter B LiDAR , the fourth parameter C LiDAR and the time scale parameter Δ of the lidar modality LiDAR . Perform spatial information interaction of different modalities on the outputs obtained after passing the first output O' HSI and the second output O' LiDAR through the linear time series selective state space block; The spatial information interaction process of different modalities is expressed as: B HSI , C HSI , Δ HSI = MLP(O' HSI ) (3) B LiDAR , C LiDAR , Δ LiDAR = MLP(O' LiDAR ) (4) y = (y HSI ⊙ z HSI ) ⊙ (y LiDAR ⊙ σ(MLP(O LiDAR )) (8) Among them, is the discrete parameter corresponding to the state equation for extracting HSI, is the discrete parameter corresponding to B HSI , ZOH is the zero-order hold, A is a continuous variable, y LiDAR is the hybrid feature of LiDAR and HSI based on the output of the linear time series selective state space block, is the discrete parameter corresponding to the state equation for extracting LiDAR, h t-1 is the system state at time t - 1, is the discrete parameter corresponding to B LiDAR , y HSI is the hybrid feature of HSI and LiDAR based on the output of the linear time series selective state space block, y is the output of the cross-modal space fusion module, ⊙ is the Hadamard product, HSI is the hyperspectral image, and LiDAR is the lidar data.
Citation Information
Patent Citations
Basic model adaptive method for multi-modal remote sensing data classification
CN117611896A