Multi-modal collaborative ground feature classification method and system based on mask contrast learning
Through the mask contrast learning method, dynamic adjustment of masking strategy and cross-entropy optimization, the problem of insufficient deep feature mining in the joint classification of HSI and LiDAR data is solved, and the classification accuracy and robustness of multimodal remote sensing data are improved.
Patent Information
- Application Number
- CN202510737756.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-04
AI Technical Summary
When existing technologies use HSI and LiDAR data for joint classification, it is difficult to effectively mine deep features and they are highly dependent on labeled samples, resulting in insufficient classification accuracy and robustness.
A multimodal collaborative feature classification method based on mask contrast learning is adopted. Through multi-scale feature extraction, CGAFT module fusion and cross-entropy loss optimization, the masking strategy is dynamically adjusted to enhance the inter-modal feature extraction capability and reduce the dependence on labeled samples.
The accuracy and robustness of multimodal remote sensing data classification are improved, especially the performance is excellent under the condition of a small number of labeled samples, and the adaptability and stability of the model in complex tasks are enhanced.
Smart Images

Figure CN120612604A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image classification, and in particular relates to a multimodal collaborative object classification method and system based on mask contrast learning. Background Art
[0002] In recent years, remote sensing technology has developed rapidly, and remote sensing imagery plays an irreplaceable role in various applications, such as disaster detection, agricultural management, and urban development planning. With the emergence of various remote sensing data, multimodal learning has gradually attracted the attention of researchers. Among these multimodal data, HSI (Hyperspectral Imaging) data provides detailed spectral information for identifying specific objects on the ground, while LiDAR (Light Detection and Ranging) data provides elevation information of the area. However, HSI faces challenges in identifying ground objects with the same spectral characteristics but different heights, and LiDAR data has difficulty distinguishing ground objects of different materials and at the same height. Joint classification using HSI and LiDAR data can fully utilize their complementary information to improve classification accuracy. Therefore, feature fusion of cross-modal data has attracted widespread attention and has been widely used in multi-source remote sensing imagery ground object classification.
[0003] In recent years, numerous machine learning techniques have been applied to the joint classification of HSI and LiDAR data, including support vector machines (SVMs), random forests, and rotation forests (RoFs). In traditional machine learning frameworks, feature extraction and selection are performed on the two different modalities separately. The features of the two modalities are then fused to form a comprehensive feature set, and finally, a classifier is used to classify the objects. However, traditional methods for applying these techniques rely on the quality of handcrafted features, lack deep feature mining, and cannot effectively fit the complex nonlinear relationships between object features in HSI and LiDAR data, which limits their application scenarios.
[0004] Deep learning-based methods have attracted much attention due to their outstanding capabilities in automatic feature extraction. However, building an effective HSI and LiDAR classification model is not easy. One of the key reasons is that deep learning-based models usually require a large number of labeled samples to achieve satisfactory accuracy, which is expensive and limited in land feature classification. In addition, these methods usually use attention mechanisms to bridge the semantic gap between modalities, focusing on coordinating feature similarities between different modalities by dynamically assigning weights to key information. The association between modalities usually depends on the same class labels in both modalities, which makes it difficult for the attention mechanism to explore deeper relationships between the two modalities.
[0005] Contrastive learning is a self-supervised learning technique designed to extract meaningful representations from unlabeled data. It typically involves using proxy tasks to distinguish between positive and negative samples, automatically acquiring sample features for model training. After extensive research, many well-known contrastive learning frameworks have been discovered, such as SimCLR, MoCo, and BYOL, that can be used to train encoders with strong representation extraction capabilities. These frameworks, such as CMC and FactorCL, have been expanded upon and expanded upon. Contrastive learning methods generally consist of two phases: pre-training and fine-tuning. During the pre-training phase, the input data is augmented with data to construct positive and negative samples as supervisory information to learn features. The high-quality feature extraction modules learned during the pre-training phase are then transferred to downstream tasks for fine-tuning. With the development of contrastive learning, researchers have proposed numerous contrastive learning methods for multimodal remote sensing data, primarily focusing on increasing sample diversity through various data augmentation strategies.
[0006] Masked Image Modeling (MIM), an advanced self-supervised learning method, learns effective visual representations by randomly masking parts of the input image and training the model to recover these masked parts based on the unmasked information. In its implementation, the input image is first segmented into a series of non-overlapping blocks of fixed size. Some of these blocks are then masked to simulate real-world scenarios with incomplete or noisy data. A deep neural network is then used to encode the masked image, enabling the model to infer the content of the occluded areas based on contextual information. This allows the model to not only understand the overall structure and local details of the image, but also capture the complex relationships between different parts.
[0007] In the field of computer vision, the combination of contrastive learning and mask image models offers a new direction for constructing representations with high instance discrimination and local perception capabilities. Contrastive learning aims to optimize the model by reducing the distance between similar inputs (positive pairs) and increasing the distance between dissimilar inputs (negative pairs) in feature space. Traditionally, positive pairs are typically obtained by augmenting the original input into two different views, while negative pairs are obtained using random sampling or memory bank techniques. While effective, this approach often relies on a fixed range of sample augmentation to generate positive pairs, which has certain limitations in flexibility.
[0008] Based on the above methods, Masked Contrastive Learning (MCL) cleverly combines the advantages of contrastive learning and masked image models. Through careful design, MCL not only leverages the masked image model to enable the model to recover the masked content based on contextual information, thereby learning rich knowledge about image structure, texture, shape, and other aspects, but also employs contrastive learning mechanisms to close the distance between similar input representations while pushing apart the representations of dissimilar inputs. This combined approach enables the model to effectively capture the internal details of the image and the complex relationships between its components, thereby forming a more compact and discriminative feature representation.
[0009] In addition, in order to further improve model performance, some studies have explored adaptive masking strategies, that is, dynamically adjusting the masking method according to the characteristics of the input samples and the learning state of the model. This is different from previous methods that only rely on contrastive learning in the pre-training stage. MCL can be implemented synchronously during the supervised training process, avoiding additional pre-training costs and improving computational efficiency. As research continues to deepen, more and more effective improvement methods have been discovered, and these advances have jointly promoted progress in the field of visual representation learning. Therefore, the MCL-based method has demonstrated excellent performance in a variety of downstream tasks, such as image classification, object detection, and semantic segmentation, demonstrating its strong generalization and expression capabilities. The development of this method provides a solid foundation and broad space for future visual analysis technologies and applications. Summary of the Invention
[0010] The present invention proposes a multimodal collaborative object classification method based on masked contrastive learning for multimodal remote sensing data classification, which reduces the reliance of deep learning models on labeled remote sensing samples, enhances their ability to extract complementary features between the two modalities, and improves the accuracy and robustness of object classification. The method includes the following steps:
[0011] S1. Acquire hyperspectral data and LiDAR point cloud data, and preprocess the hyperspectral data and the LiDAR point cloud data to obtain hyperspectral image blocks and LiDAR image blocks;
[0012] S2. Using a multi-scale feature extraction module to perform feature extraction on the hyperspectral image block and the LiDAR image block, respectively, to obtain a hyperspectral image feature block and a LiDAR image feature block;
[0013] S3, using the CGAFT module to fuse the hyperspectral image feature block and the LiDAR image feature block respectively to obtain hyperspectral image features, LiDAR image features, a first cross-modal attention map, and a second cross-modal attention map;
[0014] S4. Obtaining a first mask fusion feature and a second mask fusion feature based on the hyperspectral image feature block, the LiDAR image feature block, the first cross-modal attention map, and the second cross-modal attention map;
[0015] S5. Calculating contrast loss based on the hyperspectral image features, the LiDAR image features, the first mask fusion features, and the second mask fusion features, and updating model parameters in combination with a cross entropy function;
[0016] S6. Based on the model with updated parameters, classify the multimodal data of the objects to be classified to obtain the classification results.
[0017] Further preferably, the preprocessing method includes: spatial alignment, format unification and resampling, enhancement and normalization processing, and slicing operation.
[0018] Further preferably, the method for obtaining the hyperspectral image feature block includes:
[0019] A 3D convolution with a kernel size of 9×3×3 is used to perform a convolution operation on the hyperspectral image block. Then, a set of two-dimensional convolutions with kernel sizes of 1×1, 3×3, and 5×5 are used to simultaneously extract features. Finally, the outputs of the two-dimensional convolution are fused using element-wise addition to obtain the hyperspectral image feature block.
[0020] Further preferably, the method for obtaining the LiDAR image feature block includes:
[0021] The LiDAR image block is convolved using a two-dimensional convolution with a kernel size of 3×3. Then, a set of two-dimensional convolutions with kernel sizes of 1×1, 3×3, and 5×5 are used to simultaneously extract features. Finally, the outputs after the two-dimensional convolution are fused using element-wise addition to obtain the LiDAR image feature block.
[0022] Further preferably, the method for obtaining the hyperspectral image features and the LiDAR image features includes:
[0023] Flattening the hyperspectral image feature block and the LiDAR image feature block and concatenating them with learnable classification labels to generate a first sequence feature and a second feature sequence;
[0024] Adding a learnable position embedding to the first sequence feature and the second feature sequence to obtain the hyperspectral image feature and the LiDAR image feature.
[0025] Further preferably, the method for obtaining the first mask fusion feature and the second mask fusion feature includes:
[0026] Calculating the comprehensive activation strength of each feature position based on the first cross-modal attention map and the second cross-modal attention map;
[0027] Apply the encoder of the CGAFT module N times on the hyperspectral image feature block and the LiDAR image feature block respectively, and then input them into the FC network to obtain the feature representation F of c channels;
[0028] Compute query Q, key K, and value V from feature representation F, and compute attention graph score based on query Q, key K, and value V;
[0029] Converting the attention map score into a probability distribution, where the probability distribution is used to reflect the probability that the feature with high comprehensive activation intensity is masked; combining the Gumble max technique to guide the masking strategy in a probabilistic manner;
[0030] The hyperspectral image feature block and the LiDAR image feature block are masked based on the masking strategy to obtain the first mask fusion feature and the second mask fusion feature.
[0031] Further preferably, the method for updating model parameters based on the contrast loss includes:
[0032] l total =l CE +α·l Contrastive ,
[0033] in,
[0034] l Contrastive =l HSI-LiDAR +l HSI-MaskedHSI +l LiDAR-MaskedLiDAR ,
[0035]
[0036] Where, l total represents the total loss function; l CE represents the cross entropy function; α represents the weight of the contrast loss; l Contrastive represents the total contrast loss; l HSI-liDAR represents the HSI-LiDAR contrast loss; l HSI-MaskedHSI represents HSI-MaskedHSI contrast loss; l LiDAR-MaskedLiDAR represents the HSI-LiDAR contrast loss; τ is the temperature parameter; N b Indicates the number of images in a batch during training; represents the high-dimensional representation generated by the k-th sample in the corresponding batch, Represents the high-dimensional representation generated by the i-th sample in the corresponding batch, I k≠iIndicates that in the traversal of k, if the high-dimensional representation of the k-th sample and the high-dimensional representation of the current i-th sample are high-dimensional representations generated by different samples, it is 1, otherwise it is 0; Represents the LiDAR high-dimensional features generated by the k-th masked sample in the corresponding batch; Represents the LiDAR high-dimensional features generated by the i-th masked sample in the corresponding batch; Represents the LiDAR high-dimensional features generated by the i-th unmasked sample in the corresponding batch.
[0037] The present invention also provides a multimodal collaborative feature classification system based on mask contrast learning, comprising: a pre-training module and a classification module;
[0038] The pre-training module includes: a data processing unit, a feature extraction unit, a fusion processing unit, a mask fusion unit and a loss calculation unit;
[0039] The data processing unit is used to obtain hyperspectral data and LiDAR point cloud data, and preprocess the hyperspectral data and the LiDAR point cloud data to obtain hyperspectral image blocks and LiDAR image blocks;
[0040] The feature extraction unit is used to use a multi-scale feature extraction module to extract features from the hyperspectral image block and the LiDAR image block respectively to obtain a hyperspectral image feature block and a LiDAR image feature block;
[0041] The fusion processing unit is used to use a CGAFT module to fuse the hyperspectral image feature block and the LiDAR image feature block respectively to obtain hyperspectral image features, LiDAR image features, a first cross-modal attention map, and a second cross-modal attention map;
[0042] The mask fusion unit is used to obtain a first mask fusion feature and a second mask fusion feature based on the hyperspectral image feature block and the LiDAR image feature block and the first cross-modal attention map and the second cross-modal attention map;
[0043] The loss calculation unit is used to calculate the contrast loss based on the hyperspectral image features and the LiDAR image features and the first mask fusion features and the second mask fusion features, and update the model parameters in combination with the cross entropy function;
[0044] The classification module is used to classify the multimodal data of the ground objects to be classified based on the model with updated parameters to obtain a classification result.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] By introducing a feature fusion module based on an attention-guided probabilistic masking strategy, the present invention can dynamically mask multimodal inputs (such as hyperspectral imagery and LiDAR data). This strategy utilizes the contrast between masked and original features to enhance the generalization and robustness of the model, enabling it to more accurately capture semantic information between modalities with a small number of labeled samples, thereby improving the model's performance and accuracy in complex tasks.
[0047] In addition, the cross-attention fusion module and hierarchical contrastive learning paradigm proposed in this paper effectively enhance the feature alignment between HSI and LiDAR data, ensuring the consistency of original data and mask data at different levels. This scheme optimizes the fusion effect of cross-modal data, which not only improves the overall performance of the model, but also enhances its adaptability and robustness in different environments and application scenarios. This method can effectively alleviate the problems caused by inconsistent information between modalities, thereby improving the accuracy and stability of cross-modal learning tasks. We conducted extensive experiments on four benchmark datasets to verify the effectiveness of this invention. The results show that the method proposed in this paper can effectively solve the classification problem of limited data and labeled samples, and outperforms other state-of-the-art methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 The overall framework flow chart of the method provided by the embodiment of the present invention;
[0050] Figure 2 This is a schematic diagram of a preprocessing method according to an embodiment of the present invention;
[0051] Figure 3 This is a schematic diagram of the structure of a multi-scale feature extraction module according to an embodiment of the present invention;
[0052] Figure 4 This is a diagram of the CGAFT module architecture according to an embodiment of the present invention;
[0053] Figure 5 This is a framework flow chart of classification based on a model with updated parameters according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] Example 1:
[0057] like Figure 1 As shown, this embodiment provides a multimodal collaborative object classification method based on mask contrast learning, including the following steps:
[0058] S1. Acquire hyperspectral data and LiDAR point cloud data, and preprocess the hyperspectral data and the LiDAR point cloud data to obtain hyperspectral image blocks and LiDAR image blocks.
[0059] A further implementation is that Figure 2 As shown in FIG, the preprocessing methods include: spatial alignment, format unification and resampling, enhancement and normalization processing, and slicing operation.
[0060] Specifically, the raw hyperspectral data acquired by a hyperspectral camera is typically represented as a regular three-dimensional data cube with the dimensional structure of "height × width × number of bands." The height and width correspond to the spatial dimensions of the imaging area, while the number of bands represents the reflectance information collected within different narrow wavelength ranges. Unlike conventional color images, which only record the brightness of the red, green, and blue channels, hyperspectral data records a complete, continuous spectral curve at each pixel, covering a rich spectral response from visible light (e.g., 400 nm) to the near-infrared and even extending into the short-wave infrared (e.g., 2500 nm). This fine-grained spectral information accurately characterizes the physical composition, chemical structure, and biological properties of objects, making hyperspectral data highly discriminative in tasks such as object classification and material identification. Furthermore, due to their regular spatial arrangement, hyperspectral images are also easily aligned and fused with other types of remote sensing data. Raw data collected by LiDAR systems is typically presented as a point cloud. Unlike hyperspectral imagery, it is an irregularly distributed three-dimensional data set. Each lidar point records its three-dimensional coordinates (X, Y, Z) in space, reflecting the geometric position of an object's surface. It also includes reflection intensity—the energy of the return signal—which provides additional clues about the surface material or roughness. Depending on the lidar system configuration and application requirements, each point may also include additional attributes such as the number of echoes, scan timestamp, echo sequence, and a calculated normal vector. Lidar point clouds can meticulously depict spatial information such as topographic features, building outlines, and vegetation structure. Although their spatial distribution is typically sparse, their three-dimensional perception capabilities are extremely strong. Therefore, lidar data plays an irreplaceable role in 3D modeling, scene reconstruction, object detection, and segmentation. It naturally complements hyperspectral data, known for its spectral detail, providing rich and powerful data support for multimodal collaborative classification. Tables 1 and 2 describe the raw data structures of hyperspectral cameras and lidar, respectively.
[0061] Table 1
[0062]
[0063] Table 2
[0064]
[0065] Therefore, before performing multimodal collaborative classification of hyperspectral data and lidar point cloud data, the two types of data need to be systematically preprocessed to eliminate the differences between the original data at the spatial, format and feature levels. The first is spatial alignment. Hyperspectral data is usually organized in the form of a two-dimensional regular grid. Each pixel corresponds to a fixed area on the surface and has a clear positional relationship in the geographic coordinate system. The lidar point cloud data is a set of irregularly distributed sparse three-dimensional point clouds, and each point has its own independent spatial coordinates. Therefore, before performing multimodal fusion, it is necessary to ensure that the two are strictly corresponding in spatial position, that is, the same ground object entity can match each other in the two data sources. This embodiment uses ground control points (GCPs, Ground Control Points) for alignment. Ground control points refer to a set of feature points that can be clearly identified and whose positions are known in both hyperspectral data and lidar point cloud data, such as road intersections, isolated buildings or obvious terrain features. By calibrating these control points in both datasets and calculating a spatial transformation model based on the control point coordinates, hyperspectral imagery or lidar point cloud data can be accurately reprojected into a unified geographic coordinate system (such as the UTM projection). The final registration effect is evaluated using the root mean square error (RMSE) to ensure that the spatial deviation of the control points is within an acceptable range, thus laying a solid foundation for subsequent fusion.
[0066] This is followed by format unification and resampling. After completing the spatial registration, the difference in data organization structure between hyperspectral data and lidar point cloud data still exists: hyperspectral data has a regular two-dimensional pixel arrangement, while lidar point cloud data is an irregular, sparse three-dimensional discrete point set. In order to achieve unified data structure, the lidar point cloud is projected onto the ground plane (XY plane) and rasterized. The specific operation is to divide the study area into regular grid cells (GridCells) and count the point cloud attribute features in each cell, such as average height, maximum height, point density, average reflection intensity, etc., to generate a regular two-dimensional feature map. This processing method can make the lidar data consistent with the hyperspectral image in structure, which is convenient for pixel-by-pixel feature fusion and classification analysis.
[0067] Third, at the feature level, these two types of data need to be enhanced and normalized. Gamma correction and atmospheric correction are used for hyperspectral data to improve the quality and accuracy of the data. Gamma correction is mainly used to correct the brightness distortion caused by the nonlinear response of the sensor and adjust the brightness value of the data to better match the perception of the human eye. It is achieved through the formula To adjust the image, it is suitable for low light environment or uneven brightness; where I corrected Represents the spectral value of a pixel in the output hyperspectral image; I originalRepresents the spectral value of a pixel of the input hyperspectral image, usually normalized to the range of [0,1]. γ represents the gamma coefficient, which is used to control the intensity of the correction. Atmospheric correction restores the ground reflectivity by removing the effects of atmospheric scattering and absorption, and performs correction through the 6S model to ensure that the data reflects the actual radiation conditions of the ground rather than the interference of atmospheric factors. In addition, for the processing of hyperspectral data and lidar point cloud data, normalization is an important step to ensure data quality and improve analysis accuracy. The present invention adopts minimum-maximum normalization for hyperspectral data and lidar point cloud data. By compressing the data to a specified range (usually [0, 1]), the dimensional differences between different bands are eliminated. Specifically, through the following formula:
[0068]
[0069] Where x′ represents all pixel values of a channel of the output image; x represents all pixel values of a channel of the input image; min(x) and max(x) represent the minimum and maximum values of all pixels in a channel of the input image, respectively.
[0070] Finally, after the aforementioned organization, the hyperspectral data and LiDAR point cloud data are sliced and labeled. Typically, these data are sliced into small patches of a fixed size (e.g., 11×11 pixels), with each patch labeled by the category of its center pixel. Based on the model input requirements, the processed hyperspectral and point cloud patches are packaged into a unified format (e.g., [Batch, Height, Width, Number of Channels]) for input into the neural network model training.
[0071] S2. Use a multi-scale feature extraction module to perform feature extraction on the hyperspectral image block and the LiDAR image block respectively to obtain a hyperspectral image feature block and a LiDAR image feature block.
[0072] Hyperspectral image blocks and LiDAR image patches Input into the multi-scale feature extraction module for feature extraction, where s×s represents the block size; c1 and c2 are the original channel numbers of hyperspectral data (HSI) and laser radar point cloud data (LiDAR), such as Figure 3As shown in the architecture diagram. For the hyperspectral data branch, a 3D convolution with a kernel size of 9×3×3 is used to convolve the hyperspectral image block, and then a set of two-dimensional convolutions with kernel sizes of 1×1, 3×3, and 5×5 are used to simultaneously extract features. Finally, the outputs after the two-dimensional convolution are fused using element-wise addition to obtain the hyperspectral image feature block. For the lidar point cloud data, since it is a single-channel image, the 3D convolution used for hyperspectral data is replaced by a two-dimensional convolution with a kernel size of 3×3 to extract elevation information from the LiDAR image block. Here, the number of filters of the two convolution layers in the hyperspectral data branch is 8 and 64, respectively, while in the lidar point cloud data branch, their number is 64. In addition, in order to speed up the convergence of the model, batch normalization (BN) and LeakyReLU functions are applied sequentially after each convolution layer. After this step, the blocks from the two modalities are mapped to the hyperspectral image feature blocks respectively. and LiDAR image feature blocks , where c = 64.
[0073] S3. Use the CGAFT module to fuse the hyperspectral image feature block and the LiDAR image feature block respectively to obtain hyperspectral image features, LiDAR image features, a first cross-modal attention map, and a second cross-modal attention map.
[0074] Subsequently, two double-branch CGAFTs were used to obtain the F H and F L The global context content is obtained from the model, and the complementary relationship between the modalities is further enhanced through the cross-guided attention fusion mechanism. These features are processed by a sophisticated cross-attention fusion module and the corresponding attention map is generated. Taking the HSI branch CGAFT as an example, specifically, F H and F L Flatten and concatenate with learnable classification labels to generate first sequence features and the second characteristic sequence Where N = s 2 +1. Then, the learnable position embedding is added to F′ H and F′ L To generate hyperspectral image features and LiDAR image features Next, Z H and Z LInput M dual-branch CGAFTs in sequence to mine global contextual knowledge and enhance feature fusion between the two modalities, while obtaining the attention map between the two modalities to enhance the original input. In this embodiment, N=2 is set. Since all CGAFTs have the same structure, this embodiment takes the CGAFT of the first HSI branch as an example. Figure 4 As well as Tables 3 and 4, Table 3 gives the specific algorithm process of hyperspectral features and LiDAR image features, and Table 4 gives the specific algorithm process of obtaining the first cross-modal attention map and the second cross-modal attention map.
[0075] Table 3
[0076]
[0077]
[0078] Table 4
[0079]
[0080] S4. Obtain a first mask fusion feature and a second mask fusion feature based on the hyperspectral image feature block, the LiDAR image feature block, the first cross-modal attention map, and the second cross-modal attention map.
[0081] The combined activation levels of the first and second cross-modal attention maps generated by the CGAFT module are obtained, and the input features are dynamically masked to enhance the robustness of the model, ensuring the validity of the information in each branch. Subsequently, hierarchical contrastive learning and cross-entropy are used to fully exploit the complementary information between the different modalities and the original and masked inputs. By exchanging semantics between HSI and LiDAR, as well as between the original and masked inputs, the complementary features of the two branches can be fully explored and fully integrated.
[0082] Taking the radar point cloud data branch as an example, the first and second cross-modal attention maps are generated to calculate the comprehensive activation strength of each feature location, which is used to measure its importance in cross-modal interaction. In the cross-modal attention map (CGAttention), the query vector (Query) of the LiDAR branch is dot-producted with the key vector (Key) of the HSI branch, which inherently implements the fusion and correlation modeling of features from different modalities in the spatial dimension. The attention score obtained from this dot product not only reflects the similarity between features but also implicitly captures the complementary and synergistic relationship between the two modalities. Therefore, the response value in the cross-modal attention map can be used as an indicator of the contribution of each feature location in the fusion process. Subsequently, a mask is dynamically generated based on the activation strength of each location, perturbing features in low-response areas through random zeroing or additive Gaussian noise. This approach allows the model to focus more on high-value areas in the fused features during training, while also improving its robustness to local information loss. In addition, in order to further adapt to the characteristics of different samples or different batches, the mask ratio can also be adjusted according to the statistical characteristics of the attention distribution in the current batch, thereby achieving more flexible and targeted feature enhancement.
[0083] The radar spectral feature attention weights (i.e., radar integrated activation strength) in the spatial dimension are calculated using the multi-head cross-guided self-attention backbone based on CGAFT. H and F L Apply CGAFT encoder g θ M times, and then output it to the FC network to obtain the feature representation of c channels Each MHCGA block consists of H heads, and for the HSI branch, it is defined as:
[0084]
[0085] Where i∈1,...,h represents an attention head. Self-attention is used to compute the query Q, key K and value V from the feature representation F, d k represents the dimension of k;
[0086] Calculate the average of the attention maps on the self-attention head to obtain the attention score A of the feature:
[0087]
[0088] Only the last CGAFT block is used to calculate the most active pixels, as it inherits information learned from previous blocks. The calculated attention score A can be used as an empirical semantic richness prior to guide the multimodal feature fusion masking strategy. Our goal is to mask the spectra with the largest activation intensity to reduce reliance on them, thereby enhancing model robustness and generalization ability. Therefore, the attention score is converted into a probability distribution that reflects the probability of each spectral feature being masked:
[0089] π=softmax(A / τ prop ),
[0090] Where, τ prop is a temperature hyperparameter that controls the sharpness of the output probability. A lower temperature (less than 1) will sharpen the distribution, making it more pointed and concentrated on the most active spectrum. Therefore, setting τ prop < 1, and use the Gumblemax trick to guide the masking strategy in a probabilistic manner:
[0091] maskinds=Top-K-indices(logπ+r),
[0092] K=S×S,
[0093] r=-log(-logε),ε∈U[0,1] S×S ,
[0094] In the formula, U[0,1] is a uniform distribution, Top means sorting the input from large to small and then selecting the largest K; r represents a Gumbel distribution random variable; indices represents the index of the feature. maskinds is the feature index to be masked. These features are replaced with learnable masking marks to obtain the first mask fusion feature F Hm and the second mask fusion feature F Lm . As shown in Table 5.
[0095] Table 5
[0096]
[0097] / / Step 4: Replace the masked spectral features with learnable masking markers
[0098] F_Hm=ApplyMask(F_H,maskinds) / / Replace features in F_H according to mask index
[0099] F_Lm=ApplyMask(F_L,maskinds) / / Replace features in F_L according to mask index
[0100] Return:F_Hm,F_Lm / / Returns the masked features
[0101] Similarly, the above-mentioned H and F L Get Z H and Z L A similar process based on F Hm and F Lm Get the hyperspectral mask feature Z Masked_H and LiDAR mask feature Z Masked_L .
[0102] S5. Calculate contrast loss based on the hyperspectral image features, the LiDAR image features, the first mask fusion features, and the second mask fusion features, and update model parameters in combination with a cross entropy function.
[0103] By comparing the mask fusion features with the original data fusion features, the model is forced to perform efficient feature learning under incomplete information conditions, thereby significantly improving its robustness in the presence of noisy, missing or partially occluded data.
[0104] The learning goal of our network is to learn the similarities between different modalities of the same sample and the differences between masked and unmasked samples within the same modality. Therefore, this embodiment constructs three comparison objectives: first, cross-modal comparison between HSI and LiDAR features to promote feature consistency at the same location across different modalities; second, comparison between original HSI and masked HSI to ensure semantic consistency of features before and after mask perturbation; and third, comparison between original LiDAR and masked LiDAR to improve the stability of the LiDAR branch. Specifically, there is a cross-modal comparison loss and two intra-modal comparison losses: HSI-LiDAR comparison loss, HSI-MaskedHSI comparison loss, and LiDAR-MaskedLiDAR comparison loss. All of these comparisons are performed in a high-dimensional feature space. The features are projected through a small MLP network and then optimized using the InfoNCE loss based on cosine similarity. This process promotes deeper semantic fusion between the modalities, which not only strengthens the sharing of complementary information between the two modalities but also further enhances the semantic representation capabilities of the features. Finally, by fusing original features and mask features, we effectively achieved full mining and integration of multimodal complementarity, significantly improving the classification performance and generalization ability of the model in complex scenarios.
[0105] The three contrast losses are as follows:
[0106]
[0107] Where, l HSI-liDARrepresents the HSI-LiDAR contrast loss; l HSI-MaskedHSI represents HSI-MaskedHSI contrast loss; l LiDAR-MaskedLiDAR represents the HSI-LiDAR contrast loss; τ is the temperature parameter; N b Indicates the number of images in a batch during training; represents the high-dimensional representation generated by the k-th sample in the corresponding batch, Represents the high-dimensional representation generated by the i-th sample in the corresponding batch, I k≠i Indicates that in the traversal of k, if the high-dimensional representation of the k-th sample and the high-dimensional representation of the current i-th sample are high-dimensional representations generated by different samples, it is 1, otherwise it is 0; Represents the LiDAR high-dimensional features generated by the k-th masked sample in the corresponding batch; Represents the LiDAR high-dimensional features generated by the i-th masked sample in the corresponding batch; Represents the LiDAR high-dimensional features generated by the i-th unmasked sample in the corresponding batch.
[0108] Finally, the model parameters are updated using both the cross entropy function and the multimodal contrast loss. In this embodiment, the model refers to the network module model used in the above steps, so no additional pre-training cost is required. The formula is as follows:
[0109] l Contrastive =l HSI-LiDAR +l HSI-MaskedHSI +l LiDAR-MaskedLiDAR ,
[0110] l total =l CE +α·l Contrastive ,
[0111] Where, l total represents the total loss function; l CE represents the cross entropy function; l Contrastive represents the total contrast loss; α represents the weight of the contrast loss.
[0112] S6, freeze the network parameters and remove the contrast and mask operations; input the image to be classified into the trained multi-scale feature extraction module and CGAFT module to extract its high-dimensional semantic features; then, map the features to the target category space through the fully connected layer, and combine the Softmax activation function to achieve the category prediction of each pixel, thereby completing the pixel-level classification task of the multimodal remote sensing image. Figure 5 shown.
[0113] Example 2:
[0114] The present embodiment provides a multimodal collaborative land feature classification system based on mask contrast learning, including: a pre-training module and a classification module; the pre-training module includes: a data processing unit, a feature extraction unit, a fusion processing unit, a mask fusion unit and a loss calculation unit; the data processing unit is used to obtain hyperspectral data and lidar point cloud data, and pre-process the hyperspectral data and the lidar point cloud data to obtain hyperspectral image blocks and LiDAR image blocks; the feature extraction unit is used to use a multi-scale feature extraction module to extract features from the hyperspectral image blocks and the LiDAR image blocks respectively, to obtain hyperspectral image feature blocks and LiDAR image feature blocks; the fusion processing unit is used to use a CGAFT module to extract features from the hyperspectral image blocks and the LiDAR image blocks respectively, to obtain hyperspectral image feature blocks and LiDAR image feature blocks; The feature block and the LiDAR image feature block are fused separately to obtain hyperspectral image features and LiDAR image features as well as a first cross-modal attention map and a second cross-modal attention map; the mask fusion unit is used to obtain a first mask fusion feature and a second mask fusion feature based on the hyperspectral image feature block and the LiDAR image feature block and the first cross-modal attention map and the second cross-modal attention map; the loss calculation unit is used to calculate the contrast loss based on the hyperspectral image features and the LiDAR image features as well as the first mask fusion feature and the second mask fusion feature, and update the model parameters in combination with the cross entropy function; the classification module is used to classify the multimodal data of the ground objects to be classified based on the model after parameter update to obtain a classification result.
[0115] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A multimodal collaborative feature classification method based on mask contrast learning, characterized by: The following steps are involved: S1. Acquire hyperspectral data and LiDAR point cloud data, and preprocess the hyperspectral data and the LiDAR point cloud data to obtain hyperspectral image blocks and LiDAR image blocks; S2. Using a multi-scale feature extraction module to perform feature extraction on the hyperspectral image block and the LiDAR image block, respectively, to obtain a hyperspectral image feature block and a LiDAR image feature block; S3, using the CGAFT module to fuse the hyperspectral image feature block and the LiDAR image feature block respectively to obtain hyperspectral image features, LiDAR image features, a first cross-modal attention map, and a second cross-modal attention map; S4. Obtaining a first mask fusion feature and a second mask fusion feature based on the hyperspectral image feature block, the LiDAR image feature block, the first cross-modal attention map, and the second cross-modal attention map; S5. Calculating contrast loss based on the hyperspectral image features, the LiDAR image features, the first mask fusion features, and the second mask fusion features, and updating model parameters in combination with a cross entropy function; the model includes: a multi-scale feature extraction module and a CGAFT module; S6. Based on the model with updated parameters, classify the multimodal data of the objects to be classified to obtain the classification results.
2. The multimodal collaborative feature classification method based on mask contrast learning according to claim 1 is characterized in that: The preprocessing method includes: spatial alignment, format unification and resampling, enhancement and normalization processing, and slicing operation.
3. The multimodal collaborative feature classification method based on mask contrast learning according to claim 1 is characterized in that: The method for obtaining the hyperspectral image feature block includes: A 3D convolution with a kernel size of 9×3×3 is used to perform a convolution operation on the hyperspectral image block. Then, a set of two-dimensional convolutions with kernel sizes of 1×1, 3×3, and 5×5 are used to simultaneously extract features. Finally, the outputs of the two-dimensional convolution are fused using element-wise addition to obtain the hyperspectral image feature block.
4. The multimodal collaborative feature classification method based on mask contrast learning according to claim 1, characterized in that: The method for obtaining the LiDAR image feature block includes: The LiDAR image block is convolved using a two-dimensional convolution with a kernel size of 3×3. Then, a set of two-dimensional convolutions with kernel sizes of 1×1, 3×3, and 5×5 are used to simultaneously extract features. Finally, the outputs after the two-dimensional convolution are fused using element-wise addition to obtain the LiDAR image feature block.
5. The multimodal collaborative feature classification method based on mask contrast learning according to claim 1 is characterized in that: The method for obtaining the hyperspectral image features and the LiDAR image features includes: Flattening the hyperspectral image feature block and the LiDAR image feature block and concatenating them with learnable classification labels to generate a first sequence feature and a second feature sequence; Adding a learnable position embedding to the first sequence feature and the second feature sequence to obtain the hyperspectral image feature and the LiDAR image feature.
6. The multimodal collaborative feature classification method based on mask contrast learning according to claim 1, characterized in that: The method for obtaining the first mask fusion feature and the second mask fusion feature includes: Calculating the comprehensive activation strength of each feature position based on the first cross-modal attention map and the second cross-modal attention map; Apply the encoder of the CGAFT module M times on the hyperspectral image feature block and the LiDAR image feature block respectively, and then input them into the FC network to obtain the feature representation F of c channels; Compute query Q, key K, and value V from feature representation F, and compute attention graph score based on query Q, key K, and value V; Converting the attention map score into a probability distribution, where the probability distribution is used to reflect the probability that the feature with high comprehensive activation intensity is masked; combining the Gumble max technique to guide the masking strategy in a probabilistic manner; The hyperspectral image feature block and the LiDAR image feature block are masked based on the masking strategy to obtain the first mask fusion feature and the second mask fusion feature.
7. The multimodal collaborative feature classification method based on mask contrast learning according to claim 1, characterized in that: The method for updating model parameters based on the contrast loss includes: l total =l CE +α·l Contrastive , in, l Contrastive =l HSI-LiDAR +l HSI-MaskedHSI +l LiDAR-MaskedLiDAR , Where, l total represents the total loss function; l CE represents the cross entropy function; α represents the weight of the contrast loss; l Contrastive represents the total contrast loss; l HSI-liDAR represents the HSI-LiDAR contrast loss; l HSI-MaskedHSI represents HSI-MaskedHSI contrast loss; l LiDAR-MaskedLiDAR represents the HSI-LiDAR contrast loss; τ is the temperature parameter; N b Indicates the number of images in a batch during training; represents the high-dimensional representation generated by the k-th sample in the corresponding batch, Represents the high-dimensional representation generated by the i-th sample in the corresponding batch, I k≠i Indicates that in the traversal of k, if the high-dimensional representation of the k-th sample and the high-dimensional representation of the current i-th sample are high-dimensional representations generated by different samples, it is 1, otherwise it is 0; Represents the LiDAR high-dimensional features generated by the k-th masked sample in the corresponding batch; Represents the LiDAR high-dimensional features generated by the i-th masked sample in the corresponding batch; Represents the LiDAR high-dimensional features generated by the i-th unmasked sample in the corresponding batch.
8. A multimodal collaborative feature classification system based on mask contrast learning, the system being used to implement the method according to any one of claims 1 to 7, characterized in that: include: Pre-training module and classification module; The pre-training module includes: a data processing unit, a feature extraction unit, a fusion processing unit, a mask fusion unit and a loss calculation unit; The data processing unit is used to obtain hyperspectral data and LiDAR point cloud data, and preprocess the hyperspectral data and the LiDAR point cloud data to obtain hyperspectral image blocks and LiDAR image blocks; The feature extraction unit is used to use a multi-scale feature extraction module to extract features from the hyperspectral image block and the LiDAR image block respectively to obtain a hyperspectral image feature block and a LiDAR image feature block; The fusion processing unit is used to use a CGAFT module to fuse the hyperspectral image feature block and the LiDAR image feature block respectively to obtain hyperspectral image features, LiDAR image features, a first cross-modal attention map, and a second cross-modal attention map; The mask fusion unit is used to obtain a first mask fusion feature and a second mask fusion feature based on the hyperspectral image feature block and the LiDAR image feature block and the first cross-modal attention map and the second cross-modal attention map; The loss calculation unit is used to calculate the contrast loss based on the hyperspectral image features and the LiDAR image features and the first mask fusion features and the second mask fusion features, and update the model parameters in combination with the cross entropy function; The classification module is used to classify the multimodal data of the objects to be classified based on the model with updated parameters to obtain a classification result.
Citation Information
Patent Citations
News scene multi-level image-text retrieval method based on mask guidance information fusion
CN119441515A
SAR image and optical image deep fusion terrain classification method based on self-supervised learning
CN119851134A
Systems and methods for cross-lingual cross-modal training for multimodal retrieval
US20220383048A1
Cited By
Sugarcane flowering phase identification method and system based on multi-modal fusion
CN121434726A