Multi-modal collaborative countermeasure attack method and device based on remote sensing image classification task

By designing a multimodal collaborative adversarial attack method, the perturbation intensity and directional consistency of the multimodal remote sensing image classification model are optimized, which improves the robustness and stability of the model in the face of adversarial attacks, solves the problem of insufficient model robustness in existing technologies, and provides a more efficient attack assessment and defense solution.

CN120747741APending Publication Date: 2025-10-03WUHAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510847871.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing multimodal remote sensing image classification models are not robust enough in the face of adversarial attacks and lack systematic evaluation and effective defense mechanisms, which may have serious consequences, especially in scenarios such as target monitoring and environmental assessment.

Method used

A multimodal collaborative adversarial attack method is designed. By constructing a multimodal fusion remote sensing image classification model, adopting single-modal and multimodal attack strategies, combining early, mid-term and late fusion structures, introducing a total loss function with perturbation intensity balance and collaborative direction constraints, the perturbation generation of each modality is optimized, and multimodal adversarial samples are formed to implement the attack.

Benefits of technology

It improves the deception success rate and stability of multimodal remote sensing classification models in the face of adversarial attacks, can deceive the model more covertly with smaller disturbances, enhances the robustness and security of the model, and provides an important reference for the security testing and defense mechanism of future multimodal remote sensing intelligent systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747741A_ABST
    Figure CN120747741A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal collaborative countermeasure attack method and equipment based on a remote sensing image classification task. The method comprises the following steps: carrying out data preprocessing and model construction, obtaining remote sensing image data comprising at least two different modals, carrying out preprocessing and spatial alignment, and constructing and training a multi-modal fusion remote sensing image classification model; collaborative disturbance generation, including constructing a disturbance generation process for each mode for the multi-mode fusion remote sensing image classification model; joint optimization, including optimization of disturbance generation of each mode by taking minimization of a total loss function as a target, the total loss function including disturbance intensity constraint, confrontation effect constraint, distribution similarity constraint and cooperation direction constraint; and attack implementation: adding each optimized modal disturbance to corresponding original modal data to form a multi-modal confrontation sample, and inputting the multi-modal confrontation sample to the target model to implement attack. According to the method, the robustness performance of the multi-modal classification model under different fusion strategies can be comprehensively evaluated, and an efficient collaborative attack scheme is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence security, and in particular relates to a method for generating multimodal adversarial attacks under a multimodal remote sensing image classification task. Background Art

[0002] With the continuous advancement of remote sensing imaging technology and deep learning, remote sensing image classification, a fundamental task in remote sensing intelligent analysis, is continuously developing towards refinement and high performance. This task is essentially a pixel-level classification problem, which assigns corresponding ground feature category labels by analyzing the spectral brightness characteristics and spatial structure of each pixel in the remote sensing image. Because it can provide critical geographic information support for a variety of fields, including land use mapping, disaster response, resource management, and ecological protection, it has become one of the most representative and valuable inventions in remote sensing applications.

[0003] With the rise of deep learning, models such as convolutional neural networks have been widely introduced in remote sensing image classification tasks, achieving significant performance breakthroughs on multiple public datasets. Their powerful feature extraction and representation capabilities enable the full modeling and utilization of the complex spectral and spatial information in remote sensing images. However, despite significant progress in classification accuracy, deep models face increasing robustness issues. Research has shown that by introducing tiny perturbations, imperceptible to the human eye, into remote sensing image inputs, the model can be tricked into making incorrect judgments, thereby generating adversarial examples. These examples are not only highly deceptive but can also have serious consequences in practical applications, particularly in scenarios such as target monitoring or environmental assessment.

[0004] Numerous patents have explored the vulnerability of remote sensing image classification models to adversarial attacks under varying network structures and input modalities. These studies generally find that models struggle to maintain stable outputs in the face of various attack vectors. Furthermore, some patents attempt to introduce adversarial defense methods, such as input preprocessing, perturbation suppression, and adversarial training, to enhance the models' resilience to attacks. While these approaches have achieved initial success, they still face challenges such as performance degradation and poor generalizability.

[0005] Multimodal fusion strategies are becoming a research hotspot for improving model accuracy and robustness. Compared to single-modality remote sensing, multimodal remote sensing imagery integrates information from different sensors, such as hyperspectral imagery, high-resolution multispectral imagery, thermal infrared remote sensing imagery, lidar point clouds, and synthetic aperture radar imagery. Each modality has its own strengths in perceiving spatial structure, material reflectance, or surface geometry. By fusing these complementary data, the model's discriminative capabilities can be enhanced and the error rate reduced, leading to a more reliable and accurate classification system.

[0006] Multimodal fusion strategies are mainly divided into three types: early fusion directly splices multi-source data at the input stage, mid-term fusion integrates the feature layers after feature extraction is completed in each modality, and late fusion retains the classification results of each modality and fuses the final decision through voting or weighting. These three strategies have their own characteristics and are suitable for different task backgrounds and data structures. They are widely used in multimodal data combinations such as HSI-MSI, HSI-LiDAR, and HSI-SAR. Although these fusion models have played a significant role in improving classification accuracy, there is still a lack of in-depth research on the robustness of the model, especially its behavior in the face of adversarial attacks. Most current inventions focus on fusion structure design and accuracy optimization, ignoring the systematic analysis of the model at the security level, and lack systematic evaluation of its robustness in the face of adversarial attacks. Summary of the Invention

[0007] In order to address the problem of lack of robustness of multimodal remote sensing image classification models in the prior art, the present invention provides a complete multimodal attack evaluation experimental framework and a multimodal collaborative counter-attack scheme.

[0008] The present invention provides a multimodal collaborative anti-attack method based on remote sensing image classification tasks, which performs the following steps: Data preprocessing and model building, including acquiring remote sensing image data containing at least two different modalities, performing preprocessing and spatial alignment; building and training a multimodal fusion remote sensing image classification model; Collaborative perturbation generation, including building perturbation generation processes for each modality in a multimodal fusion remote sensing image classification model; Joint optimization involves optimizing the perturbation generation for each modality with the goal of minimizing a total loss function that incorporates perturbation strength constraints to limit the magnitude of the perturbation for each modality, adversarial effect constraints to bias the model prediction away from the true label, distribution similarity constraints to balance the perturbation strengths of different modalities, and collaborative direction constraints to ensure spatial consistency of the perturbations across the model output. The attack implementation includes adding the optimized modal perturbations to the corresponding original modal data to form multimodal adversarial samples, which are then input into the target model to implement the attack.

[0009] Moreover, the remote sensing image data are hyperspectral images, lidar images and synthetic aperture radar images.

[0010] Moreover, the multimodal fusion remote sensing image classification model adopts early fusion, mid-term fusion and / or late fusion strategies.

[0011] Moreover, when constructing the perturbation generation process for each modality, a single-modal attack strategy and / or a multi-modal attack strategy are adopted. The single-modal attack strategy generates adversarial samples by independently perturbing a single modality data, and the multi-modal attack strategy generates adversarial samples by simultaneously perturbing multiple modal data.

[0012] Moreover, the total loss function is a weighted combination of four types of constraints, and the weight of each constraint item can be adjusted according to the requirements of the attack task.

[0013] Moreover, the distribution similarity constraint is implemented by calculating the similarity of the disturbance vector intensities of different modes to constrain the similarity of the disturbance intensity distribution.

[0014] Moreover, the collaborative direction constraint is achieved by respectively calculating the difference vectors of the target model output relative to the original unperturbed input when only one modal data is perturbed, then measuring the directional similarity between the difference vectors, and minimizing the similarity measure value to promote directional consistency.

[0015] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and runnable on the processor, wherein when the processor executes the program, it implements the multimodal collaborative counter-attack method based on remote sensing image classification tasks as described above.

[0016] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal collaborative counter-attack method based on remote sensing image classification tasks as described above.

[0017] On the other hand, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the multimodal collaborative counterattack method based on remote sensing image classification tasks as described above.

[0018] Compared with the prior art, the present invention has the following advantages: Addressing a gap in the existing technology, this paper systematically investigates the performance of multimodal fusion remote sensing image classification models in the face of adversarial attacks. It designs two types of perturbation strategies, including single-modal and multimodal attacks, and combines three fusion structures: early, mid-stage, and late-stage, to construct a complete attack evaluation experimental framework. Furthermore, a multimodal collaborative adversarial attack method is designed. This method constrains the similarity of the perturbation strengths of different modalities and the similarity of the impact of each modal perturbation on the original classification results, mobilizing coordinated attacks between modalities. Compared to existing attack methods (such as FGSM and PGD, which are simply applied to multiple modalities), this method can more effectively deceive the target model (i.e., have a higher attack success rate), achieve the deception effect with smaller perturbations (more subtle), and is more stable and reliable (less likely to fail due to inter-modal perturbation conflicts). This invention expands the robustness analysis dimension of multimodal remote sensing classification systems, and also provides an important reference for the security testing and defense mechanisms of future remote sensing intelligent systems. It is conducive to revealing the possible and previously undiscovered serious security flaws of multimodal fusion models (even those with high accuracy) when facing carefully designed coordinated attacks. It sets a higher benchmark for evaluating the robustness of future multimodal models and provides key prerequisites and direct input for designing effective defense mechanisms (such as adversarial training and robust fusion strategies). BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of multimodal data preprocessing and construction according to an embodiment of the present invention.

[0020] Figure 2 It is a schematic diagram of the multimodal fusion structure of an embodiment of the present invention.

[0021] Figure 3 It is a structural diagram of the multimodal attack evaluation experiment framework according to an embodiment of the present invention.

[0022] Figure 4 Schematic diagram of a multi-modal collaborative anti-attack method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0024] This paper addresses the practical needs of multimodal remote sensing image classification systems for security analysis and robustness assessment. Addressing the issues of existing methods, such as the lack of systematic robustness analysis of multimodal models and the lack of coordination in attack mechanisms, a comprehensive multimodal attack assessment experimental framework and a multimodal collaborative counterattack method are proposed. By constructing an attack simulation mechanism for multimodal input scenarios and introducing modal balance and collaborative perturbation strategies, a more targeted and practical attack assessment of remote sensing image classification models is achieved. First, a unified preprocessing process is constructed to standardize and structurally align heterogeneous modal data from hyperspectral imagery, lidar point clouds, and synthetic aperture radar, ensuring the synchronization and consistency of modal inputs. Second, three typical multimodal fusion strategies (early, mid-term, and late) are designed and a unified fusion classification network architecture is constructed to achieve effective joint modeling and semantic discrimination of multimodal information. During the adversarial attack phase, single-modal and multimodal perturbation generation mechanisms are constructed, extending the classical attack algorithm to accommodate multimodal scenarios. By introducing a perturbation intensity balance loss and a collaborative adversarial loss, the perturbations of each modality are coordinated in amplitude and direction, enhancing the coordination of the attack and the perturbation effect on the model output. This method comprehensively evaluates the robustness of multimodal classification models under different fusion strategies and provides efficient collaborative attack solutions, providing important support for security testing and defense design of multimodal remote sensing intelligent systems.

[0025] The embodiment of the present invention provides a multi-modal collaborative anti-attack method based on remote sensing image classification tasks, which is implemented in a remote sensing data processing platform (such as PyTorch Geometric + OpenMMLab). Taking the HSI+LiDAR+SAR tri-modal classification model as an example, see Figure 1 , the complete implementation process is as follows: Step 1: Multimodal data preprocessing and construction: Collect remote sensing image data of at least two different modalities, which can be remote sensing data from different sensors such as hyperspectral HSI images, LiDAR point clouds, SAR images, etc., and uniformly preprocess and standardize the data of each modality.

[0026] The present invention further proposes that the data preprocessing step includes operations such as denoising, spectral information normalization, and patch extraction. Each data modality needs to be processed independently to ensure data accuracy and consistency.

[0027] In this step, data from multiple remote sensing sensors can first be normalized and feature aligned. The goal is to unify data from heterogeneous sources and diverse representations into a consistent input format, facilitating subsequent integration and attack in the fusion model. This step primarily processes hyperspectral imagery, lidar imagery, and synthetic aperture radar imagery, corresponding to raw observation data acquired by hyperspectral imagers (HSI), lidar systems (LiDAR), and SAR satellite sensors, respectively. During the specific operation, spectral normalization, patch construction and slicing, and modal consistency verification are performed on each modality. Spectral normalization is performed on all images. HSI typically uses min-max normalization and z-score normalization for each band. LiDAR elevation data is normalized to the [0, 1] range using a threshold interval. SAR images undergo logarithmic compression or amplitude normalization to enhance contrast. In terms of patch construction and tiling, large-scale remote sensing images are tiled using a sliding window approach. Each tile retains its positional information within the original image and is labeled with the corresponding feature category for supervised learning tasks. For modal consistency verification, standardized patches from each modality are matched and combined according to their spatial position to form a unified multimodal sample input format.

[0028] The implementation process of the embodiment is described as follows: A dataset consisting of three primary remote sensing data sources is used: hyperspectral imagery (HSI), laser radar (LiDAR), and synthetic aperture radar (SAR). Hyperspectral imagery consists of 144 continuous bands, covering a wavelength range of 364 to 1046 nanometers, with a spatial resolution of 349 rows × 1905 columns. LiDAR elevation maps are single-band grayscale images with the same spatial resolution as HSI. SAR images contain single bands and also maintain the same spatial resolution as HSI, facilitating subsequent spatial registration and joint modeling. During data preprocessing, all bands of the hyperspectral imagery are normalized to a 0-1 range to unify the input value range and ensure network training stability. LiDAR images are median filtered to remove elevation outliers, while SAR images are logarithmically transformed to enhance detail. Linear mapping is then used to compress the range to [0, 1], achieving amplitude consistency with the HSI data. In terms of image segmentation, sample blocks are constructed using a sliding window approach, with each patch size of 11×11 pixels and a stride of 1. Each sample consists of a central pixel and its surrounding neighborhood, and carries its spatial location information in the original image and the object category label, which serves as the input for supervised training.

[0029] When constructing early fusion multimodal input samples, patches from HSI, LiDAR, and SAR are aligned and channel-wise stitched based on spatial coordinates to generate a structurally consistent and semantically unified multimodal sample tensor. Invalid patches (such as out-of-bounds edges and unlabeled regions) are automatically filtered to ensure data quality in the training set. The final sample format is [C, 11, 11], where C is the total number of channels. All samples are divided into training and test sets in an 8:2 ratio to ensure consistent distribution of object categories and regional independence, thereby enhancing the classifier's generalization ability across different scenarios.

[0030] Step 2: Multimodal fusion structure construction and training: Based on the preprocessed data, a multimodal fusion method is used, including early fusion, mid-term fusion and / or late fusion models, to merge data from different modalities and train different fusion models.

[0031] like Figure 2 As shown in the figure, after completing the standardized preprocessing of multimodal remote sensing image data, the model structure construction phase begins, aiming to build a unified classification network architecture that can adapt to the feature distribution of different modal data and has high flexibility. The embodiment is based on a general multimodal remote sensing image classification baseline network design, which is oriented towards multi-source inputs from hyperspectral imagers (HSI), laser radar systems (LiDAR), and synthetic aperture radar systems (SAR), and realizes the effective fusion of different modal information within the neural network. This fusion structure relies on the collaborative construction of two core modules, namely the feature extraction module (Extract-Net) and the feature fusion module (Fusion-Net). The two modules assume the core responsibilities of information extraction, modality alignment, and fusion discrimination in the overall structure.

[0032] Extract-Net is used to extract intra-modal features from HSI, LiDAR, and SAR images, respectively. Its network structure consists of multiple sequentially stacked convolutional modules, performing convolution calculations, batch normalization, and activation functions at each stage to extract modality-specific local spatial semantics. Fusion-Net further combines high-level features from multiple modal outputs for joint modeling and classification prediction. Its internal structure typically consists of multiple 1×1 convolutional modules, combined with a pooling mechanism and a softmax layer, ultimately outputting predictions for multiple types of terrain features. The connection between these two modules can be flexibly controlled to meet the needs of different fusion strategies, supporting the modeling and implementation of three typical fusion mechanisms: early fusion, mid-term fusion, and late fusion.

[0033] The specific implementation process of the embodiment is described as follows: Extract-Net is used to extract intra-modal features from hyperspectral image (HSI), laser radar (LiDAR) and synthetic aperture radar (SAR) data respectively. The overall structure consists of four sequentially connected modules. Module one consists of a 3×3 convolution layer, batch normalization (BN), and ReLU activation function set in sequence, extracting spatial local features and outputting 16 channels, for example, outputting 7×7×16 features; module two consists of a 1×1 convolution layer, BN, ReLU activation function, and 2×2 maximum pooling operation set in sequence, further extracting and compressing features, keeping the number of channels at 16 and reducing the spatial dimension, for example, outputting 4×4×16 features; module three again uses 3×3 convolution, BN, and ReLU set in sequence to refine high-order semantic features and output 16 channels, for example, outputting 7×7×16 features; module four uses 1×1 convolution, BN, ReLU, and 2×2 maximum pooling set in sequence, and the number of output channels is expanded to 32 to provide richer semantic information for the subsequent fusion stage, for example, outputting 4×4×32 features.

[0034] Fusion-Net implements the fusion of modal features output by Extract-Net and the final category prediction. It specifically includes three key modules: Module 5 first performs preliminary feature fusion and dimensionality adjustment through 1×1 convolution, BN and ReLU activation functions, and outputs 128 feature channels, such as outputting 2×2×128 features; Module 6 further uses 1×1 convolution, BN and 2×2 average pooling to reduce the dimension to 64 channels for more compact feature expression, such as outputting 1×1×64 features; Module 7 uses 1×1 convolution and Softmax layer to map the features to the final category probability distribution, realizing the prediction of the ground object category, such as outputting 1×1×C results, where C is the total number of categories.

[0035] In the early fusion architecture, processed HSI, LiDAR, and SAR images are directly spliced ​​in the channel dimension at the input stage to construct a fused feature tensor, which is then passed as a unified input stream to Extract-Net for processing. Taking the dataset used in this invention as an example, HSI images have 144 spectral channels, LiDAR images have 1-channel elevation data, and SAR images have 1-channel data. After splicing, a four-dimensional tensor with dimensions of [batch size, 146, height, width] is formed. This tensor is directly used as input in the convolutional network to achieve joint feature extraction of the original modal information. Because splicing occurs at the very beginning of the model, the subsequent network structure does not need to distinguish between modal sources and can extract semantic representations uniformly. However, this also requires high spatial alignment accuracy between modalities and places stricter requirements on input amplitude consistency.

[0036] The mid-term fusion architecture is constructed using a dual-stream network structure, constructing three parallel Extract-Net networks to process input data from the HSI, LiDAR, and SAR modalities, respectively. Each branch network independently extracts modal features. After completing multiple convolution stages, the three output features are fused in the Fusion-Net. Fusion can be performed using channel concatenation or, depending on the experimental configuration, weighted fusion. This enables interaction between modal features in the mid-level representation space. This mid-term fusion strategy strikes a good balance between modality-independent expression and joint discrimination, offering strong scalability and suitability for scenarios with significantly different modal data features.

[0037] In the late fusion architecture, complete Extract-Net and Fusion-Net subnetworks are established for each modality. After the inputs of the three modalities are calculated through their respective full processes (in sequence through the Extract-Net and Fusion-Net subnetworks), each independently outputs a softmax probability distribution. Finally, the three probability vectors are combined at the output level, and the final classification decision is made through weighted averaging or maximum voting strategies. Late fusion achieves a high degree of module decoupling in its structure, making it suitable for environments with temporal asynchrony in modal data, inconsistent acquisition frequencies, or noise loss in a particular modality, thereby improving the robustness and flexibility of the overall system.

[0038] After the network structure is built, the model training phase relies on high-quality samples and optimization strategies to further improve the expressive power and classification performance of multimodal feature fusion. Before training, spatial alignment and amplitude normalization between modalities are uniformly performed to construct a high-quality data set. Cross entropy is used as the loss function during training. The Adam optimizer is used in the training process, the initial learning rate is set to 1e-3, and the total number of rounds is set to 150 in combination with preheating and adaptive learning rate decay strategies. In addition, the modal Dropout strategy is introduced in training to randomly block a certain modal input in some iterations to enhance the robustness of representation. Through the coordinated optimization of structural design and training strategies, the present invention demonstrates good generalization performance and classification accuracy in complex multi-source remote sensing scenarios.

[0039] Step 3: Design an Adversarial Perturbation Generation Strategy: For multimodal remote sensing image classification models, select appropriate adversarial attack strategies, including but not limited to unimodal and multimodal attacks. Unimodal attacks generate adversarial perturbations for each modality separately, while multimodal attacks perturb multiple modalities simultaneously. Adversarial examples are generated using adversarial attack algorithms (such as FGSM, PGD, MIM, and C&W) to test the model's attack resistance. Generated adversarial examples are input into a trained multimodal remote sensing image classification model to conduct the attack. The robustness and stability of the models are analyzed by comparing the performance of the models using different modal fusion strategies with that of the unimodal models. After each attack, evaluation metrics such as the change in classification accuracy, robustness index, and perturbation size are calculated and recorded.

[0040] See also Figure 3 After completing the training of the multimodal classification model and deploying the fusion structure, the present invention proposes to further construct a multimodal adversarial perturbation generation strategy that is adaptable to remote sensing scenarios. This strategy supports both unimodal and multimodal attacks, aiming to systematically evaluate the robustness of the model under different attack modes. The perturbation generation process relies on gradient-based and optimization-based adversarial attack algorithms, which precisely interfere with the model's discrimination boundary by controlling the perturbation intensity, direction, and number of active modes. In terms of attack strategy construction, the present invention preferably recommends introducing four classic adversarial attack methods: FGSM, PGD, MIM, and C&W, and adapts them for unimodal and multimodal scenarios respectively. For unimodal attacks, HSI, LiDAR, or SAR images can be perturbed independently, while the other modal input remains unchanged when generating adversarial samples. For multimodal attacks, perturbations are generated simultaneously for the HSI, LiDAR, and SAR channels, and the perturbation intensity and direction are controlled to be consistent to improve the overall attack efficiency.

[0041] The specific experimental process of the embodiment is described as follows: This paper categorizes adversarial attacks into unimodal and multimodal attacks based on the number of perturbation modalities. This example discusses a model with HSI, LiDAR, and SAR input modalities. Therefore, a unimodal attack against this model involves perturbing HSI, LiDAR, or SAR data individually to attack the model. A multimodal attack involves simultaneously and independently perturbing the HSI, LiDAR, and SAR input modalities to deceive the model.

[0042] In the single-modal attack phase, the HSI modality is first selected as the perturbation target. Taking FGSM as an example, the gradient direction is directly signed, and the original image is scrambled according to the perturbation step size to construct an adversarial sample, which is then stored in the adversarial sample set. If the PGD or MIM method is used, multiple rounds of gradient updates are performed on the sample according to the set perturbation strength and number of iterations. After each update, the perturbation is clipped to within the maximum threshold range, and a clip operation is performed in each round to ensure the legitimacy of the adversarial sample. In LiDAR and SAR channel attacks, the process is basically the same, except that the perturbation target is transferred from the spectral vector to a single-channel elevation map, and the value of each elevation pixel is modified proportionally according to the gradient direction.

[0043] For single-modal attacks on multimodal models, the present invention primarily adopts the following state-of-the-art counter-attack methods. For ease of implementation and reference, a specific implementation is introduced below: (1) Fast Gradient Signed Method (FGSM): FGSM was proposed by Goodfellow et al. It is a one-step method for generating adversarial samples. The formula is as follows:

[0044] in, is a legitimate image, It's the real label. is a sign function, is the maximum disturbance value, which controls the maximum intensity of the anti-disturbance, Represents the gradient of the objective function with respect to the input legitimate image. By searching for adversarial samples within the domain sphere of the legitimate image, the generated adversarial samples can deceive the deep learning model with small perturbations. FGSM only needs to calculate the gradient once and has a faster stacking speed.

[0045] (2) Projected Gradient Descent (PGD): PGD is an iterative version of one step. Adversarial examples generated by PGD As shown in the following formula:

[0046] in, represents the adversarial sample in the i-th iteration process, y represents the true label, Represents a sample The gradient of the loss function with respect to the input for the true label y is, is the iteration step length, is the maximum disturbance value, It is a clipping operation in each iteration, which is used to constrain the image to be within the neighborhood ball of the legal image.

[0047] (3) Momentum-based Iterative Algorithm (MIM): MIM incorporates momentum terms into the iterative process to further stabilize the update direction and alleviate local minima. This is shown in the following formula:

[0048] Where i represents the number of iterations, represents the momentum term of the i-th iteration, represents the momentum term of the i+1th iteration, α is the momentum coefficient, which is used to control the contribution of the historical gradient; is the adversarial sample of the i-th iteration, is the adversarial sample of the i+1th iteration, () is the gradient of the loss function L with respect to the input, is the iteration step size, is the maximum perturbation value, sgn() is the sign function, is the gradient vector norm. The momentum term is the weighted average of the previous gradients. By adding the momentum term, MIM can effectively solve the instability of the gradient descent process.

[0049] (4) Carlini and Wagner (C&W): The C&W attack algorithm was proposed by Carlini and Wagner and is based on optimizing the perturbation solution, as shown in the following formula:

[0050] in, Indicates the goal of minimizing the perturbation that we want to achieve during the optimization process, even if the original image and the perturbed image As similar as possible. The algorithm proposes 6 candidates for the function f(). In the present invention, the candidate function is preferably , The model input The score of the output non-target category i, Is the model input The output score of the target category t, The difference between the maximum score of the non-target class and the score of the target class is used to guide adversarial sample generation. c is a constant that serves as a trade-off between the perturbation size and the attack strength, controlling the importance of both parts in the optimization objective.

[0051] During the multimodal attack phase, gradient channels are opened for the HSI, LiDAR, and SAR modalities, and the partial derivatives of the loss function of the three modal inputs are calculated. These are then normalized and combined into an overall perturbation vector. During each iteration, a unified random perturbation initialization strategy is adopted, adding a small amount of random value to each modal input to increase sample diversity. Taking MIM as an example, a historical gradient accumulation is maintained for each of the three modalities. In each iteration, the current gradient direction is merged and updated with the momentum vector of the previous round to guide the perturbation direction to evolve in a stable direction. This operation enhances the consistency of the gradient perturbation path and improves the convergence speed of the attack. Throughout the optimization process, the three modal perturbation values ​​are individually L∞ clipped to prevent the interference amplitude of any modality from being too large and thus destroying the perceptual consistency.

[0052] To combat multimodal attacks on multimodal models, this paper further proposes the use of four state-of-the-art adversarial attack methods against single-modal attacks: improved FGSM, PGD, MIM, and C&W. These methods are extended to multimodal scenarios. The implementation details are as follows, where H represents the HSI modality, L represents the LiDAR modality, and S represents the SAR modality: (1) Multimodal FGSM attack: This method simultaneously perturbs the HSI, LiDAR, and SAR modal inputs. The perturbations are generated by a one-step reverse gradient perturbation:

[0053] in, represents the original input data of the HSI mode, represents the raw input data of the LiDAR modality, represents the raw input data of the SAR modality, 、 、 They represent the adversarial samples corresponding to each modality. sgn(•) is the symbolic function, is the maximum perturbation value, y is the true label, 、 、 Represents the gradient of the objective function corresponding to each modality with respect to the input legal image.

[0054] (2) Multimodal PGD attack: This method performs iterative gradient attacks on the three modalities of HSI, LiDAR, and SAR:

[0055] in, represents the adversarial sample of the i-th iteration of the HSI modality, represents the adversarial sample of the i-th iteration of the LiDAR modality, represents the adversarial sample of the SAR modality iteration i, 、 、 represents the adversarial sample of the i+1th iteration corresponding to each modality, is the iteration step length, is the maximum disturbance value, 、 、 is the corresponding clipping operation of each mode in each iteration process, 、 、 Represents the gradient of the objective function corresponding to each modality of the i-th iteration with respect to the input legal image.

[0056] (3) Multimodal MIM attack: This method introduces the momentum term into the gradient update process and uses a gradient-based iterative attack to perturb the three modes of HSI, LiDAR, and SAR respectively:

[0057] in, represents the momentum gradient term of the i-th iteration of the HSI mode, represents the momentum gradient term of the i-th iteration of the LiDAR mode, represents the momentum gradient term of the SAR mode iteration i, 、 、 represents the momentum gradient term of the i+1th iteration corresponding to each mode. is the iteration step size, and α is the momentum coefficient.

[0058] (4) Multimodal C&W attack: This method is based on optimizing the perturbation solution, while considering optimizing the perturbation size added to the three modalities and the adversarial loss of the multimodal network:

[0059] in, represents the original input data of the HSI mode, represents the raw input data of the LiDAR modality, represents the raw input data of the SAR modality, 、 、 Represent the adversarial samples corresponding to each modality.

[0060] On the basis of the implementation of the above-mentioned single-modal and multi-modal attack strategies, the present invention further observed that there are still two key limiting factors in the multi-modal perturbation process, namely, disturbance imbalance and disturbance conflict. These problems are particularly prominent in multi-modal data fusion, which restricts the attack effect. The former refers to the different sensitivities of different modes to disturbances, resulting in insufficient disturbance effects of some modes or waste of disturbance resources; the latter is reflected in the inconsistent gradient directions of attack effects between multiple modalities, which may offset each other, thereby weakening the overall attack effect. Therefore, the present invention subsequently proposes a multi-modal collaborative counter-attack scheme to achieve coordinated optimization of the disturbance direction and intensity of each modality, and further improve the attack efficiency and stability.

[0061] Step 4: Multimodal collaborative perturbation optimization design: This paper proposes a multimodal collaborative counterattack scheme to enhance the attack effect of the multimodal attack method. The design ideas are as follows: 1) For multimodal remote sensing image input, we construct independent adversarial perturbation generation processes for each modality, such as hyperspectral imagery (HSI), LiDAR point clouds, and SAR images. In this phase, we continuously update the perturbation information for each modality through optimization. The goal of optimization is to reduce perturbations, maximize the attack effectiveness, and shift the model classification away from the true label.

[0062] In order to achieve effective counter-attack, the present invention selects an optimization-based method and specially designs a loss function system for attack.

[0063] Among them, a key loss is disturbance loss It is used to limit the perturbation intensity and ensure that the modification of each type of image does not exceed a reasonable range, thereby preventing excessive noise from affecting image readability or data authenticity.

[0064] Another loss against loss It is used to improve the effectiveness of the attack and encourage the classification model to output incorrect prediction results after multimodal fusion. By maximizing the gap between the model prediction results and the true labels, the deceptive ability of adversarial samples is further strengthened.

[0065] 2) Increase the perturbation intensity balance loss during the optimization process, adjust the amplitude of the perturbation, and reduce the difference in perturbation intensity between each modality, so that the perturbation of each modality can appropriately affect the decision-making of the model, rather than relying solely on the data of a certain modality. Specifically, by adjusting the differences between the perturbations of each modality, the perturbation amplitudes on the hyperspectral, lidar, and synthetic aperture radar images are made closer. This step maximizes the effect of the perturbation under multimodal data fusion by iteratively optimizing the perturbation. Loss To balance the perturbation loss, the goal is to balance the perturbation size generated by each mode, G , G and G is the perturbation generated by each modality. The present invention preferably recommends using cosine similarity to measure the similarity of perturbations of different modalities. This loss can be used to constrain perturbations between different modalities to have similar distributions, and to find the optimal perturbation combination under this constraint.

[0066] 3) In the optimization process, modal collaborative confrontation loss is added, and the difference measurement item of modal classification confidence distribution is added to guide the multimodal perturbation to maintain the same perturbation direction on the model classification results, ensuring that the perturbations generated by different modalities have a synergistic effect in the discriminant space, and avoiding perturbation conflicts and mutual cancellation to the greatest extent. The collaborative adversarial loss aims to ensure that the perturbations of each modality have similar perturbation directions on the results of the multimodal classification network, thereby avoiding conflicting perturbation effects from multiple perturbations. This loss helps the algorithm find the optimal perturbation combination with minimal perturbation and a more pronounced adversarial effect.

[0067] like Figure 4 As shown in the figure, after completing the initial generation of adversarial examples, to further improve the coordination and stability of attacks in multimodal remote sensing classification tasks, the embodiment introduces a collaborative perturbation optimization mechanism to address issues such as uneven perturbation amplitudes across multiple modalities and inconsistent output deviation directions. This mechanism focuses on balancing perturbation intensity and aligning directions within the semantic discriminant space. By constructing a joint optimization strategy for multiple loss functions, each round of perturbation updates not only focuses on classification misleading itself, but also achieves overall coordination of cross-modal perturbation paths.

[0068] The optimization process for the multimodal collaborative adversarial attack is based on a trained multimodal fusion classification model. Without changing the model's structure or parameters, trainable perturbation variables are introduced for each of the three modalities: HSI, LiDAR, and SAR. In each iteration, the original multimodal input is first added to the current perturbation variable to generate a collaborative adversarial example, which is then fed into the fusion model to obtain the fused classification prediction result.

[0069] Next, a joint loss function is calculated that includes multiple constraints, including the adversarial classification loss that guides the model to output incorrect predictions. , control the disturbance amplitude not to exceed the threshold disturbance intensity loss , constraining the disturbance balance loss of each modal disturbance amplitude to be close , and the collaborative perturbation loss that maintains consistency in the output space in the adversarial direction These losses are combined in a weighted fashion to form the overall optimization objective.

[0070] The perturbation variables for each modality are then updated through backpropagation, and perturbation clipping is performed after each round to ensure legitimacy and perceptual acceptability. The optimization process continues iterating until stopping conditions, such as a maximum number of rounds, a change in classification output, or convergence of the total loss, are met. The resulting multimodal collaborative adversarial examples effectively disrupt model classification without disrupting image semantics or perceptual consistency, significantly improving adversarial attack effectiveness and cross-modal collaboration capabilities.

[0071] The specific implementation process of the embodiment is described as follows: The core idea of ​​this method is to formalize the generation process of adversarial perturbations as a multi-objective optimization problem, comprehensively considering the four major goals of perturbation intensity control, attack effectiveness, modal consistency, and interference direction coordination. The optimization objective function is defined as follows:

[0072] in, Control the disturbance amplitude, Improve attack effectiveness, Ensure the balance of modal disturbance intensity, Strengthen the consistency of disturbance direction. 、 and are the weight coefficients of each loss item, which can be flexibly adjusted according to the requirements of the attack mission.

[0073] The goal of the loss term is to generate perturbations that are as imperceptible as possible to maintain the visual naturalness of the adversarial examples and reduce the risk of human recognition. In remote sensing imagery, significant changes in the image's spectral response or terrain texture can cause erroneous responses in practical application systems, making the controllability of the perturbation particularly critical.

[0074]

[0075] in, and They represent the original hyperspectral image HSI input and the hyperspectral image HSI input after adding disturbance respectively; and They represent the original LiDAR input and the LiDAR input after adding disturbance respectively; and Respectively represent the original synthetic aperture radar SAR input and the synthetic aperture radar SAR input after adding disturbance; the present invention selects The loss ensures that the attack is visually deceptive and prevents the perturbation from excessively destroying the structural characteristics of the original data.

[0076] This is the core of the attack optimization, and its goal is to disrupt the model's classification behavior to the greatest extent possible. It optimizes the model's output for adversarial samples to deviate from the true label, thereby achieving misclassification.

[0077]

[0078] in, Represents the hyperspectral image HSI input after adding disturbance; Represents the LiDAR input after adding disturbance; represents the synthetic aperture radar SAR input after adding disturbance; f Represents a multimodal remote sensing classification model; represents the true label of the sample; Represents the classification loss function, here it is the cross entropy loss.

[0079] In the balanced perturbation scheme, the disturbance magnitude generated by each mode is first calculated. The perturbation intensity of different modes is measured by the norm. The present invention defines the perturbation generated on the hyperspectral image HSI as , the modal disturbance generated on the radar point cloud LiDAR is , the modal disturbance generated on the synthetic aperture radar SAR is In this invention, the function of obtaining the disturbance vector is named G , and the perturbation vector is calculated as:

[0080] in is the disturbance vector of the HSI mode, is the original input data of HSI, is a parameter in the model optimization process. Similarly, is the perturbation vector of the LiDAR modality, is the raw input of LiDAR data, is the disturbance vector of the SAR mode, is the raw input of SAR data.

[0081] Since the dimensions of the perturbation vectors of HSI, LiDAR and SAR are different, this paper introduces a conversion function T to map the perturbation vectors of LiDAR and SAR images to the same dimensions as the HSI hyperspectral image, so as to facilitate subsequent calculation and optimization. The radar point cloud data LiDAR and synthetic aperture radar data SAR are transformed by the conversion function T into and The formula is as follows:

[0082] in, Represents the dimension of HSI hyperspectral image, represents the dimension of the LiDAR image. The transformation function T() expands the LiDAR image to the same dimension as the HSI image. Since LiDAR images have only one channel, while hyperspectral images have multiple channels depending on their spectral range, the present invention employs a channel stacking approach, whereby the single LiDAR image channel is continuously superimposed to form a perturbation with the same number of HSI channels. This expansion and transformation allows the perturbation vectors of HSI and LiDAR to be compared and analyzed for similarity within the same dimension, facilitating further analysis and processing. This is crucial for achieving modal perturbation balance. The same principle applies to the dimensionality of SAR images.

[0083] The purpose of balanced perturbation loss is to make the perturbations of the HSI hyperspectral modality, LiDAR radar point cloud, and SAR synthetic aperture radar modality consistent in magnitude, resulting in a more balanced perturbation. By minimizing the difference in perturbation strength between the different modalities, the perturbations of the three modalities are balanced, further improving the attack effectiveness. In general, the loss function formula for the balanced perturbation of the present invention is as follows:

[0084] in, Represents a vector Norm, used to measure the strength of the disturbance.

[0085] To implement the collaborative countermeasure solution, the present invention defines three disturbance effect change vectors, which are used to represent the attack effects of HSI, LiDAR and SAR respectively. Assume that the output of the multimodal model is ,in 、 and are the original HSI image, the original LiDAR image and the original SAR image respectively. , without disturbance and In the case of multimodal model output is , a single perturbation , without disturbance and In this case, the multimodal model output is , a single perturbation , no disturbance and In this case, the multimodal model output is The present invention defines the change vector of the perturbation attack effect as the difference vector between the output of the model after perturbation and the output of the model under the original data:

[0086] in, 、 and They respectively represent the change vector of the disturbance attack effect on HSI disturbance, the change vector of the disturbance attack effect on LiDAR disturbance, and the change vector of the disturbance attack effect on SAR disturbance.

[0087] The present invention uses the cosine similarity function to constrain three disturbance attack effect vectors, namely the disturbance effect change vector of the HSI modality, the disturbance effect change vector of the LiDAR modality, and the disturbance attack effect change vector of the SAR modality. Cosine similarity is used to measure the directional consistency of two disturbance attack effect vectors, so that the attack effect directions of the three modalities tend to be consistent during the loss optimization process, and avoid mutual cancellation conflicts. According to the three disturbance effect change vectors given above 、 and , the present invention uses cosine similarity to measure its directional consistency, and the definition of cosine similarity is as follows:

[0088] in 、 and They are the change vectors of HSI, LiDAR and SAR perturbation attacks respectively. Norm. The cosine similarity value ranges from -1 to 1, where the closer it is to 1, the more consistent the direction of the perturbation attack effect change vectors of the two are.

[0089] The collaborative adversarial scheme uses collaborative adversarial loss to ensure that the three modalities can work together, improving the effectiveness of multimodal adversarial attacks. The specific optimization formula for the collaborative adversarial loss is as follows:

[0090] in, In order to cooperate with the loss, during the optimization process, the optimization goal is to minimize , so that the three modal optimization directions remain consistent.

[0091] Based on the aforementioned multimodal collaborative perturbation optimization design, an actual attack can be implemented. Based on the previously preprocessed multimodal remote sensing data (HSI, LiDAR, SAR) and a trained multimodal fusion classification model, a multimodal collaborative adversarial attack process is implemented. Under a white-box attack setting, this invention constructs a joint loss function and employs an optimization algorithm to iteratively update the perturbation, effectively deceiving the fusion model. To ensure the stealth, balance, and coordination of the attack, the invention designs perturbation-limiting loss, adversarial loss, perturbation-balancing loss, and collaborative adversarial loss. The perturbation direction and amplitude of each type of data in the multimodal input are independently controlled to achieve the attack objective.

[0092] The specific implementation process of the embodiment is described as follows: The experimental data used in this paper comes from three sources: hyperspectral imagery (HSI) with 145 valid bands and a spatial size of 145×145, normalized to [0,1] after Z-score normalization; LiDAR imagery, a single-channel elevation map registered with the HSI, with the raw elevation values ​​processed by median filtering and then normalized by linear stretching; and SAR imagery, a single intensity channel, which is first logarithmically transformed to enhance texture information and then normalized to the range [0,1]. Sample construction uses a sliding window with a window size of 11×11 and a step size of 1. For each central pixel, the corresponding HSI, LiDAR, and SAR patches are extracted and stitched together to form a unified trimodal input tensor. The resulting stitching results in 147 sample channels. The resulting valid samples are ultimately divided into training and test sets in an 8:2 ratio. A multimodal classification model is constructed using a mid-term fusion architecture as an example. The model consists of three input branches, one for processing HSI, LiDAR, and SAR data. Modal features are extracted and concatenated in a fusion layer, followed by two fully connected layers for classification. The output dimension is the number of categories. Based on the mid-term fusion model trained on the above dataset, a multimodal collaborative adversarial attack is performed on the model. Perturbation variables are constructed for HSI, LiDAR, and SAR, respectively, with the initial perturbation set to zero. The total loss function constructed is:

[0093] The weights of each loss item are set to: α=1.0, β=0.8, γ=0.5. Among them, the perturbation loss by Expression; Coping with Loss Make the prediction deviate from the true label and tend to the target label; perturb the balance loss Expressed as cosine similarity, it is used to balance the intensity of different modal disturbances; collaborative adversarial loss The model logits output direction consistency is constrained.

[0094] The optimization process uses the Adam optimizer, with an initial learning rate of 0.01 and 100 iterations. In each iteration, the model output under the current perturbation state is calculated, the total loss is back-propagated, and the perturbation is updated, while the perturbation sample value is clipped to [0, 1].

[0095] The disturbance after optimization of the coordinated attack method 、 and , so that the model can deal with the adversarial sample image after adding disturbance 、 and The perturbations are small enough to be imperceptible to the human eye. Without significantly altering the input image, they can effectively mislead the fusion model. The introduction of the collaborative adversarial loss significantly improves the consistency of multimodal perturbations in the discriminant space, preventing the model from relying on a single modality for "self-healing," demonstrating excellent attack stability and cross-modal adaptability.

[0096] For simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because certain steps can be performed in other orders or simultaneously according to the embodiments of the present invention. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0097] In specific implementation, the method proposed in the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. System devices that implement the method, such as computer-readable storage media that store the corresponding computer program of the technical solution of the present invention and computer equipment that runs the corresponding computer program, should also be within the scope of protection of the present invention.

[0098] The following embodiments describe the electronic device provided by the present invention. The electronic device described below and the multimodal collaborative anti-attack method based on the remote sensing image classification task described above can be referenced to each other.

[0099] The electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call logic instructions in the memory to execute a multimodal collaborative counterattack method based on a remote sensing image classification task, which primarily includes the software processing portion of the aforementioned steps.

[0100] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0101] On the other hand, an embodiment of the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the software processing part of the multimodal collaborative anti-attack method based on remote sensing image classification tasks provided by the above methods.

[0102] On the other hand, an embodiment of the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the software processing part of the multimodal collaborative anti-attack method based on remote sensing image classification tasks provided by the above-mentioned methods.

[0103] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0104] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multimodal collaborative counterattack method based on remote sensing image classification tasks, characterized by: Follow these steps: Data preprocessing and model building, including acquiring remote sensing image data containing at least two different modalities, performing preprocessing and spatial alignment; building and training a multimodal fusion remote sensing image classification model; Collaborative perturbation generation, including building perturbation generation processes for each modality in a multimodal fusion remote sensing image classification model; Joint optimization involves optimizing the perturbation generation for each modality with the goal of minimizing a total loss function that incorporates perturbation strength constraints to limit the magnitude of the perturbation for each modality, adversarial effect constraints to bias the model prediction away from the true label, distribution similarity constraints to balance the perturbation strengths of different modalities, and collaborative direction constraints to ensure spatial consistency of the perturbations across the model output. The attack implementation includes adding the optimized modal perturbations to the corresponding original modal data to form multimodal adversarial samples, which are then input into the target model to implement the attack.

2. The multimodal collaborative anti-attack method based on remote sensing image classification task according to claim 1 is characterized by: The remote sensing image data includes hyperspectral images, lidar images and synthetic aperture radar images.

3. The multimodal collaborative anti-attack method based on remote sensing image classification task according to claim 1 is characterized by: The multimodal fusion remote sensing image classification model adopts early fusion, mid-term fusion and / or late fusion strategies.

4. The multimodal collaborative anti-attack method based on remote sensing image classification task according to claim 1 is characterized by: When constructing the perturbation generation process for each modality, a unimodal attack strategy and / or a multimodal attack strategy are adopted. The unimodal attack strategy generates adversarial samples by independently perturbing a single modality data, and the multimodal attack strategy generates adversarial samples by simultaneously perturbing multiple modal data.

5. The multimodal collaborative anti-attack method based on remote sensing image classification task according to claim 1 is characterized by: The total loss function is a weighted combination of four types of constraints, and the weight of each constraint item can be adjusted according to the requirements of the attack task.

6. The multimodal collaborative anti-attack method based on remote sensing image classification task according to claim 1 is characterized by: The distribution similarity constraint is implemented by calculating the similarity of the disturbance vector intensities of different modes to constrain the similarity of the disturbance intensity distribution.

7. The multimodal collaborative anti-attack method based on remote sensing image classification task according to claim 1 is characterized by: The collaborative direction constraint is achieved by respectively calculating the difference vectors of the target model output relative to the original unperturbed input when only one modal data is perturbed, then measuring the directional similarity between the difference vectors, and minimizing the similarity measure value to promote directional consistency.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a multimodal collaborative counter-attack method based on a remote sensing image classification task as described in any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements a multimodal collaborative counterattack method based on a remote sensing image classification task as described in any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that: When the computer program is executed by a processor, it implements a multimodal collaborative counterattack method based on a remote sensing image classification task as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Disturbance perception-based multi-modal remote sensing ground feature image classification method and system

    CN121962782A

  • Multi-modal remote sensing feature image classification method and system based on disturbance perception

    CN121962782B