Multi-modal alignment enhancement migration attack method oriented to VLP model

By enhancing the images and text of the VLP model multiple times, building a many-to-many alignment relationship, and using the BERT model for reverse attack enhancement, the problem of insufficient cross-modal adversarial attack in the existing technology is solved, and efficient attacks of adversarial samples between different models are achieved.

CN120219875APending Publication Date: 2025-06-27YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510276059.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the multimodal adversarial attack against VLP models, the cross-model migration is insufficient and the multimodal alignment relationship cannot be effectively utilized, resulting in poor effectiveness of adversarial samples when migrating between different models.

Method used

By enhancing images and text multiple times, a many-to-many image-text alignment relationship is constructed, and reverse attack enhancement is combined with the BERT model, mixed text sets are generated and optimized to improve cross-model attack capabilities against samples.

Benefits of technology

It significantly improves the cross-model migration of adversarial samples, enhances the attack success rate of adversarial samples between different VLP models, and is suitable for a variety of VLP model architectures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219875A_ABST
    Figure CN120219875A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal alignment enhancement migration attack method for a VLP model, and the method comprises the steps: carrying out the multi-time enhancement of an image in a selected clean sample, obtaining an enhanced image, carrying out the reverse attack enhancement of each text in the clean sample, obtaining an enhanced text, combining the enhanced text with an original text, obtaining a mixed sample set, and carrying out the recognition of the mixed sample set. And guiding optimization of the mixed sample set by the enhanced image to obtain an adversarial sample, and then optimizing the adversarial image by taking the adversarial sample as guidance to obtain the adversarial sample. According to the method, through image enhancement and text enhancement, a many-to-many image-text alignment relationship is constructed, and the cross-model attack ability of the adversarial sample is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence security, and more specifically, relates to a multi-modal alignment enhanced transfer attack method for VLP models. Background Art

[0002] In recent years, deep learning technology has been widely applied in various fields and has become the core driving force in the research of artificial intelligence. Among them, the Vision-Language Pre-training (VLP) model, as the current mainstream deep learning model architecture, has demonstrated excellent performance in many fields. For example, image recognition, object detection, multi-modal visual question answering, image-text retrieval, etc. VLP shows great potential in solving complex problems and improving the level of intelligence. However, with the expansion of the application scope of VLP, potential security hazards have gradually emerged.

[0003] Among them, the emergence of adversarial examples poses a huge threat to VLP models. Adversarial examples are generated by adding fine-grained perturbations to the original input data. These perturbations are usually almost imperceptible to the human eye but can significantly mislead the predictions of VLP models. For example, in an image classification task, a tiny pixel-level modification may cause the VLP model to misclassify an image into a completely different category. This phenomenon indicates that VLP models may exhibit great vulnerability when facing adversarial attacks, thus posing a severe challenge to their reliability in security-critical applications. Adversarial examples have invisibility and cross-model transferability. Invisibility means that adversarial examples are visually not much different from clean examples, and cross-model transferability means that attackers can generate adversarial examples by using the information of a local white-box proxy model, and this sample can attack an unseen black-box target model. As the most threatening type of attack in current adversarial attacks, cross-model transfer attacks make it extremely challenging for deep learning models to be deployed in some tasks with high security requirements because attackers can achieve attacks without accessing the target model.

[0004] Currently, adversarial attacks can be divided into single-modal adversarial attacks and multi-modal adversarial attacks. Single-modal adversarial attacks are mainly designed for single-modal models and do not perform well on multi-modal models because they do not fully consider multi-modal interaction characteristics. Although recent multi-modal attack methods for VLP models attempt to jointly perturb the image and text modalities, they are still limited by one-to-one or one-to-many alignment patterns, resulting in insufficient cross-model transferability of the generated adversarial examples. Specifically manifested as:

[0005] (1) Limited alignment: relying on fixed one-to-one or one-to-many image-text pairs, unable to cover the many-to-many alignment characteristics during the training of VLP models;

[0006] (2) Overfitting risk: The generated adversarial samples are overly adapted to the white-box proxy model and are difficult to transfer to the black-box target model.

[0007] Therefore, there is an urgent need for an attack method that can make full use of the multi-modal alignment relationship and improve the cross-model transferability of adversarial samples to discover the vulnerabilities of VLP models. Summary of the Invention

[0008] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a multi-modal alignment enhanced transfer attack method for VLP models. By image enhancement and text enhancement, a many-to-many image-text alignment relationship is constructed to enhance the cross-model attack ability of adversarial samples.

[0009] To achieve the above-mentioned invention purpose, the multi-modal alignment enhanced transfer attack method for VLP models of the present invention includes the following steps:

[0010] S1: Select a multi-modal clean sample according to actual needs, denote the image as x I , and the corresponding text as x T,d , where d = 1, 2,..., D, and D represents the number of texts;

[0011] S2: Enhance the image x I N times to obtain N enhanced images N = 2D - 1;

[0012] S3: Perform reverse attack enhancement on the text to obtain the enhanced text corresponding to each text x T,d The specific method of reverse attack enhancement is as follows: Determine K keywords from each text x

[0013] respectively, and perform masking processing on each keyword to obtain K masked texts T,d Input the masked text into the pre-trained BERT model, predict P synonyms that best match the keywords, and replace the K keywords with different combination schemes to obtain Q alternative texts, Q = (P + 1) - 1; K - 1;

[0014] Randomly select D enhanced images from N - 1 enhanced images as the reference image set represents the d-th reference image, and then use the following formula to screen out the enhanced text set from the Q alternative texts of each text x T,d

[0015]

[0016] Among them, f T represents the text encoder in the target VLP model, and f I represents the image encoder in the target VLP model; || || represents taking the norm;

[0017] S4: Merge the enhanced text and the original text x T,d to obtain a mixed text set Optimize the mixed text set to obtain an optimized mixed text set Extract the optimized text corresponding to the enhanced text as the adversarial text The optimization objective expression of the mixed text set is as follows:

[0018]

[0019] Among them, represents a reference image set composed of N - 1 enhanced images and the original image x I ;

[0020] S5: Guided by D adversarial samples iteratively optimize the image perturbation ε I to guide the generation of adversarial images The expression of the generation target is as follows:

[0021]

[0022] Among them, represents an image set composed of the adversarial image and its N - 1 enhanced images ;

[0023] The multi-modal alignment enhanced transfer attack method of the present invention for the VLP model enhances the images in the selected clean samples multiple times to obtain enhanced images, performs reverse attack enhancement on each text in the clean samples to obtain enhanced texts, merges the enhanced texts and the original texts to obtain a mixed sample set, optimizes the mixed sample set with the enhanced images to obtain adversarial samples, and then optimizes the adversarial images guided by the adversarial samples, thereby obtaining adversarial samples.

[0024] The present invention has the following beneficial effects:

[0025] 1) The present invention generates enhanced texts by enhancing images and replacing keyword vocabularies, constructs a multi-modal alignment relationship, simulates the VLP model training process, and significantly improves the cross-model transferability of adversarial samples;

[0026] 2) The present invention jointly optimizes image and text perturbations to enhance the adversarial effect of cross-modal interaction;

[0027] 3) Through experimental verification, the present invention improves the success rate of transfer attacks for different models and can be applied to alignment-based z (CLIP) and fusion-based (ALBEF, TCL) VLP models, with wide applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flowchart of the specific implementation of the multi-modal alignment enhanced transfer attack method for VLP models according to the present invention;

[0029] Figure 2 is an example diagram of the multi-modal alignment enhanced transfer attack method for VLP models in this embodiment;

[0030] Figure 3 is an example diagram of reverse attack enhancement in this embodiment;

[0031] Figure 4 is an example diagram of determining keywords in this embodiment;

[0032] Figure 5 is the attack success rate of the present invention under different numbers of enhanced texts in this embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0033] The following describes the specific implementation of the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.

[0034] Embodiment

[0035] Figure 1 is a flowchart of the specific implementation of the multi-modal alignment enhanced transfer attack method for VLP models according to the present invention. Figure 2 is an example diagram of the multi-modal alignment enhanced transfer attack method for VLP models in this embodiment. As Figure 1 and Figure 2 shown, the specific steps of the multi-modal alignment enhanced transfer attack method for VLP models according to the present invention include:

[0036] S101: Select clean samples:

[0037] Select a multi-modal clean sample according to actual needs, and denote the image as x I , and the corresponding text as x T,d , where d = 1, 2,..., D, and D represents the number of texts.

[0038] S102: Image enhancement:

[0039] For image x I Perform N enhancements to obtain N enhanced images N = 2D - 1.

[0040] The image encoder f of the VLP model I Generally, it is a Vision Transformer (ViT) or a Convolutional Neural Network (CNN). To enable the enhanced images to adapt to image encoders of different architectures, the present invention constructs a Random Augmentation Pool (RAP) to diversify the enhancement of images. The specific method is as follows: Set several enhancement operations according to the image encoder in the target VLP model to form a random augmentation pool. Each time an enhancement is performed, randomly select more than one enhancement operation from the random augmentation pool to enhance image x I Perform enhancement. In this embodiment, the enhancement operations include image patch swapping, patch shuffling, scaling, flipping, and padding. Finally, image x I Will generate multiple enhanced image variants, thus realizing the expansion of image diversity.

[0041] S103: Reverse attack enhancement:

[0042] To generate accurate text and construct a many-to-many text-image relationship, the present invention performs Reverse Attack Augmentation (RAA) on the text to obtain the enhanced text corresponding to each text x T,d Figure 3 Is an example diagram of reverse attack enhancement in this embodiment. As Figure 3 Shown, the specific method of reverse attack enhancement in the present invention is as follows:

[0043] First, determine K keywords from each text x T,d . In this embodiment, the importance of each word in the text is evaluated to determine the keywords. Figure 4 Is an example diagram of determining keywords in this embodiment. As Figure 4 Shown, denote the current text as x T , the number of words is M. For each word w T In text x j Perform masking to obtain the masked text

[0044]

[0045] Among them, j = 1, 2,..., M, and MASK represents a preset mask.

[0046] The masked text ​Input together with the original text x T into the text encoder f of the target VLP model T to obtain text embeddings, and calculate the KL divergence of the two text embeddings as the masked word w j in the text x T 's importance score

[0047]

[0048] where KL() represents calculating the KL divergence of the two text embeddings.

[0049] Select the K words with the highest importance scores from the M words as the keywords of the current text x T 's keywords.

[0050] After determining the keywords, use the BERT model to replace the keywords, and generate accurate text descriptions for each image by maximizing the image-text similarity. That is, mask each keyword separately to obtain K masked texts Input the masked texts into the pre-trained BERT (Bidirectional Encoder Representation from Transformers) model, predict P synonyms that best match the keywords, and replace the K keywords with different combination schemes to obtain Q candidate texts, Q=(P + 1) K -1. The BERT model is a self-encoding model based on Transformer, which can learn language representations from large-scale unsupervised text data. The core feature of BERT is to use a bidirectional attention mechanism to comprehensively capture the relationship between each word in the sentence and its context. Through the Masked Language Model (MLM), BERT randomly masks some words in the sentence and then predicts these masked words based on the context information. Compared with traditional unidirectional language models, the words predicted by BERT can better integrate into the context semantics of the sentence, making the generated text more fluent and natural, and significantly improving the performance of the model in natural language processing tasks.

[0051] Next, select D enhanced images from the N - 1 enhanced images as the reference image set denote the d-th reference image, and then use the following formula to screen out the enhanced text set from the Q candidate texts of each text x T,d 's Q candidate texts

[0052]

[0053] Among them, f T represents the text encoder in the target VLP model, and f I represents the image encoder in the target VLP model. || || represents taking the norm. This formula calculates the similarity between the image embedding and the text embedding, and generates text that can more accurately describe the image by maximizing this similarity by modifying the words. In this embodiment, the reference image is the D enhanced images with the highest similarity to the original image x I among the N - 1 enhanced images.

[0054] S104: Generate adversarial text:

[0055] Multiple enhanced image - text pairs are generated through the previous steps. These enhanced image - text pairs, together with the original image - text pair, enrich the alignment information in the embedding space, provide sufficient representation ability for the many - to - many image - text relationship, and indicate that the embeddings of multiple images can match the embeddings of multiple texts. In the present invention, the original text and the enhanced text are mixed, and together with the enhanced image, they are input into the target VLP model to obtain the image embedding and the text embedding, and the text perturbation is optimized with the image embedding as the guide. The optimization goal is to make the adversarial text far away from the enhanced image embedding in the embedding space, that is, to reduce the alignment score. The specific method is as follows:

[0056] Merge the enhanced text and the original text x T,d to obtain a mixed text set Optimize the mixed text set to obtain an optimized mixed text set Extract the optimized text corresponding to the enhanced text from it as the adversarial text The optimization target expression of the mixed text set is as follows:

[0057]

[0058] Among them, represents the reference image set composed of N - 1 enhanced images and the original image x I

[0059] In this embodiment, the optimization of the mixed text set adopts the BERT - Attack model. The BERT - Attack model is a generative adversarial model composed of two BERTs. Its idea is to use one BERT as an adversarial generator to generate adversarial samples, and use another BERT as the model to be attacked.

[0060] S105: Generate adversarial images:

[0061] ​Next, cross-modal guidance is used for adversarial image generation. Guided by the initial adversarial sample embedding text and aiming to minimize the image-text similarity, the image perturbation is iteratively optimized to generate adversarial images, making the adversarial images far from the adversarial text in the feature space. The specific method is as follows:

[0062] With D adversarial samples as the guidance, iteratively optimize the image perturbation ε I to guide the generation of adversarial images . The expression of the generation target is as follows:

[0063]

[0064] where represents the image set composed of the adversarial image and its N - 1 enhanced images .

[0065] In this embodiment, the iterative optimization of the image perturbation ε I adopts the PGD (Projected Gradient Descent) algorithm. The PGD algorithm is an attack method widely used in the field of adversarial sample generation, which generates adversarial perturbations based on the gradient of the model loss function with respect to the input sample. Its core idea is to iteratively modify the input sample along the direction of the loss function gradient (for maximizing the loss), so that the model makes incorrect predictions for the modified sample and increases the loss of the model as much as possible.

[0066] Through the above operations, it has been ensured that the adversarial image is far from the adversarial text. In practical applications, based on the generated adversarial image, the adversarial text can be further optimized to make the adversarial text even farther from the adversarial image. The specific method is as follows:

[0067] With the adversarial image as the guidance, iteratively optimize the text perturbation ε T,d to guide the optimization of the generation of the adversarial text . The expression of the generation target is as follows:

[0068]

[0069] where represents the text set composed of D optimized adversarial texts , represents the adversarial image and its D - 1 enhanced images .

[0070] In this embodiment, the iterative optimization of the adversarial text also adopts the BERT - Attack model.

[0071] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention. In this embodiment, each multi-modal clean sample includes 1 image and 5 texts. In this embodiment, 2 existing single-modal attacks and 2 multi-modal attacks are selected as comparison methods to compare with the present invention, namely PGD, BERT-Attack, Co-Attack, and SGA. On the Flickr30K dataset, cross-model performance tests of adversarial samples are carried out for VLP models with different architectures, and the evaluation index is the attack success rate. The higher the success rate, the better the attack effect. Table 1 is a comparison table of the evaluation indexes of the present invention and the comparison methods for attacking the ALBEF model and the TCL model. Table 1 is a comparison table of the evaluation indexes of the present invention and the comparison methods for attacking the CLIP(ViT) model and the CLIP(CNN) model.

[0072]

[0073] Table 1

[0074]

[0075] Table 2

[0076] As shown in Table 1 and Table 2, the italic part is the white-box attack, that is, the test model is a white-box model, and its parameters, structure, gradients, etc. can be obtained. As the transferability of the adversarial samples increases, the attack success rate of the white-box attack usually decreases slightly. This decrease is attributed to the improvement of the perturbation generalization ability, avoiding overfitting to the white-box proxy model. In the table, PGD and BERT-Attack are single-modal attacks, and their attack success rates are lower than those of the multi-modal attacks Co-Attack, SGA, and the present invention. Co-Attack and SGA are single-to-one and single-to-many multi-modal adversarial attack strategies respectively. They generate perturbations through limited image-text alignment information, so the transferability of the generated adversarial samples is limited. In contrast, the present invention significantly improves the attack success rate of the transfer attack by using many-to-many relationships in the embedding space to enrich the alignment information. The present invention achieves a high attack success rate in most cases. For example, when using CLIP(ViT) to attack CLIP(CNN), compared with the state-of-the-art results obtained by SGA, the present invention increases the attack success rate of image-text retrieval R@1 by 12.26%.

[0077] To verify the effectiveness of the enhanced text reverse attack of the present invention, a comparative experiment is designed in this embodiment. In the generated mixed text set, different numbers of enhanced texts and original texts are selected and merged to obtain a mixed sample set. Figure 5 is the attack success rate of the present invention in this embodiment under different numbers of enhanced texts. As Figure 5As shown, compared with the method without enhanced text, the method proposed by the present invention has achieved an improvement in the attack success rate on different models.

[0078] Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A multi-modal alignment enhanced transfer attack method for VLP model, characterized in that: The following steps are involved: S1: Select a multimodal clean sample according to actual needs, and let the image be x I , the corresponding text is x T,d , where d = 1, 2, ..., D, D represents the number of texts; S2: for image x I Perform N enhancements to obtain N enhanced images S3: Perform reverse attack enhancement on the text to obtain each text x T,d Corresponding enhanced text The specific method of reverse attack enhancement is: From each text x T,d Determine K keywords respectively, perform mask processing on each keyword respectively, and obtain K masked texts Mask Text Input the pre-trained BERT model, predict the P synonyms that best match the keywords, and replace the K keywords with different combinations to obtain Q candidate texts, Q = (P + 1) K -1; Randomly select D enhanced images from N-1 enhanced images as the reference image set Denotes the dth reference image, and then uses the following formula to get the value from each text x T,d The enhanced text set is selected from the Q candidate texts Among them, f T represents the text encoder in the target VLP model, f I represents the image encoder in the target VLP model, || || represents the norm; S4: Enhanced text and the original text x T,d Merge to get a mixed text set Mixed text sets Optimize to get optimized mixed text set Extract enhanced text from The corresponding optimized text is used as the adversarial text The optimization target expression of the mixed text set is as follows: in, Represents N-1 enhanced images and the original image x I The reference image set constructed; S5: With D adversarial examples To guide, iteratively optimize the image perturbation ε I , guided adversarial image Generate, the expression of the target is as follows: in, Represents adversarial images and its N-1 enhanced images Composed set of images.

2. The multimodal alignment enhanced transfer attack method according to claim 1, characterized in that: The method for image enhancement in step S2 is: according to the image encoder in the target VLP model, a plurality of enhancement operations are set to form a random enhancement pool, and one or more enhancement operations are randomly selected from the random enhancement pool each time the image x is enhanced. I Make enhancements.

3. The multimodal alignment enhanced transfer attack method according to claim 1, characterized in that: The enhancement operations include image block exchange, block reordering, scaling, flipping, and padding.

4. The multimodal alignment enhanced transfer attack method according to claim 1, characterized in that: The specific method of keyword recognition in step S3 is: Remember the current text is x T , the number of words is M, for text x T Each word w j Mask and get the mask text Wherein, j = 1, 2, ..., M, MASK represents a preset mask; Mask Text With the original text x T The text encoder f is input to the target VLP model together T Get the text embedding from the text embedding, and calculate the KL divergence of the two text embeddings as the masked word w j In the text x T Important score Among them, KL() means calculating the KL divergence of two text embeddings; Select K words with the largest importance scores from M words as the current text x T Keywords.

5. The multimodal alignment enhanced transfer attack method according to claim 1, characterized in that: The reference image in step S3 is the N-1 enhanced images that are the same as the original image x I The D enhanced images with the highest similarity.

6. The multimodal alignment enhanced transfer attack method according to claim 1, characterized in that: The optimization of the mixed text set in step S4 adopts the BERT-Attack model.

7. The multimodal alignment enhanced transfer attack method according to claim 1, characterized in that: The image perturbation ε in step S5 I The iterative optimization adopts the PGD algorithm.

8. The multi-modal alignment enhanced transfer attack method according to claim 1, characterized in that: The step S6 also includes optimizing the adversarial text, specifically by: To guide, iteratively optimize the text perturbation ε T,d , guiding the optimization of adversarial text Generate, the expression of the target is as follows: in, Denotes D optimized adversarial texts The text set that constitutes Represents adversarial images and its D-1 enhanced images Composed set of images.

9. The multi-modal alignment enhanced transfer attack method according to claim 8, characterized in that: The adversarial text optimization adopts the BERT-Attack model.