Visual language model vulnerability determination method and device based on shared feature attack
By building source and target models in the visual language model, generating adversarial samples and optimizing perturbations, identifying and enhancing shared adversarial features, the problem of insufficient robustness of visual language model in the prior art is solved, and better migration capabilities and attack performance are achieved.
Patent Information
- Application Number
- CN202510278688.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-10
AI Technical Summary
In identifying and mitigating security threats to visual language models against attacks, the prior art relies too much on specific model features of the source model, limiting its migration capabilities, and ignoring the polarity of the features, resulting in insufficient robustness.
By building source and target models, aggression samples are generated and perturbed, aggression features are obtained and their contribution to output are calculated, model enhancement is used to achieve shared adversarial features, and spatially and frequency domain enhancement are performed. Finally, the enhancement results are substituted into the attack algorithm for perturbation to identify vulnerabilities in the visual language model.
This method significantly improves the robustness and migration capabilities of visual language models, and can show better attack performance on different models, data sets, and tasks, thereby more fully identifying vulnerabilities.
Smart Images

Figure CN120124073A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of network security technology, and in particular, to a method and device for determining vulnerabilities of a vision-language model based on shared feature attacks. Background Art
[0002] Vision-Language Large Models (LVLMs) have attracted significant attention due to their remarkable performance in various multimodal tasks, including image captioning, visual question answering, and multimodal dialogue. Despite these achievements, LVLMs still face severe security challenges. LVLMs process visual and text inputs from different domains, which increases the risk of attackers manipulating the inputs to mislead the model. Therefore, it is necessary to fully identify the robustness of LVLMs and potential vulnerabilities before deployment. However, related methods rely too much on specific model features of the source model, which limits their transferability, and distort all features indiscriminately, ignoring the polarity of features, i.e., the difference between positive and negative features. Therefore, developing a method and device for determining vulnerabilities of a vision-language model based on shared feature attacks to effectively overcome the above defects in the related technologies has become an urgent technical problem in the industry. The present invention has broad application value in the field of artificial intelligence, aiming to study and improve the robustness of multimodal large models, identify and mitigate potential security threats posed by adversarial attacks to artificial intelligence systems. By evaluating the security of the model under different attack scenarios, this technology can provide theoretical support for the design and optimization of defense mechanisms. At the same time, the present invention helps to enhance the interpretability of multimodal large models, analyze the decision-making process of the model using adversarial attack means, reveal its potential vulnerabilities and biases, and thus promote the research on the security and credibility of the model. Summary of the Invention
[0003] In view of the above problems existing in the prior art, the embodiments of the present invention provide a method and device for determining vulnerabilities of a vision-language model based on shared feature attacks.
[0004] In a first aspect, the embodiments of the present invention provide a method for determining vulnerabilities of a vision-language model based on shared feature attacks, including: constructing a source model and a target model, generating adversarial samples on the source model, generating perturbations including optimizations; obtaining adversarial features and calculating the contribution of each adversarial feature to the output, implementing shared adversarial features using model enhancement, and performing spatial enhancement and frequency domain enhancement on the shared adversarial features; substituting the spatial enhancement result and the frequency domain enhancement result into an attack algorithm to perturb the shared adversarial features, and obtaining vulnerabilities of the vision-language model.
[0005] Based on the above content of the method embodiment, in the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, the constructing a source model and a target model includes:
[0006] Source model: y = S(x, t)
[0007] Target model: y = U(x, t)
[0008] Among them, y is the response of the vision-language large model; x is the original image of the visual input; t is the prompt; S is the source model; U is the target model.
[0009] Based on the content of the above method embodiments, the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, generating adversarial samples on the source model includes:
[0010] x adv = x + δ
[0011] Among them, x adv is the adversarial sample; δ is the visual perturbation.
[0012] Based on the content of the above method embodiments, the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, generating an optimized perturbation includes:
[0013]
[0014] ||δ|| p ≤ ε Λ U(x + δ, t) ≠ y'
[0015] Among them, argmax is the symbol for maximizing the optimization; L is the language modeling loss; || || p is the p-norm symbol; ε is the perturbation budget; Λ is the intersection; y' is the label.
[0016] Based on the content of the above method embodiments, the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, obtaining adversarial features and calculating the contribution of each adversarial feature to the output includes:
[0017]
[0018] p = h φ (f θ (x))
[0019] Among them, is the contribution of the projected feature to the output; p i is the projected feature; p′ i is the reference projected feature; x i ' is the reference image; n is the path integral step; x i is the image; h φ is the projection encoder; f θ is the image encoder.
[0020] Based on the content of the above method embodiments, the method for determining vulnerabilities of a vision - language model based on shared - feature attacks provided in the embodiments of the present invention, and performing spatial enhancement and frequency - domain enhancement on the shared adversarial features, includes:
[0021] T s (x) = x + ηx'
[0022] T f (x) = F -1 (F l (x)+(1 - α)F h (x)+αF h (x'))
[0023]
[0024] Among them, T s is the spatial - domain enhancement transformation; T f is the frequency - domain enhancement transformation; F -1 is the inverse discrete Fourier transform; F l is the low - frequency component; F h is the high - frequency component; α is the frequency - domain enhancement intensity control variable; x' is the randomly sampled image; η is the spatial enhancement intensity control variable; F is the discrete Fourier transform; u is the x - axis of the spectrogram; v is the y - axis of the spectrogram; w is the column index of the image; j is the imaginary unit; e is the natural constant; h is the row index of the image; H is the height of the input image; W is the width of the input image.
[0025] Based on the content of the above method embodiments, the method for determining vulnerabilities of a vision - language model based on shared - feature attacks provided in the embodiments of the present invention, substituting the spatial enhancement result and the frequency - domain enhancement result into the attack algorithm to perturb the shared adversarial features, includes:
[0026]
[0027] S = |X′ S |
[0028] F = |X′ f |
[0029]
[0030] ||x adv - x|| p ≤ε
[0031] Among them, G is the enhanced integral gradient; x′ S is the randomly sampled spatially enhanced image; X′ S is the set of spatially enhanced images; IG p is the integral gradient of the projected feature; x'f is the frequency-domain enhanced image of random sampling; X' f is the set of frequency-domain enhanced images; L SAF is to generate the shared adversarial feature loss; P is the projected feature of the input image; p' is the projected feature of the reference image.
[0032] In a second aspect, an embodiment of the present invention provides a visual language model vulnerability determination device based on shared feature attack, including: a first main module for constructing a source model and a target model, generating adversarial samples on the source model, and generating optimized perturbations; a second main module for obtaining adversarial features and calculating the contribution of each adversarial feature to the output, using model enhancement to achieve shared adversarial features, and performing spatial enhancement and frequency-domain enhancement on the shared adversarial features; a third main module for substituting the spatial enhancement result and the frequency-domain enhancement result into an attack algorithm to perturb the shared adversarial features and obtain the vulnerability of the visual language model.
[0033] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0034] at least one processor, at least one memory, and a communication interface; wherein,
[0035] the processor, the memory, and the communication interface communicate with each other;
[0036] the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the method for determining the vulnerability of the visual language model based on shared feature attack provided by any one of the various implementation manners in the first aspect.
[0037] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method for determining the vulnerability of the visual language model based on shared feature attack provided by any one of the various implementation manners in the first aspect.
[0038] The method and device for determining the vulnerability of the visual language model based on shared feature attack provided by the embodiments of the present invention explore the feature extraction patterns of LVLMs, identify the features shared among various models and most effective for adversarial attacks, and interfere with them; the proposed attack algorithm performs cross-model attacks and exhibits good transferability; experiments show that compared with related technologies, the methods for determining the vulnerability of the visual language model based on shared feature attack proposed in the various embodiments of the present invention show better attack performance on different models, datasets, and tasks, and thus can more fully confirm the vulnerabilities. Description of the Drawings
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] Figure 1 Schematic flow chart of the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided by an embodiment of the present invention;
[0041] Figure 2 Schematic structural diagram of the device for determining vulnerabilities of a vision-language model based on shared feature attacks provided by an embodiment of the present invention;
[0042] Figure 3 Schematic physical structure diagram of an electronic device provided by an embodiment of the present invention;
[0043] Figure 4 Schematic visualization effect diagram of the original image and its corresponding enhanced image provided by an embodiment of the present invention;
[0044] Figure 5 Schematic visualization effect diagram of t-SNE from InstructBLIP visual features provided by an embodiment of the present invention;
[0045] Figure 6 Schematic diagram of the attack success rate effect of the VQA task on each dataset in white-box and black-box models provided by an embodiment of the present invention;
[0046] Figure 7 Schematic diagram of the CLIP score results obtained by testing the Flickr30K and MSCOCO datasets in the image captioning task on white-box and black-box models provided by an embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention. Additionally, the technical features in each embodiment or individual embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the structure composition mode, but must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention. If there are step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0048] An embodiment of the present invention provides a method for determining vulnerabilities in a vision-language model based on shared feature attacks. Refer to Figure 1 , the method includes: constructing a source model and a target model, generating adversarial samples on the source model, generating perturbations including optimizations; obtaining adversarial features and calculating the contribution of each adversarial feature to the output, implementing shared adversarial features using model enhancement, and performing spatial enhancement and frequency-domain enhancement on the shared adversarial features; substituting the spatial enhancement result and the frequency-domain enhancement result into an attack algorithm to perturb the shared adversarial features, and obtaining vulnerabilities in the vision-language model.
[0049] Based on the content of the above method embodiment, as an optional embodiment, in the method for determining vulnerabilities in a vision-language model based on shared feature attacks provided by the embodiments of the present invention, the constructing the source model and the target model includes:
[0050] Source model: y = S(x, t) (1)
[0051] Target model: y = U(x, t) (2)
[0052] Where y is the response of the vision-language large model; x is the original image of the visual input; t is the prompt; S is the source model; U is the target model.
[0053] Specifically, let y = F d (h p (f θ (x)), t), f θ is an image encoder, h P is a projector, F dIt is a large language model (LLM). For convenience, the source model is represented by Equation (1), and the target LVLM is represented by Equation (2).
[0054] Based on the content of the above method embodiments, as an alternative embodiment, in the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, generating adversarial examples on the source model includes:
[0055] x adv = x + δ (3)
[0056] where x adv is the adversarial example; δ is the visual perturbation.
[0057] Specifically, in the black-box setting, the attacker knows nothing about the target LVLM, including its architecture and parameters. The focus of the work is on target-free transfer-based attacks. As shown in Equation (3), the adversarial example is represented, where δ represents a carefully designed visual perturbation. The attacker's goal is to generate an adversarial example x adv on the source model S to cause the target model U to crash.
[0058] Based on the content of the above method embodiments, as an alternative embodiment, in the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, generating an optimized perturbation includes:
[0059]
[0060] ||δ|| p ≤ ε Λ U(x + δ, t) ≠ y' (5)
[0061] where argmax is the symbol for maximizing the optimization; L is the language modeling loss; || || p is the p-norm symbol; ε is the perturbation budget; Λ is the intersection; y' is the label. Similar to previous work, the L p norm perturbation budget constraint ||x adv - x|| p ≤ ε is used to control the imperceptibility of the adversarial example.
[0062] Based on the content of the above method embodiments, as an alternative embodiment, in the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, obtaining adversarial features and calculating the contribution of each adversarial feature to the output includes:
[0063]
[0064] p = h φ (f θ (x)) (7)
[0065] Among them, is the contribution of the projection feature to the output; p i is the projection feature; p' i is the reference projection feature; x i ' is the reference image; n is the path integral step; x i is the image; h φ is the projection encoder; f θ is the image encoder.
[0066] Specifically, given the input x ∈ X, the feature map g: X → R n Construct g(x) = (g 1 ,..., g n ). The contribution of the correct prediction at the feature g(x) is a vector A(g(x)) = (a 1 ,..., a n ) ∈ R n . If there exists a feature g(x) such that |a i | > ρ (ρ > 0), then this feature is easily disturbed by adversarial perturbations. This feature is defined as an adversarial feature. To obtain adversarial features, measure the contribution of the feature g(x) to the final output of the LVLMs. There have been many studies on the contribution A(g(x)) of features to model predictions. Simple feature values and gradients roughly reflect A(g(x)). In particular, the gradient reflects the contribution and polarity of the feature. Simply multiplying the feature value by the gradient does not meet the sensitivity requirement. The lack of sensitivity causes the gradient to focus on irrelevant features, making it difficult to identify adversarial features. To address this limitation, Integrated Gradients (IG) is used as an indicator of feature contribution. The IG definition for the projector feature p i includes:
[0067]
[0068] Among them, L is the language modeling loss, and x' i represents the baseline image used as a reference for measuring the contribution. Let represent the contribution of the projector feature to the output, then equations (6) and (7) are obtained.
[0069] Based on the content of the above method embodiments, as an optional embodiment, in the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided by the embodiments of the present invention, the spatial enhancement and frequency-domain enhancement of the shared adversarial features are performed, including:
[0070] T s (x) = x + ηx' (8)
[0071] T ff(x) = F -1 (F l (x)+(1 - α)F h (x)+αF h (x')) (9)
[0072]
[0073] where T s is the spatial domain enhancement transformation; T f is the frequency domain enhancement transformation; F -1 is the inverse discrete Fourier transform; F l is the low - frequency component; F h is the high - frequency component; α is the frequency domain enhancement degree control parameter; x' is the randomly sampled image; η is the spatial enhancement degree control parameter; F is the discrete Fourier transform; u is the x - axis of the frequency spectrum diagram; v is the y - axis of the frequency spectrum diagram; w is the column index of the image; j is the imaginary unit; e is the natural constant; h is the row index of the image; H is the height of the input image; W is the width of the input image.
[0074] Specifically, as Figure 4 shown (the first column represents the original image, and the rest shows different model enhancements applied to the original image, where the first row applies spatial enhancement and the second row applies frequency enhancement), different models f(T(x)) output correct responses because there are inherent shared features between them. To adopt shared adversarial features, a feasible idea is to obtain more source models to mine their common properties. The collected LVLMs are usually inefficient and time - consuming. Therefore, model enhancement is used to achieve this diversity, and its formal definition includes: Let x ∈ X be the input with its true label y true ∈ Y, p represents the prompt, f(x) represents a vision - language model f: X → Y with the language modeling loss L(x, y). If there exists a loss - preserving transformation T(), then a new model f'(x) = f(T(x)) is derived from the original model f. This derivation of the model is called model enhancement. Based on the concept of model enhancement, a new enhancement strategy is proposed to extract different feature representations g(x) that simulate features from different models. Specifically, different gradients of the projector features are extracted using T(x). By integrating this information, overfitting features are neutralized while shared adversarial features are retained. Since the features extracted by spatial enhancement are too similar, overfitting features cannot be fully neutralized in the spatial domain. Spatial transformation cannot be converted into significantly different enhancement models because the differences in the frequency domain are ignored and the generality of features between different models cannot be simulated. Therefore, the frequency domain is extended for model enhancement.
[0075] In spatial enhancement, the input image x is modified by adding a randomly sampled image x' with a certain intensity, thereby introducing information from other images. Under the interference of other additional information, overfitting features are easily changed, while inherent shared features exhibit their robustness. The spatial domain enhancement transformation is defined as shown in Equation (8). The enhanced images in the spatial domain simulate little model diversity. Therefore, frequency enhancement is proposed. In fact, naturally trained models are prone to overfitting to high-frequency features during the training phase. By mixing high-frequency information from the features of other images, the overfitting of adversarial samples to the source model is reduced. The frequency domain enhancement transformation is defined as shown in Equation (9).
[0076] Based on the content of the above method embodiments, as an optional embodiment, in the method for determining vulnerabilities of a vision-language model based on shared feature attacks provided in the embodiments of the present invention, substituting the spatial enhancement result and the frequency domain enhancement result into the attack algorithm to perturb the shared adversarial features includes:
[0077]
[0078] S = |X′ S | (12)
[0079] F = |X′ f | (13)
[0080]
[0081] ||x adv -x|| p ≤ ε (15)
[0082] where G is the integrated gradient of enhancement; x′ S is the spatially enhanced image obtained by random sampling; X′ S is the set of spatially enhanced images; IG p is the integrated gradient of the projected feature; x' f is the frequency domain enhanced image obtained by random sampling; X′ f is the set of frequency domain enhanced images; L SAF is the loss for generating shared adversarial features; P is the projected feature of the input image; p' is the projected feature of the reference image.
[0083] Specifically, the SAF attack algorithm is as shown in Equation (11), and adversarial features are identified by calculating the contribution of each feature. Adversarial features are divided into shared features and overfitting features. Then, shared adversarial features are adopted through spatial and frequency enhancements. Using the enhanced integrated gradient as a guide, the shared adversarial features are perturbed. As Figure 5As shown (left half: clear image and adversarial samples generated by SAF with an update strategy. Right half: clear image and adversarial samples generated by SAF without an update strategy), even when features are damaged, LVLMs still extract similar semantics of the clear image. This is due to the powerful attention mechanism that obtains information from adjacent features of different modalities and damaged features. To adapt to the robustness brought by this powerful mechanism, a dynamic update strategy is proposed. Specifically, the enhanced comprehensive gradient is updated every M iterations to fully disrupt the currently identified shared adversarial features.
[0084] The method for determining vulnerabilities of visual language models based on shared feature attacks provided by the embodiments of the present invention explores the feature extraction patterns of LVLMs, identifies the features shared among various models that are most effective against adversarial attacks, and interferes with them; the proposed attack algorithm conducts cross-model attacks, showing good transferability; experiments show that, compared with related technologies, the methods for determining vulnerabilities of visual language models based on shared feature attacks proposed in the embodiments of the present invention exhibit better attack performance on different models, datasets, and tasks, thus enabling more thorough confirmation of vulnerabilities.
[0085] The method for experimentally verifying each embodiment of the present invention, dataset: In the experiment, five datasets widely used in different downstream tasks were used. For the Visual Question Answering (VQA) task, the VQAv2 dataset, VizWiz dataset, and OKVQA dataset were adopted. Since the visual question answering task is more challenging, the accuracy of some models is inherently low on this task, which makes it difficult to evaluate the effectiveness and transferability of attacks. Therefore, 1000 correctly predicted samples were randomly selected from each dataset for testing. For the image captioning task, 1000 images were selected from the Flickr30K dataset and the MSCOCO dataset respectively for experimental evaluation. Model: Five modern vision-language large models were adopted, including InstructBLIP, BLIP-2 (OPT), BLIP-2 (FlanT5), MiniGPT-v2, and LLaVA. MiniGPT-v2 and LLaVA bridge the modality gap between visual and language information through linear layers, while InstructBLIP and BLIP-2 achieve this purpose through Qformer. Baseline and implementation details: The feature-based attack method NAA, the multi-modal attack method end-to-end attack, the CLIP-based attack, and the Visual Token attack (VTattack) were adopted as the baseline methods of the experiment. Following the common settings in [reference], the maximum perturbation size was set to 8 / 255, and the Projected Gradient Descent (PGD) algorithm was applied for 500 iterations of optimization. The learning rate was set to 0.5 / 255. Evaluation metrics: For the Visual Question Answering (VQA) task, the evaluation method was adopted, which is inefficient and only applicable to small batches of data, and is obviously impractical for large-scale experimental data. The accuracy was accurately evaluated by averaging the matching degree between the answers generated by the model and different subsets of human annotations. The Attack Success Rate (ASR) of the visual question answering task is the decrease in accuracy. For the image captioning task, it was proposed that the traditional evaluation method performs poorly when evaluating large models. Therefore, the CLIP score was adopted to evaluate the attack performance, and this score is used to measure the similarity between the image and the text.
[0086] Use various methods to generate adversarial samples for the target downstream task on InstructBLIP or BLIP-2 (OPT), and evaluate their performance on the five selected models. The results of visual question answering are as Figure 6 shown (the adversarial samples were generated by InstructBLIP and BLIP-2 (OPT) and evaluated on five state-of-the-art large vision-language models. The results with the gray background represent the proposed SAF attack. * indicates a white-box attack. The best results are highlighted in bold, while the overall best performance is marked in red), and the results of image captioning are as Figure 7Shown (adversarial examples are generated on InstructBLIP and BLIP-2 (OPT) and tested on five state-of-the-art large vision-language models), is a summary of some research results. As Figure 6 Shown, transfer attack experiments were conducted on the VQAv2, VizWiz, and OKVQA datasets respectively. The results show that the shared semantic space created by CLIP is different from the feature space of large vision-language models (LVLMs), which leads to the least desirable results. The end-to-end attack (E2E) performs well on the source model but has poor transferability. The feature-based attack method NAA has poor overall attack performance, even worse than the overfitted E2E attack. Therefore, traditional transfer attacks are no longer applicable. Due to the arbitrary destruction of features by the Visual Token Attack (VT-Attack), the attack success rate (ASR) on the source model is sometimes even worse than that on other models. In contrast, although the proposed Shared Adversarial Feature Attack (SAF) achieves comparable attack effects on the source model, it significantly improves the attack performance on other models. As shown by the red highlighted part, SAF shows excellent overall performance on all five models. In addition, an in-depth analysis of the experimental results was carried out. The models are shown to be vulnerable on some datasets. Specifically, when conducting experiments on the VizWiz dataset, a significant increase in the overall attack success rate was observed. When using the VQAv2 dataset, the large vision-language models (LVLMs) show strong robustness, making it more difficult to succeed in the attack and more difficult to transfer between different models. In addition, adversarial examples generated by InstructBLIP are generally better than those generated by BLIP-2 (OPT). It is speculated that the performance of adversarial examples is related to the performance of the model. InstructBLIP was systematically tuned on 13 vision-language datasets, which makes it superior to BLIP-2, and the adversarial perturbations more comprehensively disrupt semantic information during the optimization process. This explains why it is more difficult to achieve attack transfer for models like MiniGPT-v2 and LLaVA that contain more high-level semantic information. Therefore, it is recommended to use more powerful models to generate adversarial examples. Figure 7 Shows that when performing image captioning tasks on the Flickr30K and MSCOCO datasets, the method is superior to the baseline method. Similar advantages and generality are demonstrated in different tasks, not limited to a single task.
[0087] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention can be encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, an embodiment of the present invention provides a visual language model vulnerability determination device based on shared feature attack, and this device is used to execute the visual language model vulnerability determination method based on shared feature attack in the above method embodiments. Refer to Figure 2 , the device includes: a first main module, which is used to implement building a source model and a target model, generating adversarial samples on the source model, and generating including optimized perturbations; a second main module, which is used to implement obtaining adversarial features and calculating the contribution of each adversarial feature to the output, using model enhancement to achieve shared adversarial features, and performing spatial enhancement and frequency domain enhancement on the shared adversarial features; a third main module, which is used to implement substituting the spatial enhancement result and the frequency domain enhancement result into an attack algorithm to perturb the shared adversarial features, and obtaining vulnerabilities of the visual language model.
[0088] The visual language model vulnerability determination device based on shared feature attack provided by the embodiment of the present invention adopts Figure 2 several modules among them, explores the feature extraction patterns of LVLMs, identifies the features shared among various models and most effective for adversarial attacks, and interferes with them; uses the proposed attack algorithm for cross-model attacks, showing good transferability; experiments show that, compared with related technologies, the method for determining visual language model vulnerabilities based on shared feature attack proposed in each embodiment of the present invention shows better attack performance on different models, datasets and tasks, so as to be able to more fully confirm vulnerabilities.
[0089] It should be noted that the device in the device embodiment provided by the present invention, in addition to being used to implement the method in the above method embodiment, can also be used to implement the method in other method embodiments provided by the present invention. The difference is only in setting corresponding functional modules, and its principle is basically the same as that of the above device embodiment provided by the present invention. As long as those skilled in the art, on the basis of the above device embodiment, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and on the premise of ensuring the practicability of the technical solutions, the device in the above device embodiment can be improved, so as to obtain corresponding device type embodiments for implementing the methods in other method type embodiments. For example:
[0090] Based on the content of the above device embodiment, as an optional embodiment, the visual language model vulnerability determination device based on shared feature attack provided in the embodiment of the present invention further includes: a first sub-module, which is used to implement the building of the source model and the target model, including:
[0091] Source model: y = S(x, t)
[0092] Target model: y = U(x, t)
[0093] Wherein, y is the response of the vision - language large model; x is the original image of the visual input; t is the prompt; S is the source model; U is the target model.
[0094] Based on the content of the above - mentioned device embodiment, as an alternative embodiment, the vision - language model vulnerability determination device provided in the embodiments of the present invention based on shared - feature attack further includes: a second sub - module for generating adversarial samples on the source model, including:
[0095] x adv = x + δ
[0096] Wherein, x adv is the adversarial sample; δ is the visual perturbation.
[0097] Based on the content of the above - mentioned device embodiment, as an alternative embodiment, the vision - language model vulnerability determination device provided in the embodiments of the present invention based on shared - feature attack further includes: a third sub - module for generating an optimized perturbation, including:
[0098]
[0099] ||δ|| p ≤ ε Λ U(x + δ, t) ≠ y'
[0100] Wherein, argmax is the symbol for maximizing the optimization; L is the language modeling loss; |||| p is the p - norm symbol; ε is the perturbation budget; Λ is the intersection; y' is the label.
[0101] Based on the content of the above - mentioned device embodiment, as an alternative embodiment, the vision - language model vulnerability determination device provided in the embodiments of the present invention based on shared - feature attack further includes: a fourth sub - module for obtaining adversarial features and calculating the contribution of each adversarial feature to the output, including:
[0102]
[0103] p = h φ (f θ (x))
[0104] Wherein, is the contribution of the projected feature to the output; p i is the projected feature; p' i is the reference projected feature; x i' is the reference image; n is the path integral step; x i is the image; h φ is the projection encoder; f θ is the image encoder.
[0105] Based on the content of the above device embodiment, as an optional embodiment, the visual language model vulnerability determination device based on shared feature attack provided in the embodiments of the present invention further includes: a fifth sub-module, configured to implement spatial enhancement and frequency domain enhancement of the shared adversarial features, including:
[0106] T s (x) = x + ηx'
[0107] T f (x) = F -1 (F l (x) + (1 - α)F h (x) + αF h (x'))
[0108]
[0109] where, T s is the spatial domain enhancement transformation; T f is the frequency domain enhancement transformation; F -1 is the inverse discrete Fourier transform; F l is the low-frequency component; F h is the high-frequency component; α is the frequency domain enhancement degree control quantity; x' is the randomly sampled image; η is the spatial enhancement degree control quantity; F is the discrete Fourier transform; u is the x-axis of the spectrogram; v is the y-axis of the spectrogram; w is the column index of the image; j is the imaginary unit; e is the natural constant; h is the row index of the image; H is the height of the input image; W is the width of the input image.
[0110] Based on the content of the above device embodiment, as an optional embodiment, the visual language model vulnerability determination device based on shared feature attack provided in the embodiments of the present invention further includes: a sixth sub-module, configured to implement perturbing the shared adversarial features by substituting the spatial enhancement result and the frequency domain enhancement result into the attack algorithm, including:
[0111]
[0112] S = |X′ S |
[0113] F = |X′ f |
[0114]
[0115] ||x adv - x||p ≤ ε
[0116] Among them, G is the enhanced integral gradient; x' S is the spatially enhanced image randomly sampled; X' S is the set of spatially enhanced images; IG p is the integral gradient of the projected feature; x' f is the frequency-domain enhanced image randomly sampled; X' f is the set of frequency-domain enhanced images; L SAF is the generated shared adversarial feature loss; P is the projected feature of the input image; p' is the projected feature of the reference image.
[0117] The method of the embodiment of the present invention is implemented relying on an electronic device. Therefore, it is necessary to introduce the relevant electronic device. For this purpose, an embodiment of the present invention provides an electronic device, as Figure 3 shown, the electronic device includes: at least one processor, a communication interface, at least one memory, and a communication bus. Among them, the at least one processor, the communication interface, and the at least one memory complete communication with each other through the communication bus. The at least one processor can call the logical instructions in the at least one memory to execute all or part of the steps of the methods provided by the foregoing method embodiments.
[0118] In addition, when the logical instructions in the foregoing at least one memory are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the method embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0120] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0121] The flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present invention. Based on this understanding, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0122] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "comprising..." do not preclude the presence of additional identical elements in the process, method, article or device comprising the said elements. For any similar expressions such as "predetermined threshold", "preset threshold", etc., if no specific value is indicated, those of ordinary skill in the art can determine their specific values through simple experiments or corresponding debugging.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for determining visual language model vulnerabilities based on shared feature attack, characterized in that: include: Construct source models and target models, generate adversarial samples on the source models, and generate optimized perturbations; obtain adversarial features and calculate the contribution of each adversarial feature to the output, use model enhancement to realize shared adversarial features, and perform spatial and frequency domain enhancement on the shared adversarial features; substitute the spatial enhancement results and frequency domain enhancement results into the attack algorithm to perturb the shared adversarial features and obtain the vulnerabilities of the visual language model.
2. The method for determining a visual language model vulnerability based on a shared feature attack according to claim 1, characterized in that: The constructing of the source model and the target model includes: Source model: y = S (x, t) Target model: y = U (x, t) Among them, y is the response of the visual language model; x is the original image of the visual input; t is the prompt; S is the source model; U is the target model.
3. The method for determining a visual language model vulnerability based on a shared feature attack according to claim 2, characterized in that: The generating of adversarial samples on the source model includes: x adv =x+δ Among them, x adv is the adversarial sample; δ is the visual perturbation.
4. The method for determining a visual language model vulnerability based on a shared feature attack according to claim 3 is characterized in that: The generating includes optimizing the perturbations, including: ||d|| p ≤εΛU(x+δ,t)≠y' Among them, argmax is the symbol for optimizing to the maximum value; L is the language modeling loss; || || p is the p-norm symbol; ε is the perturbation budget; Λ is the intersection; y' is the label.
5. The method for determining a visual language model vulnerability based on a shared feature attack according to claim 4, characterized in that: The obtaining of adversarial features and calculating the contribution of each adversarial feature to the output includes: p=hφ(fθ(x)) in, is the contribution of the projected features to the output; p i is the projection feature; p' i is the reference projection feature; x' i is the reference image; n is the path integration step; x i is an image; h φ is the projection encoder; f θ For image encoder.
6. The method for determining visual language model vulnerabilities based on shared feature attack according to claim 5, characterized in that: The shared adversarial features are spatially enhanced and frequency-domain enhanced, including: T s (x)=x+ηx' T f (x)=F -1 (F l (x)+(1-α)F h (x)+αF h (x')) Among them, T s is the spatial domain enhancement transform; T f is the frequency domain enhancement transform; F -1 is the inverse discrete Fourier transform; F l is the low frequency component; F h is the high frequency component; α is the frequency domain enhancement control amount; x' is the randomly sampled image; η is the spatial enhancement control amount; F is the discrete Fourier transform; u is the x-axis of the spectrum graph; v is the y-axis of the spectrum graph; w is the column index of the image; j is the imaginary unit; e is a natural constant; h is the row index of the image; H is the height of the input image; W is the width of the input image.
7. The method for determining visual language model vulnerabilities based on shared feature attacks according to claim 6, characterized in that: Substituting the spatial enhancement result and the frequency domain enhancement result into the attack algorithm to perturb the shared adversarial feature includes: S=|X’ S | F=|X’ f | ||x adv -x|| p ≤ε Where G is the enhanced integrated gradient; x' S is a randomly sampled spatially enhanced image; X' S For spatially enhanced image collection; IG p is the integrated gradient of the projected feature; x' f is a randomly sampled frequency domain enhanced image; X' f is the frequency domain enhanced image set; L SAF is to generate shared adversarial feature loss; P is the projected feature of the input image; p' is the projected feature of the reference image.
8. A visual language model vulnerability determination device based on shared feature attack, characterized in that: include: The first main module is used to build a source model and a target model, generate adversarial samples on the source model, and generate perturbations including optimization; The second main module is used to obtain adversarial features and calculate the contribution of each adversarial feature to the output, use model enhancement to realize shared adversarial features, and perform spatial and frequency domain enhancement on the shared adversarial features. The third main module is used to substitute the spatial enhancement results and frequency domain enhancement results into the attack algorithm to perturb the shared adversarial features and obtain the vulnerabilities of the visual language model.
9. An electronic device, characterized in that: include: At least one processor, at least one memory and a communication interface; wherein, The processor, memory and communication interface communicate with each other; The memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which cause a computer to execute the method of any one of claims 1 to 7.