A method for generating cross-protocol generalization adversarial samples of an automatic test image evaluation system
Patent Information
- Application Number
- CN202610825671.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-01
AI Technical Summary
[0004]然而,现有图像自动评测系统的对抗鲁棒性仍存在不足,表现在其评测结果可能受到输入图像中细微扰动的影响
Smart Images

Figure CN122676285A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence security and computer vision, specifically to a method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images. Background Technology
[0002] With the development of applications such as image generation, content moderation, and security assessment, automatic image evaluation systems have been gradually developed to automatically evaluate image quality, image-text consistency, image security, and content compliance. Automatic image evaluation systems primarily refer to those built upon multimodal large language models. These systems use multimodal large language models, such as LLaVA and Qwen-VL, as the review body, receiving the image to be evaluated, the evaluation task type, and the evaluation criteria as input. The evaluation criteria consist of reference text or security policies, and the evaluation task types include image quality evaluation, image-text consistency evaluation, image security evaluation, and content compliance evaluation. The automatic image evaluation system then organizes the evaluation image and evaluation criteria into evaluation inputs that can be processed by its internal multimodal large language model, and outputs a score, judgment conclusion, or natural language review reasoning from the internal multimodal large language model. This type of automatic image evaluation system based on multimodal large language models can reduce the cost of manual annotation and improve the efficiency of large-scale evaluation. Its evaluation results can be used for scenarios such as image generation result quality assessment, generated content screening, image content review, training data review, and security compliance assessment.
[0003] Existing automatic image evaluation systems based on multimodal large language models typically employ different evaluation protocols. These protocols refer to the rules by which the automatic image evaluation system, after receiving the image to be evaluated, the evaluation task type, and the evaluation criteria, organizes, decomposes, or packages these elements to construct the actual evaluation input for the internal multimodal large language model, and processes the model's output. Specifically, these protocols include input organization methods, prompt word templates, question generation methods, answer parsing methods, and score aggregation rules. Common evaluation protocols mainly fall into three categories: evaluation methods based on response label probability, evaluation methods based on fine-grained semantic decomposition, and evaluation methods based on natural language review prompts. Evaluation methods based on response label probability typically convert the evaluation criteria into a single binary question, input the image to be evaluated and the question into an internal multimodal large language model, and obtain the evaluation score based on the probability of the model generating positive or negative response labels. Evaluation methods based on fine-grained semantic decomposition decompose the evaluation criteria into multiple fine-grained sub-questions, input the image to be evaluated and each sub-question into an internal multimodal large language model, and then aggregate the model's answers to each sub-question to obtain the final evaluation result. Evaluation methods based on natural language review prompts encapsulate the evaluation criteria into complete review prompt words, and input them along with the image to be evaluated into an internal multimodal large language model, so that the model can directly generate scoring text, judgment conclusions or review reasons.
[0004] However, existing automatic image evaluation systems still lack sufficient adversarial robustness, as their evaluation results can be affected by subtle perturbations in the input image. While attackers cannot modify the reference text, security policies, or the evaluation protocol itself, they can apply imperceptible perturbations to the image being evaluated, causing the automatic image evaluation system to output artificially high image quality scores or image-text consistency scores, or to incorrectly classify risky images as safe or compliant. Existing general adversarial attack methods targeting multimodal large language models are mostly geared towards tasks such as image understanding or text generation, typically relying on specific prompt text, specific output targets, or specific task interfaces, making it difficult to adapt to the differences between different evaluation protocols in automatic image evaluation systems. Since different evaluation protocols differ in input organization, prompt word templates, question generation methods, answer parsing methods, and score aggregation rules, a method for generating adversarial examples with strong cross-evaluation protocol generalization capabilities is needed. This method would ensure that the generated adversarial examples maintain stable influence under different evaluation protocols, allowing for testing and revealing the robustness deficiencies of such automatic image evaluation systems in automatic evaluation scenarios, and providing a foundation for subsequent defense design. Summary of the Invention
[0005] The purpose of this invention is to provide a method for generating cross-protocol generalized adversarial examples for an automatic image evaluation system. First, multiple binary semantic questions for verifying image content are generated based on evaluation criteria composed of reference text or security policies, and semantic constraints are constructed accordingly to improve the probability of positive judgments output by the automatic image evaluation system. Second, high-scoring proxy samples are selected according to a proxy evaluation protocol, and high-scoring representation centers are calculated, and representation constraints are constructed accordingly. Finally, image perturbation is treated as a variable to be optimized. While keeping the structure and parameters of the multimodal large language model in the automatic image evaluation system unchanged, the perturbation values of each pixel in each channel of the image perturbation are updated through multiple rounds of gradient iteration based on the total loss function composed of semantic constraints and representation constraints, thereby obtaining image perturbations that satisfy the perturbation constraints and generating adversarial examples. To achieve the above objective, the technical solution process adopted by this invention is as follows:
[0006] A method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images, characterized by comprising the following steps:
[0007] A method for generating cross-protocol generalized adversarial examples for an automatic image evaluation system, wherein the automatic image evaluation system refers to a multimodal large language model. The automatic evaluation system, which serves as the review body, is characterized by the following steps in its generation method:
[0008] (1.1) Input and optimization preparation: Determine the image to be evaluated The evaluation task type and evaluation criteria; wherein the evaluation criteria are provided in the reference text. or security policy The evaluation task types include at least one of image quality evaluation, image-text consistency evaluation, image security evaluation, and content compliance evaluation; initialization with the image to be evaluated. Image perturbation of the same size And represent the perturbed image as ;
[0009] (1.2) Generation of binary semantic query sets: Determine the target content description used to construct semantic constraints based on the evaluation task type. Description of the target content Semantic component decomposition is performed, and the resulting semantic components are transformed into binary semantic queries that can be answered in affirmative or negative form, thus obtaining a set of binary semantic queries. ;
[0010] (1.3) Semantic constraint loss construction: Set up positive response tags to represent the validity of the target semantics. The perturbed image With the binary semantic query set The binary semantic queries in the input are the review subject of the automatic image evaluation system, namely the multimodal large language model. According to the multimodal large language model Output positive response flag Conditional probabilities are used to construct semantic constraint loss. ;
[0011] (1.4) Determination of high-resolution representation centers: Setting up a set of proxy evaluation protocols Using multiple proxy samples without added image perturbations, the performance of these proxy samples in various proxy evaluation protocols is calculated. The evaluation scores are used to select high-scoring proxy samples; these high-scoring proxy samples are then extracted into the multimodal large language model. The internal state representation is used to determine the high-scoring representation center corresponding to each agent evaluation protocol. ;
[0012] (1.5) Characterization of constraint loss construction: according to the aforementioned agent evaluation protocols Extract the perturbated image In the multimodal large language model Internal state representation in Based on internal state representation With the high-resolution characterization center The distance between them is used to construct the constraint loss. ;
[0013] (1.6) Image perturbation optimization: Preserve the multimodal large language model With the structure and parameters unchanged, the image is perturbed. As variables to be optimized, the semantic constraint loss is jointly minimized under preset perturbation constraints. and the characterization constraint loss And according to the total loss function Image perturbation The gradient is iteratively updated multiple times to obtain the final image perturbation. ;
[0014] (1.7) Adversarial example generation and evaluation: Perturb the final image Applied to the image to be evaluated This yields adversarial examples for an automatic image evaluation system. ; Adversarial samples of the image Compared with the original image to be evaluated Each image is input into the same target image automatic evaluation system to obtain adversarial examples. Corresponding adversarial evaluation score and the original image to be evaluated. Corresponding original evaluation score And determine the adversarial sample of the image based on the difference between the two. The degree of influence on the evaluation results of the automatic target image evaluation system.
[0015] The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images, as described above, is characterized in that the process of generating the binary semantic query set in step (1.2) includes:
[0016] (2.1) When the evaluation task type is image quality evaluation or image-text consistency evaluation, obtain the corresponding reference text. and the reference text Determined as target content description When the evaluation task type is image security evaluation or content compliance evaluation, the corresponding security policy is obtained. and in accordance with security policies Based on the prohibited content, risk categories, or compliance requirements stipulated in the regulations, generate or select safe alternative descriptions that conform to the security strategy from preset harmless scenario descriptions as the target content descriptions. ;
[0017] (2.2) Description of the target content Semantic component decomposition is performed to extract at least one semantic component from the target object, object quantity, object attributes, spatial relationships, and scene background; each extracted semantic component is then converted into a binary semantic query. And enable the binary semantic query A positive response indicates that the corresponding semantic component is true, resulting in a binary semantic query set. ,in, This indicates the number of binary semantic queries.
[0018] The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images, as described above, is characterized in that the semantic constraint loss construction process in step (1.3) includes:
[0019] (3.1) Set up a positive response tag to indicate that the target semantics are true. The affirmative response flag Includes at least one of "yes" or "yes";
[0020] (3.2) The perturbed image Each with binary semantic query set Each binary semantic query in The images are combined into text-image input pairs, and these text-image input pairs are then input into the multimodal large language model of the automatic image evaluation system. ;
[0021] (3.3) Obtain the multimodal large language model Output positive response flag conditional probability And calculate the semantic constraint loss based on multiple conditional probabilities. :
[0022] (1)
[0023] The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images, as described above, is characterized in that the process of determining the high-resolution representation center in step (1.4) includes:
[0024] (4.1) Set up a set of proxy evaluation protocols ,in Indicates the number of proxy evaluation protocols; the proxy evaluation protocols The evaluation method can be rewritten and constructed based on the evaluation method based on response tag probability, the evaluation method based on fine-grained semantic decomposition, and the evaluation method based on natural language review prompts. The methods include rewriting the evaluation basis into a binary question that is different from the target evaluation protocol expression but semantically similar, adopting a sub-question decomposition method that is different from the target evaluation protocol, or constructing review prompt words that are different from the target evaluation protocol expression but consistent with the evaluation goal.
[0025] (4.2) For any agent evaluation protocol Obtain a proxy sample set consisting of publicly available image datasets or undisturbed images with the same distribution as the current evaluation task, and evaluate them according to the proxy evaluation protocol. The evaluation input is constructed using surrogate samples and evaluation criteria, and then fed into the multimodal large language model in the automatic image evaluation system. During this forward evaluation process, the evaluation score corresponding to each proxy sample is obtained, and the internal state representation corresponding to each proxy sample is recorded at the representation extraction position before the score is generated.
[0026] (4.3) Select the highest-scoring proxy samples based on the evaluation scores to form the proxy evaluation agreement. The corresponding high-scoring proxy sample set and read from the recorded internal state representation Internal state representation of each high-scoring surrogate sample ,in, This represents a high-scoring proxy sample. This represents the multimodal large language model in the automatic image evaluation system;
[0027] (4.4) For the high-scoring proxy sample set Internal state representation of each high-scoring surrogate sample The average calculation yields the agent evaluation agreement. Corresponding high-resolution characterization center ,in This indicates that the automatic image evaluation system is in the proxy evaluation protocol. The internal representation center corresponding to the formation of high-scoring evaluation results.
[0028] The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images, as described above, is characterized in that the process of constructing the representation constraint loss in step (1.5) includes:
[0029] (5.1) The perturbed image The evaluation criteria are based on the agency evaluation agreement. The various agent evaluation agreements in The corresponding evaluation input is constructed and fed into the multimodal large language model of the automatic image evaluation system. Subsequently, the internal state representation corresponding to the perturbed image is extracted at the same representation extraction location as the high-resolution surrogate sample, denoted as... ;
[0030] (5.2) Calculate the internal state representation of the image after perturbation. High-resolution characterization centers corresponding to each agent's evaluation agreement The mean square error distance between them is used to calculate the constraint loss. :
[0031] (2)
[0032] The cross-protocol generalization adversarial example generation method for an automatic evaluation system of test images, as described above, is characterized in that the image perturbation optimization process in step (1.6) includes:
[0033] (6.1) Loss based on semantic constraints and characterizing constraint loss Construct the total loss function:
[0034] (3)
[0035] in, These are the weighting parameters used to balance semantic constraints and representational constraints;
[0036] (6.2) At the preset disturbance amplitude Solving the final image perturbation under constraints ,in, Indicates image perturbation The change in any pixel channel does not exceed the preset perturbation range. ;
[0037] (6.3) Preserve the aforementioned multimodal large language model The structure and parameters remain unchanged, only the image is perturbed. As the variable to be optimized, the image perturbation is initialized to satisfy... The image is subjected to random perturbations, and the total loss function is calculated through backpropagation in each iteration to affect the image perturbation. The gradient; with the first Taking the round of iteration as an example, the image perturbation is updated using the projective gradient descent method:
[0038] (4)
[0039] in, Indicates the iteration step size. Represents a symbolic function. This means performing element-wise projection on the updated image perturbation tensor to maintain the perturbation values of each pixel channel. Within the range;
[0040] (6.4) Repeat the loss calculation, gradient calculation, perturbation update and projection processing until the preset number of iterations or the total loss function converges to obtain the final image perturbation. .
[0041] Compared with existing technologies, this invention decomposes the target content description into multiple binary semantic queries and uses positive response tags to construct semantic constraints, enabling image perturbation to be optimized around the target semantic content. Simultaneously, this invention filters high-scoring proxy samples based on a proxy evaluation protocol, extracts the internal state representations of these high-scoring proxy samples, calculates high-scoring representation centers, and then uses representation constraints to make the internal representations of the perturbated image closer to the internal representations of the high-scoring regions. Through these methods, the image adversarial examples generated by this invention do not depend on the scoring rules of a specific evaluation protocol, and can be applied to different evaluation tasks and different evaluation protocols in image automatic evaluation systems based on multimodal large language models, providing a basis for improving the robustness of such image automatic evaluation systems. Attached Figure Description
[0042] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0043] Figure 1This diagram illustrates the input and output of the automatic image evaluation system based on a multimodal large language model as described in this invention, as well as the relationship between the system and the evaluation protocol and the multimodal large language model as the review subject.
[0044] Figure 2 This is a schematic diagram of the overall process of a cross-protocol generalized adversarial sample generation method for an automatic evaluation system of test images according to the present invention. Detailed Implementation
[0045] As previously mentioned, this invention proposes a cross-protocol generalized adversarial sample generation method for an automatic evaluation system of test images. To make the purpose, technical solution and advantages of this method clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.
[0046] (1) Overall Framework: The automatic image evaluation system described in this invention and its relationship with the evaluation protocol and the multimodal large language model as the evaluation subject are as follows: Figure 1 As shown. This invention revolves around image perturbation optimization, and the overall framework is as follows. Figure 2 As shown, after acquiring the image to be evaluated, seven steps are executed sequentially: input and optimization preparation, binary semantic query generation, semantic constraint construction, high-scoring representation center determination, representation constraint construction, image perturbation optimization, and adversarial example generation and evaluation. Specifically, the binary semantic query generation and semantic constraint construction branches correspond to the semantic constraint branch, which transforms the target content of interest in the evaluation task into a binary semantic query and constrains the perturbated image to approximate the target semantics at the output response level. The high-scoring representation center determination and representation constraint construction branches correspond to the representation constraint branch, which estimates the high-scoring representation region through a proxy evaluation protocol and high-scoring proxy samples, and constrains the perturbated image to approximate the representation of that region at the internal representation level, thereby improving the score of adversarial examples. Using the semantic constraints and representation constraints constructed in the above two branches respectively, joint optimization is performed in the image perturbation optimization step to generate the final adversarial example and obtain the adversarial evaluation result.
[0047] Specifically, the semantic constraint branch does not directly rely on the final scoring rules of a particular evaluation protocol. Instead, it focuses on the semantic discrimination content that is of common concern to different evaluation tasks, providing constraints on the semantic and response levels of the model output for image perturbation optimization. The representation constraint branch utilizes surrogate evaluation protocols to select high-scoring surrogate samples and extracts their internal state representations within the automatic image evaluation system, providing a reference at the internal representation level. Finally, under preset perturbation constraints, the semantic and representation constraints are jointly optimized to obtain the final image perturbation and generate an adversarial image that can generalize across evaluation protocols.
[0048] (2) Input and Optimization Preparation: Obtain the image to be evaluated and initialize the image perturbation to be optimized using random noise. Specifically, let the image to be evaluated be... The image perturbation to be optimized is initialized as random noise. The image after perturbation is The image perturbation To compare with the image to be evaluated An optimizable perturbation tensor of the same size, where each element corresponds to the change in a single pixel in the image to be evaluated. For example... Figure 1 As shown, along with the image to be evaluated, the input also includes the evaluation criteria and the evaluation task type: the evaluation criteria are derived from the reference text. or security policy The evaluation task types include at least one of image quality evaluation, image-text consistency evaluation, image security evaluation, and content compliance evaluation.
[0049] Binary semantic query generation: Based on the evaluation task type, determine the target content description used to construct semantic constraints, and decompose the target content description into semantic components. The specific process of constructing the target content description and decomposing its semantic components is as follows:
[0050] The first step, if the evaluation task type is image quality evaluation or image-text consistency evaluation, is to obtain the corresponding reference text. And identify it as the target content description If the evaluation task type is image security evaluation or content compliance evaluation, then obtain the corresponding security policy. and in accordance with security policies Based on the prohibited content, risk categories, or compliance requirements stipulated in the regulations, generate or select safe alternative descriptions that conform to the security policy from a pre-set harmless scenario description library as the target content description. The above methods enable the description of the target content. Corresponding to the core discrimination criteria in different evaluation tasks, it provides semantic reference for subsequent image perturbation generation. Thus, subsequent steps can construct image perturbations around a unified target content description, making the perturbation generation process no longer dependent on the scoring rules of a specific evaluation protocol, but rather oriented towards the semantic discrimination content that is of common concern to different evaluation tasks, thereby improving the generalization of adversarial samples across evaluation protocols.
[0051] The second step is to describe the target content. Semantic component decomposition is performed to extract semantic information such as target object, number of objects, object attributes, spatial relationships, and scene background. Each extracted semantic component is then transformed into a binary semantic query that can be answered affirmatively or negatively by a multimodal large language model in an automatic image evaluation system, resulting in a set of binary semantic queries. ,in, This indicates the number of binary semantic queries; each binary semantic query... Used to point to the target content description One of the discriminable semantic components, and for A positive response indicates that the corresponding semantic component is true. For example, when the target content describes... When the target message is "three ceramic cups are placed on the kitchen countertop", a binary semantic query could be "Does the image contain three cups?", "Are the cups made of ceramic?", or "Are the cups located on the kitchen countertop?", etc. The above process describes the target content. The image automatic evaluation system decomposes the data into basic semantic facts that are verified in different evaluation protocols. This allows the subsequent image perturbation optimization process to no longer directly target the scoring template of a certain evaluation protocol, but instead to construct visual evidence around these basic semantic facts that can be perceived and verified by the multimodal large language model in the image automatic evaluation system.
[0052] (3) Semantic constraint construction: Set up positive response tags to indicate that the target semantics are true. For example, "Yes" or "Yes". The perturbed image... Each with binary semantic query set Each binary semantic query in The input pairs are combined into image-text pairs and input into a multimodal large language model in an automatic image evaluation system. The output response of the multimodal large language model for each binary semantic query is obtained. Positive responses from the model output are then marked. The conditional probability is denoted as ,in, This represents the multimodal large language model in the automatic image evaluation system. The conditional probability is used to measure the positive response label. When used as the target response, the perturbed image Binary semantic query The strength of visual evidence provided by the semantic fact it points to. The semantic constraint loss is calculated and defined based on multiple conditional probabilities:
[0053] (5)
[0054] The semantic constraint loss, as a semantic component in the total loss function, serves as image perturbation in the subsequent joint optimization process. Provides gradient directions that make multimodal large language models more likely to output positive responses, i.e., perturbation of the image. Under various binary semantic queries, there is a greater tendency to output a positive response, thus making Achieve stable score improvements under different evaluation protocols.
[0055] (4) Determination of high-scoring representation centers: To construct representation constraints, high-scoring representation centers are first calculated on the surrogate evaluation protocol using surrogate samples. The surrogate evaluation protocol and surrogate samples without image perturbation are introduced to avoid directly relying on the specific prompt templates, scoring rules, or score extraction methods of the target evaluation protocol, thereby reducing the risk of overfitting the representation constraints to a single evaluation protocol and providing a stable reference representation for generalization across evaluation protocols.
[0056] Specifically, a set of proxy evaluation protocols is set up. Among them, the evaluation agreements of each agent Independent of the target evaluation protocol used in the testing phase, this proxy evaluation protocol organizes, decomposes, or packages the images and criteria to be evaluated in a manner similar to but different from the target protocol. This constructs the actual evaluation input fed into the internal multimodal large language model, thereby estimating the image automatic evaluation system's internal representation preferences for high-scoring images using different scoring methods. (Each proxy evaluation protocol...) The evaluation method can be rewritten and constructed based on response tag probability, fine-grained semantic decomposition, and natural language review prompts. The methods include rewriting the evaluation criteria into binary questions that are different from the target evaluation protocol but semantically similar, adopting a sub-question decomposition method that is different from the target evaluation protocol, or constructing review prompt words that are different from the target evaluation protocol but consistent with the evaluation goal.
[0057] For any proxy evaluation protocol First, a proxy sample set is obtained, consisting of undisturbed images from a publicly available image dataset or images with the same distribution as the current evaluation task. These proxy samples are not used as the original images for generating adversarial examples; they are only used for offline estimation of high-scoring representation regions. Then, the proxy evaluation protocol is followed. The evaluation input is constructed using surrogate samples and evaluation criteria, and then fed into the multimodal large language model in the automatic image evaluation system. In this forward evaluation process, the evaluation score corresponding to each proxy sample is obtained, and the internal state representation of each proxy sample is recorded at the representation extraction position before the score is generated. Finally, based on the evaluation scores, the highest-scoring proxy samples are selected to form the proxy evaluation protocol. The corresponding high-scoring proxy sample set and read from the recorded internal state representation The internal state representations corresponding to each high-scoring proxy sample are used for subsequent high-scoring representation center calculations.
[0058] In the above process, the following principles are adopted for obtaining the internal state representation: Since the multimodal large language model has usually completed the fusion encoding of the visual information of the proxy sample and the evaluation input text information before generating the scoring result, judgment conclusion, or review text, and formed an internal state for predicting subsequent evaluation results, the internal state extraction position before scoring generation is taken as the representation extraction position. The preferred representation extraction position is after the last evaluation prompt word and immediately adjacent to the decoding position before the generation of the scoring result, judgment conclusion, or review text. The internal state at this position can represent the internal evaluation judgment formed by the multimodal large language model before outputting the evaluation result, and is denoted as the internal state representation. ,in, This represents a high-scoring proxy sample. This represents the multimodal large language model in the automatic image evaluation system. For the set... The agent evaluation protocol is obtained by averaging the internal state representations of each high-scoring agent sample. Corresponding high-resolution characterization centers:
[0059] (6)
[0060] in, Used to represent the automatic image evaluation system in the proxy evaluation protocol The center of the internal representation region corresponding to the formation of high-scoring evaluation results is used as a reference for the subsequent construction of representation constraints.
[0061] (5) Representation constraint construction: The previous step was used to construct the representation constraints for each agent evaluation protocol. Determine the high-resolution characterization center This step extracts the perturbed image using the same protocol and at the same representation extraction location. The internal state representation of , and constrain it to be close to the corresponding Specifically, the perturbed image The evaluation criteria are based on the agency evaluation agreement. The various agent evaluation agreements in Construct the corresponding evaluation input and input it into the multimodal large language model of the automatic image evaluation system. Then, the internal state representation is extracted at the same score generation location as the high-scoring proxy sample, denoted as... Due to the high-resolution characterization center It is based on the agency evaluation agreement High-scoring proxy sample set The internal state characterization is obtained by averaging, therefore and They reside in the same internal state space; the distance between them is used to represent the perturbed image. In the proxy evaluation agreement The degree of deviation relative to the high-resolution representation region. The representation constraint loss is calculated and defined based on the mean square error distance between the perturbed image's internal state representation and multiple high-resolution representation centers:
[0062] (7)
[0063] The representation constraint loss, as a representation component in the total loss function, becomes image perturbation in the subsequent joint optimization process. Provide the gradient direction that makes the internal state representation of the perturbed image closer to the center of the high-resolution representation, that is, make the perturbed image At the internal representation level, it approximates the high-scoring representation centers under each proxy evaluation protocol. The representation constraints do not directly depend on the scoring rules of the target evaluation protocol, but rather utilize the internal representation region formed by high-scoring samples under the proxy evaluation protocol to provide a cross-evaluation protocol representation reference for perturbation optimization.
[0064] Image perturbation optimization: obtaining semantic constraint loss and characterizing constraint loss Then, construct the total loss function:
[0065] (8)
[0066] in, These are the weight parameters used to balance semantic constraints and representational constraints. Furthermore, regarding the perturbation amplitude... Solving for the optimal image perturbation under constraints by minimizing This causes the perturbation image to tend to produce the target response at the output level and to tend to approximate the representation region corresponding to the high-scoring sample at the internal representation level.
[0067] (9)
[0068] in, Indicates image perturbation The change in any pixel channel does not exceed the preset perturbation range. During the optimization process, the multimodal large language model is maintained. The parameters remain unchanged, only the image is perturbed. As variables to be optimized, first, the image perturbation is initialized to satisfy... The random perturbation; in each iteration, the perturbed image is... Input multimodal large language model The semantic constraint loss, representation constraint loss, and total loss function are calculated, and the total loss function is used to calculate the effect of image perturbation on the image through backpropagation. The gradient is used to update the perturbation values of each pixel channel in the image perturbation. The gradient is then used to update the perturbation values of each pixel channel in the image perturbation. Taking the round of iteration as an example, the image perturbation is updated using the projective gradient descent method:
[0069] (10)
[0070] in, Indicates the iteration step size. Indicates the first The gradient of the total loss function with respect to image perturbation during each iteration. Represents a symbolic function. This means performing element-wise projection on the updated image perturbation tensor so that the perturbation values of each pixel channel remain constant. Within the range; simultaneously, for the perturbed image The pixel values are constrained element-wise to keep them within the range of legal image values. The loss calculation, gradient acquisition, perturbation update, and projection processing are repeated until a preset number of iterations is reached or the total loss function converges, yielding the final image perturbation. .
[0071] (6) Adversarial example generation and evaluation: Perturb the final image Applied to the image to be evaluated This yields adversarial image samples for input into the target image automatic evaluation system. .Will The image is input into an automatic target image evaluation system to obtain an adversarial evaluation score or adversarial evaluation result. Meanwhile, the unperturbed image to be evaluated... Input the same target image into an automatic evaluation system to obtain the original evaluation score or the original evaluation result. .according to and The differences between them determine the extent to which the generated adversarial examples affect the evaluation results of the automatic target image evaluation system. Compared to If a change occurs in a preset direction or a preset magnitude, it indicates that the adversarial example can alter the output of the automatic image evaluation system under the corresponding evaluation task, thereby completing the generation of the adversarial example and its effect evaluation.
[0072] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0073] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for generating cross-protocol generalized adversarial examples for an automatic image evaluation system, wherein the automatic image evaluation system refers to a multimodal large language model. The automatic evaluation system, which serves as the review body, is characterized by the following steps in its generation method: (1.1) Input and optimization preparation: Determine the image to be evaluated. The evaluation task type and evaluation criteria; wherein the evaluation criteria are provided in the reference text. or security policy The evaluation task types include at least one of image quality evaluation, image-text consistency evaluation, image security evaluation, and content compliance evaluation; initialization with the image to be evaluated. Image perturbation of the same size And represent the perturbed image as ; (1.2) Generation of binary semantic query sets: Determine the target content description used to construct semantic constraints based on the evaluation task type. Description of the target content Semantic component decomposition is performed, and the resulting semantic components are transformed into binary semantic queries that can be answered in affirmative or negative form, thus obtaining a set of binary semantic queries. ; (1.3) Semantic constraint loss construction: Set up positive response tags to represent the validity of the target semantics. The perturbed image With the binary semantic query set The binary semantic queries in the input are the review subject of the automatic image evaluation system, namely the multimodal large language model. According to the multimodal large language model Output positive response flag Conditional probabilities are used to construct semantic constraint loss. ; (1.4) Determination of high-resolution representation centers: Setting up a set of proxy evaluation protocols Using multiple proxy samples without added image perturbations, the performance of these proxy samples in various proxy evaluation protocols is calculated. The evaluation scores are used to select high-scoring proxy samples; these high-scoring proxy samples are then extracted into the multimodal large language model. The internal state representation is used to determine the high-scoring representation center corresponding to each agent evaluation protocol. ; (1.5) Characterization of constraint loss construction: according to the aforementioned agent evaluation protocols Extract the perturbated image In the multimodal large language model Internal state representation in Based on internal state representation With the high-resolution characterization center The distance between them is used to construct the constraint loss. ; (1.6) Image perturbation optimization: Preserve the multimodal large language model With the structure and parameters unchanged, the image is perturbed. As variables to be optimized, the semantic constraint loss is jointly minimized under preset perturbation constraints. and the characterization constraint loss And according to the total loss function Image perturbation The gradient is iteratively updated multiple times to obtain the final image perturbation. ; (1.7) Adversarial example generation and evaluation: Perturb the final image Applied to the image to be evaluated This yields adversarial examples for an automatic image evaluation system. ; Adversarial samples of the image Compared with the original image to be evaluated Each image is input into the same target image automatic evaluation system to obtain adversarial examples. Corresponding adversarial evaluation score and the original image to be evaluated. Corresponding original evaluation score And determine the adversarial sample of the image based on the difference between the two. The degree of influence on the evaluation results of the automatic target image evaluation system.
2. The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images according to claim 1, characterized in that, The process of generating the binary semantic query set in step (1.2) includes: (2.1) When the evaluation task type is image quality evaluation or image-text consistency evaluation, obtain the corresponding reference text. and the reference text Determined as target content description When the evaluation task type is image security evaluation or content compliance evaluation, the corresponding security policy is obtained. and in accordance with security policies Based on the prohibited content, risk categories, or compliance requirements stipulated in the regulations, generate or select safe alternative descriptions that conform to the security strategy from preset harmless scenario descriptions as the target content descriptions. ; (2.2) Description of the target content Semantic component decomposition is performed to extract at least one semantic component from the target object, number of objects, object attributes, spatial relationships, and scene background; each extracted semantic component is then converted into a binary semantic query. And enable the binary semantic query A positive response indicates that the corresponding semantic component is true, resulting in a binary semantic query set. ,in, This indicates the number of binary semantic queries.
3. The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images according to claim 1, characterized in that, The process of constructing the semantic constraint loss in step (1.3) includes: (3.1) Set up a positive response tag to indicate that the target semantics are true. The affirmative response flag Includes at least one of "yes" or "yes"; (3.2) The perturbed image Each with binary semantic query set Each binary semantic query in The images are combined into text-image input pairs, and these text-image input pairs are then input into the multimodal large language model of the automatic image evaluation system. ; (3.3) Obtain the multimodal large language model Output positive response flag conditional probability And calculate the semantic constraint loss based on multiple conditional probabilities. : (1)。 4. The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images according to claim 1, characterized in that, The process of determining the high-resolution characterization center in step (1.4) includes: (4.1) Set up a set of proxy evaluation protocols ,in Indicates the number of proxy evaluation protocols; the proxy evaluation protocols The evaluation method can be rewritten and constructed based on the evaluation method based on response tag probability, the evaluation method based on fine-grained semantic decomposition, and the evaluation method based on natural language review prompts. The methods include rewriting the evaluation basis into a binary question that is different from the target evaluation protocol expression but semantically similar, adopting a sub-question decomposition method that is different from the target evaluation protocol, or constructing review prompt words that are different from the target evaluation protocol expression but consistent with the evaluation goal. (4.2) For any agent evaluation protocol Obtain a proxy sample set consisting of publicly available image datasets or undisturbed images with the same distribution as the current evaluation task, and evaluate them according to the proxy evaluation protocol. The evaluation input is constructed using surrogate samples and evaluation criteria, and then fed into the multimodal large language model in the automatic image evaluation system. During this forward evaluation process, the evaluation score corresponding to each proxy sample is obtained, and the internal state representation corresponding to each proxy sample is recorded at the representation extraction position before the score is generated. (4.3) Select the highest-scoring proxy samples based on the evaluation scores to form the proxy evaluation agreement. The corresponding high-scoring proxy sample set and read from the recorded internal state representation Internal state representation of each high-scoring surrogate sample ,in, This represents a high-scoring proxy sample. This represents the multimodal large language model in the automatic image evaluation system; (4.4) For the high-scoring proxy sample set Internal state representation of each high-scoring surrogate sample The average calculation yields the agent evaluation agreement. Corresponding high-resolution characterization center ,in This indicates that the automatic image evaluation system is in the proxy evaluation protocol. The internal representation center corresponding to the formation of high-scoring evaluation results.
5. The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images according to claim 1, characterized in that, The process of constructing the characterization constraint loss in step (1.5) includes: (5.1) The perturbed image The evaluation criteria are based on the agency evaluation agreement. Various agent evaluation agreements in The corresponding evaluation input is constructed and fed into the multimodal large language model of the automatic image evaluation system. Subsequently, the internal state representation corresponding to the perturbed image is extracted at the same representation extraction location as the high-resolution surrogate sample, denoted as... ; (5.2) Calculate the internal state representation of the image after perturbation. High-resolution characterization centers corresponding to each agent's evaluation agreement The mean square error distance between them is used to calculate the constraint loss. : (2)。 6. The method for generating cross-protocol generalized adversarial examples for an automatic evaluation system of test images according to claim 1, characterized in that, The image perturbation optimization process described in step (1.6) includes: (6.1) Loss based on semantic constraints and characterizing constraint loss Construct the total loss function: , (3) in, These are the weighting parameters used to balance semantic constraints and representational constraints; (6.2) At the preset disturbance amplitude Solving the final image perturbation under constraints ,in, Indicates image perturbation The change in any pixel channel does not exceed the preset perturbation range. ; (6.3) Preserve the aforementioned multimodal large language model The structure and parameters remain unchanged, only the image is perturbed. As the variable to be optimized, the image perturbation is initialized to satisfy... The image is subjected to random perturbations, and the total loss function is calculated through backpropagation in each iteration to affect the image perturbation. The gradient; with the first Taking the round of iteration as an example, the image perturbation is updated using the projective gradient descent method: , (4) in, Indicates the iteration step size. Represents a symbolic function. This involves element-wise projection of the updated image perturbation tensor to maintain the perturbation values of each pixel channel. Within the range; (6.4) Repeat the loss calculation, gradient calculation, perturbation update and projection processing until the preset number of iterations or the total loss function converges to obtain the final image perturbation. .