Large model evaluation method and device based on adversarial attack
Through structural causal models, analyze the confounding effects of large models, screen key samples and generate adversarial samples, solve the problem of unable to effectively identify large model vulnerabilities in the existing technology, realize the automation and systematization of large model intelligence level evaluation, and improve the robustness and credibility of the model.
Patent Information
- Application Number
- CN202510629074.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Testing methods based on existing data sets in the prior art cannot effectively identify specific vulnerabilities or weak links of large models, and cannot automate and systematic the evaluation of large models' intelligent level tests.
By using pre-constructed structural causal model, the confounding effect of confounding factors on the prediction results of large models through confounding paths is analyzed, key samples affected by confounding effects are screened, and adversarial samples are generated through black or white box methods to evaluate the large model.
It realizes the automation and systematization of large-scale intelligent level test evaluation, can effectively identify specific vulnerabilities or weak links of the model, and improves the robustness and credibility of the model.
Smart Images

Figure CN120144484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large model evaluation method and device based on adversarial attack. Background Art
[0002] Large models represented by ChatGPT have breakthrough natural language understanding and generation capabilities, but the training and application of large models are still in the early stages and face major bottlenecks. Whether it is self-training domain large models, domain adaptation and fine-tuning based on open source large models, or commercial large models, they all face the key issue of how to evaluate the intelligence level of large models. Fast and realistic intelligence level evaluation can not only support agile parameter tuning and version iteration to improve training efficiency, but also has important value for the technical selection and credibility assurance of large models.
[0003] Evaluation methods based on existing datasets evaluate the performance of the model based on a fixed set of test data. This method can help us understand the overall performance of the large model to a certain extent. However, due to the complex internal mechanisms of pre-trained large models and the characteristics of high-dimensional input space, traditional testing methods often cannot fully cover the performance of the model under all possible inputs. In addition, testing methods based on existing datasets often cannot effectively identify specific vulnerabilities or weak links in the model.
[0004] Therefore, how to provide automated and systematic domain large-scale model intelligence level test evaluation data and evaluation methods is a major practical problem that needs to be solved urgently. Summary of the invention
[0005] The present invention provides a large model evaluation method and device based on adversarial attack, which is used to solve the defect that the testing method based on the existing data set in the prior art cannot effectively identify the specific loopholes or weak links of the model, and realizes the automation and systematization of the large model intelligence level test evaluation data and evaluation method.
[0006] The present invention provides a large model evaluation method based on adversarial attack, comprising the following steps.
[0007] Use pre-built structural causal models to analyze the confounding effects of confounding factors on the prediction results of large models through confounding pathways; Based on the analysis results of the structural causal model, by comparing the output differences of different prompts, the key samples affected by the confounding effect are screened; For the key samples, generating adversarial samples by using a black box method or a white box method; The large model is evaluated using the adversarial sample.
[0008] According to the large model evaluation method based on adversarial attack provided by the present invention, the construction process of the structural causal model specifically includes: Define the variables of the large model and the causal relationships between the variables; Based on the causal relationships, draw a causal diagram DAG; Establish the structural equations of the variables to mathematically describe the causal relationships between the variables.
[0009] According to the large model evaluation method based on adversarial attack provided by the present invention, analyzing the confounding effect of the confounding factor on the prediction result of the large model through the confounding path specifically includes: Determine the confounding factor and the corresponding confounding path of the confounding factor through the causal diagram DAG; Eliminate the bias by means of controlling variables or interventions to block the confounding path; Through the structural equation, calculate the difference in effect estimation before and after blocking the confounding path, and determine the calculation result of the difference in effect estimation as the confounding effect of each confounding factor on the prediction result of the large model through the confounding path.
[0010] According to the large model evaluation method based on adversarial attack provided by the present invention, the different prompts include simple prompts and complex prompts. Among them, the simple prompt only includes the task name and generalization statement, and the complex prompt includes one or more of detailed task description, solution steps, potential interference items, constraint conditions, and examples; Correspondingly, based on the analysis result of the structural causal model, by comparing the output differences of different prompts, screen the key samples affected by the confounding effect, specifically including: Splice the same input sample with the simple prompt and the complex prompt respectively to form two input sequences; Input the two input sequences into the large model to generate prediction results, and record the probability distribution of each prediction result; Calculate the confidence difference between the two prediction results. If the calculation result of the confidence difference is greater than the preset threshold, determine the input sample as a key sample.
[0011] According to the large model evaluation method based on adversarial attack provided by the present invention, the black box method includes: Traverse all positions of the input sample, replace the word at the i-th position with a mask to form a new input sample ; Determine the importance score of each position of the input sample through the model output: , wherein, is the position importance score of the i-th position of the input sample, It is the encoded vector at the i-th position in the input sample, where i = 1 to n, and n is the total number of positions in the input sample. It is the probability distribution before replacement. It is the probability distribution after replacing the i-th position. Select the input sample position with the highest score for character-level or semantic-level perturbation.
[0012] According to the large model evaluation method based on adversarial attack provided by the present invention, the white-box method includes: Generate the optimal perturbation vector based on the theory of gradient-based adversarial sample generation ; Traverse all positions of the input sample, replace the word at the i-th position with a mask to form a new input sample ; Compare the vector obtained after encoding with the optimal perturbation vector and select the input sample position with the smallest distance as the optimal perturbation position.
[0013] According to the large model evaluation method based on adversarial attack provided by the present invention, the adversarial samples include the following types: Adversarial prompts, injecting interference into the task instructions; Adversarial content, modifying key vocabulary in the input sample.
[0014] According to the large model evaluation method based on adversarial attack provided by the present invention, using the adversarial samples, the metrics for evaluating the large model include one or more of the following: The decrease in accuracy, the attack success rate, semantic consistency, robustness, prompt sensitivity, semantic understanding and generalization ability, defense ability, and causal interpretability.
[0015] The present invention also provides a large model evaluation device based on adversarial attack, including the following modules: The confounding effect analysis module is used to analyze the confounding effect of confounding factors on the prediction results of the large model through the confounding path by using a pre-constructed structural causal model; The key sample screening module is used to screen key samples affected by the confounding effect by comparing the output differences of different prompts based on the analysis results of the structural causal model; The adversarial sample generation module is used to generate adversarial samples for the key samples by black-box method or white-box method; The evaluation module is used to evaluate the large model by using the adversarial samples.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for evaluating a large model based on adversarial attack as described in any one of the above is implemented.
[0017] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for evaluating a large model based on adversarial attack as described in any one of the above is implemented.
[0018] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for evaluating a large model based on adversarial attack as described in any one of the above is implemented.
[0019] A method and device for evaluating a large model based on adversarial attack provided by the present invention utilize a pre-constructed structural causal model to analyze the confounding effect of confounding factors on the prediction results of the large model through confounding paths; based on the analysis results of the structural causal model, by comparing the output differences of different prompts, key samples affected by the confounding effect are screened; for the key samples, adversarial samples are generated by black-box methods or white-box methods; and the large model is evaluated using the adversarial samples. The present invention analyzes the confounding effect of the large model and reduces the influence of confounding factors through causal theory, thereby finding key samples in the dataset; for the key samples, methods for generating adversarial samples in black-box and white-box scenarios are proposed, and the adversarial samples are used for evaluating the large model, which can more effectively evaluate the robustness of the large model. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It is a schematic flowchart of the method for evaluating a large model based on adversarial attack provided by the present invention.
[0022] Figure 2 It is a schematic diagram of the structural causal model of the information extraction process based on the large model provided by the present invention.
[0023] Figure 3 It is a schematic flowchart of the generation of adversarial samples provided by the present invention.
[0024] Figure 4 It is a schematic diagram of the structure of the device for evaluating a large model based on adversarial attack provided by the present invention.
[0025] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0027] The present invention will be specifically described below with reference to the accompanying drawings of the specification. The specific operation methods in the method embodiments can also be applied to the apparatus embodiments or system embodiments. In the description of the present invention, unless otherwise specified, "at least one" includes one or more. "A plurality" means two or more. For example, at least one of A, B, and C includes: A alone, B alone, A and B existing simultaneously, A and C existing simultaneously, B and C existing simultaneously, and A, B, and C existing simultaneously. In the present invention, " / " means "or". For example, A / B may represent A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0028] The present invention will be specifically described below in combination with the specific embodiments.
[0029] In some specific implementation schemes of the present invention, as Figure 1 shown, the present solution provides a large model evaluation method based on adversarial attacks, including: Step 100: Analyze the confounding effect of confounding factors on the prediction results of the large model through the confounding path by using a pre-constructed structural causal model; Based on the analysis result of the structural causal model, screen the key samples affected by the confounding effect by comparing the output differences of different prompts; For the key samples, generate adversarial samples by using the black-box method or the white-box method; Use the adversarial samples to evaluate the large model.
[0030] It should be noted that in the existing large model evaluation schemes, the scale of the test data set is large, and it is impossible to generate corresponding adversarial samples for each sample. The evaluation methods based on the existing data sets cannot fully cover the performance of the model under all possible inputs, and the test methods based on the existing data sets usually cannot effectively identify specific vulnerabilities or weak links of the model.
[0031] Therefore, the present invention analyzes the confounding effect of the large model and reduces the influence of the confounding factors through causal theory, so as to find the key samples in the data set; for the key samples, an adversarial sample generation method in black-box and white-box scenarios is proposed, and the generated adversarial samples are used for large model evaluation, so as to more effectively evaluate the robustness of the large model.
[0032] In some possible embodiments of the present invention, the construction process of the structural causal model specifically includes: Define the variables of the large model and the causal relationships between the variables; Based on the causal relationships, draw a causal graph DAG; Establish the structural equations of the variables to mathematically describe the causal relationships between the variables.
[0033] Specifically, this embodiment provides a construction method of a structural causal model, which reveals the confounding effect of confounding factors such as pre-trained knowledge (K) on the model output (such as backdoor path interference) by constructing a structural causal model (Structural Causal Model, SCM).
[0034] Specifically, in this embodiment, the causal inference theory is introduced, and a structural causal model SCM of the bias effect of the large model is proposed to explore the formation mode and occurrence mechanism of the pre-trained knowledge bias effect; on this basis, the samples that are more likely to cause prediction errors of the large model are located for perturbation, forming a data set that more severely tests the robustness of the large model.
[0035] In possible embodiments, the structural causal model expresses the causal relationships between key elements in the process of processing downstream tasks. Taking the classification paradigm as an example, there are five model variables in the SCM: Variable 1: The prior knowledge K in the large model; Variable 2: The encoder E of the large model; Variable 3: The feature representation X generated by the large model; Variable 4: The answer generation constraint V (such as the set of candidate labels provided to the model or the constraints given in the prompt); Variable 5: The final prediction result Y.
[0036] In some possible embodiments of the present invention, analyzing the confounding effect of confounding factors on the prediction results of the large model through the confounding path specifically includes: Determining the confounding factors and the corresponding confounding paths of the confounding factors through the causal diagram DAG; Eliminating the bias by means of controlling variables or interventions to block the confounding path; Through the structural equation, calculating the difference in effect estimates before and after blocking the confounding path, and determining the calculation result of the difference in effect estimates as the confounding effect of each confounding factor on the prediction results of the large model through the confounding path.
[0037] Specifically, this embodiment provides an implementation manner for analyzing the confounding effect. The causal confounding effect brought by confounding factors such as pre-trained knowledge can be discovered by using the structural causal model.
[0038] In a possible embodiment, as Figure 2 shown, there are the following paths in the SCM: V←K→E means that the large model knowledge directly affects the encoder and the answer generation constraint.
[0039] E→X means that the sample features are generated by the encoder.
[0040] X→Y←V means that the final prediction result jointly determines the predicted label by the features and the answer generation constraint.
[0041] According to the causal inference theory, Figure 2 there is a backdoor path in : X←E←K→V→Y. This path means that there is a confounding factor K in the causal relationship from X to Y. Therefore, when calculating the conditional probability P(Y|X), the influence of K must be considered. That is, although the internal knowledge of the large model brings performance improvement, it also plays the role of a confounding factor.
[0042] Specifically, the backdoor path: X←E←K→V→Y indicates that the pre-trained knowledge (K) as a confounding factor may lead to model prediction bias. Therefore, causal intervention (such as adjusting the prompt constraint V) is needed to weaken the confounding influence and locate the samples vulnerable to interference.
[0043] It can be seen from the above analysis that the causal confounding effect brought by pre-trained knowledge can be discovered by using the structural causal model. And the direct manifestation of this confounding effect is that it is prone to misleading the output of the large model in some cases. The degree of "being prone to misleading" can be reflected by the change in the output results brought by simple prompts and complex (or more semantically rich) prompts. On this basis, a means of locating key data samples is provided, that is, perturbing the key data samples.
[0044] In some possible embodiments of the present invention, the different prompts include simple prompts and complex prompts. Among them, the simple prompt only contains the task name and a generalization statement, and the complex prompt includes one or more of a detailed task description, solution steps, potential interference items, constraints, and examples; Accordingly, based on the analysis results of the structural causal model, by comparing the output differences of different prompts, key samples affected by the confounding effect are screened, specifically including: Splice the same input sample with a simple prompt and a complex prompt respectively to form two input sequences; Input the two input sequences into the large model to generate prediction results, and record the probability distribution of each prediction result; Calculate the confidence difference between the two prediction results. If the calculation result of the confidence difference is greater than the preset threshold, then determine the input sample as a key sample.
[0045] Specifically, this embodiment provides an implementation manner for screening key samples affected by the confounding effect, comparing prompts with different information contents (P1 simple prompt and P2 complex prompt), and screening samples with significant output differences.
[0046] It can be understood that for different prediction tasks, the defined simple prompts and complex prompts are also different.
[0047] In a possible embodiment, in an entity recognition task, P1 only provides the task name and a simple description (relying on the pre-trained knowledge of the model), and P2 provides a detailed task description and solution details (fully activating the model's capabilities). If the outputs of the same sample under P1 and P2 are inconsistent, it is determined as a key sample (the prediction is unstable due to the interference of pre-trained knowledge).
[0048] Specifically, in the entity recognition task of this embodiment, if the given input sample is: s, and the entity type candidate set is: l, the prompts for the entity recognition task are as follows: Prompt 1: Perform an entity recognition task to identify words (entities) with specific meanings in the text, mainly including personal names, place names, organization names, proper nouns, etc., and mark the words to be recognized in the text sequence. Select the most likely result from the label set for output. The input sample is: s, and the candidate label set is: l.
[0049] Prompt 2: Perform an entity recognition task to sequentially identify and label words (entities) with specific meanings in the text. An entity is a type of word with a specific meaning, an objectively existing concrete thing, usually referring to an actually existing organization or institution that plays a role. Entities mainly include personal names, place names, organization names, proper nouns, etc. Mark the words to be recognized in the text sequence. In entity recognition, special attention needs to be paid to special situations such as polysemy of entities, nested entities, or spaced entities. Given the candidate set of labels l, for each selected word, select the most likely entity type result from the label set for output. The input sample is: s. The candidate label set is: l.
[0050] Specifically, when locating key data samples, given a sample, two different prompts are designed to guide the large model to output results. The first prompt P1 only contains less task information, such as the task name and a simple task description. The purpose of P1 is to explore activating the pre-trained knowledge of the large model only through simple prompt information and using the knowledge it remembers to understand and solve the task. The second prompt P2 includes a detailed description of the current task and the details to note when solving the task. The purpose of P2 is to explore activating and prompting the pre-trained knowledge of the large model more fully by providing richer task information, so as to better solve the current task. Taking the entity recognition task as an example: P1 and P2 will get two outputs for the same sample. Among all sample data, by using prompts with two different information contents, find samples where the simple prompt is different from the complex prompt. Such samples, as samples with a large semantic distance from the large model, can be used for perturbation enhancement at the end point.
[0051] For the above key samples, the present invention generates corresponding adversarial samples. Different from the method of generating image adversarial samples, in the text field, since the input text is discrete data, directly adding interference to the whole text usually results in meaningless results. Therefore, in the field of text adversarial sample generation, usually first find the position of the word (phrase) that has the greatest impact on the model in the input sample, and perform character-level and word-level interference on the text data at this position, such as using strategies like spelling mistakes, typos, synonyms, etc., to generate adversarial samples. Therefore, how to find the key position that has the greatest impact on the model is the key problem in text adversarial sample generation. To address this problem, this project analyzes it in two scenarios: black box and white box.
[0052] In some possible implementation manners of the present invention, the black box method includes: Traverse all positions of the input sample, replace the word at the i-th position with a mask to form a new input sample ; Determine the importance score of each position of the input sample through the model output: (1), wherein, is the position importance score of the i-th position of the input sample, is the encoding vector of the i-th position in the input sample, i = 1~n, and n is the total number of positions in the input sample, is the probability distribution before replacement, is the probability distribution after replacing the i-th position; Select the input sample position with the highest score for character-level or semantic-level perturbation.
[0053] Specifically, this embodiment provides an implementation manner of a black-box method. Traverse the input sample positions, replace them with [MASK] and observe the change in the output probability. By calculating the position importance score, select the position with the highest score for perturbation (such as typo correction, synonym replacement).
[0054] Specifically, in the black-box scenario, usually only the output result and its corresponding probability can be obtained. Therefore, this embodiment finds the key position that has the greatest impact on the model through the degree of change in its probability. Specifically, traverse all positions of the input sample, replace the word at the i-th position with MASK, thereby forming a new input , and obtain the importance score of each position through the change in the model output p. The importance score of the i-th position The calculation method is as follows: , According to the obtained importance score , select the key positions: (2), wherein, , represents the influence degree of the i-th position on the model output, and y is the corresponding label.
[0055] Perform interference with different granularities on the j position to generate adversarial samples.
[0056] In some possible implementation manners of the present invention, the white-box method includes: Generate an optimal perturbation vector based on the theory of gradient-based adversarial sample generation ; Traverse all positions of the input sample, replace the word at the i-th position with a mask to form a new input sample ; Compare the vector obtained after encoding with the optimal perturbation vector and select the input sample position with the smallest distance as the optimal perturbation position.
[0057] Specifically, this embodiment provides an implementation of a white-box method to generate an optimal perturbation vector (based on gradient optimization), compare the similarity between the [MASK] position encoding and and select the position with the largest perturbation amplitude.
[0058] Specifically, in the black-box scenario, this embodiment finds the optimal perturbation position by analyzing the model loss and its gradient changes. Specifically, first, an optimal perturbation vector is generated based on the theory of adversarial example generation based on gradients ; secondly, all positions of the input sample are traversed, and the word at the i-th position is replaced with MASK to form a new input ; finally, the vector obtained after encoding is compared with the optimal perturbation vector, and the position with the smallest distance is selected as the optimal perturbation position.
[0059] (3).
[0060] The adversarial example pairs obtained based on the above method have the largest perturbation amplitude to the model and can more effectively evaluate the robustness of the large model.
[0061] In some possible implementation manners of the present invention, the adversarial examples include the following types: Adversarial prompts, injecting interference into the task instructions; Adversarial content, modifying keyword vocabulary in the input sample.
[0062] Specifically, this embodiment provides an implementation manner of adversarial examples, which may include adversarial instructions and adversarial content. Among them, adversarial prompts refer to injecting interference into the task instructions, and adversarial content refers to modifying keyword vocabulary in the input sample. The interference granularity includes character level, word level, sentence level, etc.
[0063] In a possible embodiment, as Figure 3 shown, the input of the large model is divided into two parts: a prompt and the main content. The prompt part is used to tell the model the task type or the type of answer required, and the main content part refers to the text content that the model is expected to process. According to the interference position, the robustness test data set can be divided into three parts: adversarial prompts, adversarial content, and adversarial input. Among them, adversarial prompts are to generate adversarial prompts for the prompt, and the main content remains the clean main content; adversarial content is to generate adversarial main content for the main content, and the prompt remains the clean prompt; adversarial input is to generate adversarial prompts and adversarial main content for both the prompt and the main content. According to the perturbation granularity, the interference can be divided into character level, word level, sentence level, and mixed level. The interference tools include TextBugger, TextFooler, CheckList, and artificial synthesis, etc. Exemplarily, seeFigure 3 In the example part, in the input part of the clean sample, the clean prompt and the clean main content are "Translate the following sentence into English" and "I am very happy today", respectively. By performing character-level, word-level, sentence-level, or mixed-level perturbations on the clean sample, adversarial prompts, adversarial content, and adversarial inputs can be generated. Among them, the adversarial prompt includes the adversarial prompt "Translate the following sentence into English <perturbation content>" and the clean main content "I am very happy today", the adversarial content includes the clean prompt "Translate the following sentence into English" and the adversarial main content "I am very happy today <perturbation content>", and the adversarial input includes the adversarial prompt "Translate the following sentence into English <perturbation content>" and the adversarial main content "I am very happy today <perturbation content>".
[0064] In possible embodiments, the generated adversarial samples are used to evaluate the robustness of the model and test the stability and anti-interference ability of the model under input perturbations.
[0065] Specifically, character-level perturbations (such as typos, spelling mistakes) are used to test the sensitivity of the model to subtle text changes; semantic-level perturbations (such as synonym replacement, sentence pattern adjustment) are used to evaluate the model's understanding ability of semantic consistency; adversarial prompts (such as misleading task instructions) are used to verify the model's resistance ability to instruction interference. The quantitative indicators include the decrease in accuracy: the degree of improvement in the model's prediction error rate after perturbation, and the output consistency: the range of fluctuations in the results of the same problem under different perturbations.
[0066] In possible embodiments, the generated adversarial samples are used to expose potential vulnerabilities of the model and reveal the weak links of the model through targeted attacks.
[0067] Specifically, such as knowledge bias vulnerabilities: aiming at the pre-training knowledge confounding effect (such as backdoor path interference), generating misleading samples to verify whether the model over-relies on prior knowledge. Semantic understanding vulnerabilities: testing the model's parsing ability for complex semantics such as polysemy and entity nesting. Provide specific directions for model tuning (such as adjusting the pre-training data distribution, enhancing the generalization ability of specific tasks).
[0068] In possible embodiments, the generated adversarial samples are used to verify the effectiveness of causal analysis and test the conclusions of the structural causal model (SCM) through adversarial samples.
[0069] Specifically, first, causal model localization, identifying samples vulnerable to confounding factors (such as pre-training knowledge K) through SCM. Adversarial sample generation, specifically generating perturbations for these samples and observing whether the model fails as predicted by the causal model. Verify the accuracy of causal inference and form a closed loop of "theoretical analysis → practical verification".
[0070] In the embodiments of the present invention, through the generation and application of adversarial samples, not only the depth and efficiency of large model evaluation are improved, but also the model is directly promoted to evolve in a more robust, more trustworthy, and more secure direction.
[0071] In some possible embodiments of the present invention, the metrics for evaluating the large model using the adversarial samples include one or more of the following: The decrease in accuracy rate, the attack success rate, semantic consistency, robustness, prompt sensitivity, semantic understanding and generalization ability, defense ability, and causal interpretability.
[0072] Specifically, this embodiment provides an implementation manner of the metrics for evaluating the large model. The above multi-dimensional evaluation metrics together constitute a comprehensive evaluation system for the intelligence level of the large model, overcome the pre-training knowledge bias, improve the model stability and credibility, and provide technical support for the research and development of large models in key fields.
[0073] Specifically, for the evaluation of robustness, the core evaluation objective is to evaluate the stability and anti-interference ability of the model in the face of input perturbations or adversarial attacks. The adversarial sample test is to detect whether the model prediction results change significantly by generating character-level, word-level, or sentence-level adversarial samples (such as typos, synonym replacements). The key metrics include the perturbation amplitude tolerance: the consistency of the model output after perturbation (such as the decrease in accuracy rate), and the backdoor path interference resistance ability: the resistance effect of the model to the confounding effect of pre-training knowledge (analyzed through a causal model).
[0074] For prompt sensitivity, the evaluation objective is to measure the response difference of the model to prompts with different information contents, reflecting the degree of its dependence on pre-training knowledge. The test method is to compare the output results of simple prompts (P1) and complex prompts (P2). The key metrics include output consistency, that is, the consistency of the prediction results of the same sample under P1 and P2 (the greater the difference, the higher the prompt sensitivity), and the degree of knowledge bias, that is, the adaptability of the model to complex task descriptions (such as whether it can overcome the misguidance of pre-training knowledge).
[0075] For the evaluation of semantic understanding and generalization ability, the evaluation objective is to verify the depth of the model's understanding of complex semantics, polysemy, entity nesting, etc. The test methods include key sample tests, observing whether the model can accurately identify samples with distant semantic distances (such as polysemy, rare entities), and cross-domain generalization, evaluating the model performance under unseen tasks or data distributions.
[0076] For the evaluation of defense capabilities, the evaluation objective is to analyze the anti-interference ability of the model under different attack scenarios. In the black-box scenario, adversarial samples are generated only based on the change of output probability to test the model's defense against unknown attacks. In the white-box scenario, the model gradient is used to generate the optimal perturbation to evaluate the model's resistance to the risk of internal mechanism exposure. The key metrics include the attack success rate: the probability that the adversarial sample successfully misleads the model, and the perturbation detection ability: whether the model can identify and filter adversarial inputs.
[0077] For the evaluation of causal interpretability, the evaluation objective is to analyze the causal logic of the model's prediction results through a structural causal model (SCM). The testing methods include confounding effect localization: identifying the interference path of the pre-trained knowledge (K) on the prediction result (Y) (such as the backdoor path X←E←K→V→Y), and the causal intervention effect: by adjusting the answer generation constraint (V), verifying whether the model can weaken the influence of the confounding factor.
[0078] In a specific embodiment, the large model evaluation method based on adversarial attack provided by this solution includes the following steps: Step 1: Construct an SCM to analyze the interference path of the pre-trained knowledge on entity type annotation; Step 2: Input the sample "Peking University was founded in 1898", and use P1 (simple prompt) and P2 (complex prompt) respectively to obtain the output; P1 output: Label "Peking University" as "organization name"; P2 output: Label "1898" as "time" and determine it as a key sample (the time entity recognition ability is not activated by P1); Step 3: Generate an adversarial sample "1898 Mou" at the position of "1898" to test whether the model can still correctly identify it as "time"; Step 4: Calculate the decrease in accuracy (from 95% to 70%) to verify the insufficient robustness of the model.
[0079] The large model evaluation method based on adversarial attack provided by the present invention screens key samples through a causal model, reduces the consumption of full-scale test resources, improves the evaluation efficiency, combines backdoor path analysis and adversarial attack, exposes the model's knowledge bias and semantic understanding defects, and accurately locates vulnerabilities; supports both black-box (commercial model evaluation) and white-box (internal optimization) scenarios, and adapts to high-security requirements In some specific implementation schemes of the present invention, such as Figure 4 shown, this solution provides a large model evaluation device based on adversarial attack. The device includes: A confounding effect analysis module 41, configured to use a pre-constructed structural causal model to analyze the confounding effect of the confounding factor on the prediction result of the large model through the confounding path; The key sample screening module 42 is used to screen key samples affected by confounding effects by comparing the output differences of different prompts based on the analysis results of the structural causal model; The adversarial sample generation module 43 is used to generate adversarial samples for the key samples by black-box methods or white-box methods; The evaluation module 44 is used to evaluate the large model by using the adversarial samples.
[0080] The large model evaluation device based on adversarial attack provided by the embodiments of the present invention has the same implementation principle and beneficial effects as those of the large model evaluation method based on adversarial attack shown in the above embodiments. For the implementation principle and beneficial effects of the large model evaluation method based on adversarial attack shown in the above embodiments, reference can be made, and details will not be elaborated here.
[0081] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the large model evaluation method based on adversarial attack, and the method includes: analyzing the confounding effect of the confounding factor on the prediction result of the large model through the confounding path by using a pre-constructed structural causal model; screening key samples affected by the confounding effect by comparing the output differences of different prompts based on the analysis results of the structural causal model; generating adversarial samples for the key samples by black-box methods or white-box methods; and evaluating the large model by using the adversarial samples.
[0082] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0083] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the large model evaluation method based on adversarial attacks provided by the above-mentioned various methods. The method includes: using a pre-constructed structural causal model to analyze the confounding effect of confounding factors on the prediction results of the large model through confounding paths; based on the analysis results of the structural causal model, by comparing the output differences of different prompts, screening key samples affected by the confounding effect; for the key samples, generating adversarial samples through black-box methods or white-box methods; and using the adversarial samples to evaluate the large model.
[0084] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the large model evaluation method based on adversarial attacks provided by the above-mentioned various methods. The method includes: using a pre-constructed structural causal model to analyze the confounding effect of confounding factors on the prediction results of the large model through confounding paths; based on the analysis results of the structural causal model, by comparing the output differences of different prompts, screening key samples affected by the confounding effect; for the key samples, generating adversarial samples through black-box methods or white-box methods; and using the adversarial samples to evaluate the large model.
[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A large model evaluation method based on adversarial attack, characterized in that: include: Use pre-built structural causal models to analyze the confounding effects of confounding factors on the prediction results of large models through confounding pathways; Based on the analysis results of the structural causal model, by comparing the output differences of different prompts, the key samples affected by the confounding effect are screened; For the key samples, generating adversarial samples by using a black box method or a white box method; The large model is evaluated using the adversarial sample.
2. The large model evaluation method based on adversarial attack according to claim 1 is characterized in that: The construction process of the structural causal model specifically includes: Define the variables of the large model and the causal relationships between the variables; Based on the causal relationship, a causal graph DAG is drawn; A structural equation of each of the variables is established to mathematically describe the causal relationship between the variables.
3. The large model evaluation method based on adversarial attack according to claim 2 is characterized in that: Analyze the confounding effects of confounding factors on the prediction results of large models through confounding paths, including: Determine the confounding factors and the confounding paths corresponding to the confounding factors through the causal graph DAG; Eliminate bias by controlling variables or intervening to block the confounding pathway; The structural equation is used to calculate the difference in effect estimates before and after blocking the confounding pathway, and the calculation result of the difference in effect estimates is determined as the confounding effect of each confounding factor on the prediction result of the large model through the confounding pathway.
4. The large model evaluation method based on adversarial attack according to claim 1 is characterized in that: The different prompts include simple prompts and complex prompts, wherein the simple prompts only include the task name and a generalized statement, and the complex prompts include one or more of the detailed task description, solution steps, potential interference items, constraints, and examples; Accordingly, based on the analysis results of the structural causal model, by comparing the output differences of different prompts, key samples affected by the confounding effect are screened, including: The same input sample is concatenated with simple prompts and complex prompts to form two input sequences; Input the two input sequences into the large model to generate prediction results, and record the probability distribution of each prediction result; The confidence difference between the two prediction results is calculated, and if the calculated result of the confidence difference is greater than a preset threshold, the input sample is determined as a key sample.
5. The large model evaluation method based on adversarial attack according to claim 1 is characterized in that: The black box approach includes: Traverse all positions of the input sample and replace the word at the i-th position with the mask to form a new input sample ; Determine the importance score of each position of the input sample through the model output: , in, Score the position importance of the i-th position of the input sample, is the encoding vector of the i-th position in the input sample, i=1~n, n is the total number of positions of the input sample, is the probability distribution before replacement, is the probability distribution after replacing the i-th position; The input sample positions with the highest scores are selected for character-level or semantic-level perturbations.
6. The large model evaluation method based on adversarial attack according to claim 1 is characterized in that: The white-box approach includes: Generating optimal perturbation vectors based on the theory of adversarial sample generation based on gradient ; Traverse all positions of the input sample and replace the word at the i-th position with the mask to form a new input sample ; Will The encoded vector and the optimal perturbation vector After comparison, the input sample position with the smallest distance is selected as the optimal perturbation position.
7. The large model evaluation method based on adversarial attack according to claim 1 is characterized in that: The adversarial samples include the following types: Adversarial prompts, which inject distractions into task instructions; Adversarial content modifies key words in input samples.
8. The large model evaluation method based on adversarial attack according to claim 1 is characterized in that: Using the adversarial sample, the indicators for evaluating the large model include one or more of the following: The accuracy drop, attack success rate, semantic consistency, robustness, cue sensitivity, semantic understanding and generalization ability, defense capability and causal interpretability.
9. A large model evaluation device based on adversarial attack, characterized in that: include: The confounding effect analysis module is used to analyze the confounding effects of confounding factors on the prediction results of large models through confounding pathways using pre-built structural causal models; A key sample screening module, used to screen key samples affected by confounding effects by comparing output differences of different prompts based on the analysis results of the structural causal model; An adversarial sample generation module, used to generate adversarial samples for the key samples by a black box method or a white box method; An evaluation module is used to evaluate the large model using the adversarial sample.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the large model evaluation method based on adversarial attack as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Text poisoning detection method for multi-channel manufacturing industry data
CN117473391A
Track prediction method under scene fusion based on space-time structure causal model
CN117933397A
Artificial intelligence model security automatic evaluation method oriented to general service scene
CN118627059A
Black box large language model testing method based on adversarial sample migration
CN119204158A
Method and device for training debiased multi-modal large language model for medical care
CN119361165A
Cited By
Causal effect robustness test method based on adversarial disturbance
CN122388478A
A method for testing causality effect robustness based on adversarial perturbation
CN122388478B