Artificial intelligence model safety benchmark evaluation method

By defining the benchmark evaluation scheme and iterative formula to generate adversarial samples, combined with the weighted normalization method, the problems of missing benchmarks and inconsistent indicators in the security evaluation of artificial intelligence models are solved, and the comparability and reliability of the evaluation results are achieved.

CN120296733AActive Publication Date: 2025-07-11DATA SPACE RES INST

Patent Information

Application Number
CN202510353850.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing technology lacks a benchmark definition for the security evaluation of artificial intelligence models, too many evaluation indicators and lack of comprehensive scoring methods, and the inactive sample generation process is not dynamic enough, resulting in inconsistent and unreliable evaluation results.

Method used

By defining a benchmark evaluation scheme, including attack parameter lists and iterative formulas, generating adversarial samples, combining benchmark models and data sets, white box and black box evaluations are supported, and comprehensive scoring is adopted using a weighted normalization method.

Benefits of technology

The horizontal and vertical comparability of the evaluation results is achieved, which reduces the difficulty of user selection, improves the evaluation efficiency and credibility, and enhances the reliability of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296733A_ABST
    Figure CN120296733A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an artificial intelligence model safety benchmark evaluation method. According to the method, benchmark evaluation on the security of the algorithm model is realized, and quantitative analysis on the security of the evaluated model is realized on the theoretical basis by giving a clear definition of benchmark evaluation, so that the evaluation result has significance in comparison with other models in the transverse direction and comparison with models of different versions in the longitudinal direction; different attack intensities and differences of different evaluation indexes in the evaluation process are comprehensively considered, and a comprehensive score calculation formula of multi-dimensional indexes is given, so that a score result is intuitive and easy to understand in numerical value, the model selection difficulty of the public in practical application of an artificial intelligence model is greatly reduced, and through an adversarial sample dynamic generation mechanism, the evaluation accuracy is improved. The labor time cost in the adversarial sample generation process is greatly reduced, the possibility that the evaluated model obtains the adversarial sample is avoided, and the safety and credibility of the evaluation result are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a method for benchmarking the security of an artificial intelligence model. Background Art

[0002] Artificial intelligence is one of the most rapidly developing fields in today's technology, and it is widely used in various industries. With the popularity of large models in the past two years, people have gradually begun to realize the security issues of artificial intelligence models in their respective application scenarios. At the same time, to accurately describe the security of artificial intelligence models, researchers have proposed a variety of evaluation metrics in many dimensions. For example, there are no less than 50 evaluation dimensions for the current security evaluation of large models, and the relationships between each evaluation dimension are complex. In practical applications, the following problems exist:

[0003] (1) Lack of a clear definition of benchmark evaluation. When currently evaluating the security of artificial intelligence models, due to the lack of a clear definition of a unified benchmark evaluation, different evaluation agencies or individuals often use different attack methods, data sets, or evaluation algorithms, etc., resulting in inconsistent inputs in the evaluation process. The evaluation results cannot be compared horizontally with other models and cannot be compared vertically with different versions of the same model. The evaluation results can only qualitatively indicate that there are certain problems with the security of the evaluated model, but cannot quantitatively analyze the security of the evaluated model.

[0004] (2) Too many evaluation metrics and lack of comprehensive scoring of multi-dimensional metrics. When evaluating the security of artificial intelligence models, a certain model often leads in some evaluation metrics and lags in other evaluation metrics. Currently, the specific scores of the evaluated model in each evaluation dimension are generally listed, and the model user selects according to the business characteristics by himself / herself. However, due to the large number of evaluation metrics and the complex relationships between the metrics, users often cannot intuitively distinguish the high and low security levels between models from the overall perspective. For example, in the ethical security dimension, a certain model may score 30 points and be in the leading position in the industry, while in the numerical calculation dimension, 90 points may be in the leading position in the industry. And the score value ranges of some evaluation dimensions are from 0 to 1 (such as accuracy), and the score value ranges of some evaluation dimensions are from 0 to 100. At the same time, under a certain attack intensity, the evaluation of the model may be 50 points, but if the attack intensity is increased, the evaluation of the model may only be 30 points. Therefore, a comprehensive scoring method is needed, which comprehensively considers different attack intensities and the differences of different evaluation metrics in the evaluation process, and gives a comprehensive score that intuitively reflects the security of the evaluated model.

[0005] (3) Adversarial examples lack a dynamic generation mechanism. When evaluating the security of an artificial intelligence model, adversarial examples need to be generated first, and the generation process of adversarial examples is the attack process on the artificial intelligence model. In previous attack processes, the structural characteristics of the model to be evaluated were often not considered, and fixed adversarial examples were used for evaluation. On the one hand, the generation process of adversarial examples mostly relies on manual annotation, which is time-consuming and laborious. On the other hand, the model to be evaluated can obtain adversarial examples by reserving backdoors and conduct targeted training based on the known adversarial examples, resulting in unreliable evaluation results.

[0006] Therefore, to solve the above problems, how to invent a security benchmark evaluation method for artificial intelligence models has become a very important technical research topic in this field. Summary of the Invention

[0007] The purpose of the present invention is to provide a security benchmark evaluation method for artificial intelligence models to solve the problems in the above background technology, including the lack of a clear definition of benchmark evaluation, too many evaluation indicators, the lack of comprehensive scoring of multi-dimensional indicators, and the lack of a dynamic generation mechanism for adversarial examples in the current security evaluation of large models.

[0008] To achieve the above purpose, the present invention provides a security benchmark evaluation method for artificial intelligence models, including the following steps:

[0009] S1. Benchmark scheme judgment: According to the application scenario to which the model to be evaluated belongs, judge whether there is a benchmark evaluation scheme; if not, establish an attack scheme and a benchmark evaluation scheme.

[0010] S2. Attack scheme establishment: Define a list of attack parameters, including the learning rate σ, the upper limit of perturbation amplitude ε, and the maximum number of iterations T. The larger the value of the attack parameter, the greater the attack intensity.

[0011] S3. Benchmark evaluation scheme construction: Include a benchmark model, a benchmark data set, an associated attack scheme, and evaluation indicators. The benchmark model is a model recognized by the industry, and the benchmark data set contains original samples and their correct labels.

[0012] S4. Dynamic generation of adversarial examples: Generate an adversarial example set \(x\) through an iterative formula adv , and the formula is:

[0013]

[0014] where is the adversarial example generated in the \(t\)-th iteration; is the adversarial example generated in the (t + 1)-th round of iteration; θ is the model weight; σ is the step size (learning rate) for each iteration, controlling the magnitude of each update; sign is the sign function, used to obtain the direction of the gradient; y is the correct result label of the sample x; is the partial derivative of the loss function J with respect to ; Clip x,ε is the clipping function, ensuring that the generated adversarial example is within the ε-neighborhood of the original sample;

[0015] S5. Benchmark model evaluation: Load the benchmark model weights and calculate the scores score of each evaluation metric i,j , including accuracy (ACC) and log-likelihood deviation (ALDP);

[0016] S6. Evaluation type judgment: Support white-box evaluation and black-box evaluation. For white-box evaluation, the source code of the model to be evaluated is required, and for black-box evaluation, only adversarial examples are used;

[0017] S7. Attack and calculation of the model to be evaluated: Generate adversarial examples according to the evaluation type and calculate the scores of the evaluation metrics;

[0018] S8. Comprehensive score calculation: After normalizing the scores of the evaluation metrics, calculate the final benchmark score SScore through the weighted formula. The formula is:

[0019]

[0020] where Score is the comprehensive score of the model to be evaluated, and Score base is the score of the benchmark model.

[0021] As a further improvement of this technical solution, after iterating T times through the formula for generating adversarial examples in step S4, an adversarial example corresponding to a sample in the benchmark dataset will be generated. After traversing each sample x in the benchmark dataset, an adversarial example set will be obtained through the above formula;

[0022] Since each benchmark evaluation scheme is associated with multiple attack schemes, after replacing the attack parameters in the formula for generating adversarial examples in step S4 with the attack parameter values in the corresponding attack scheme, an adversarial example set for the corresponding attack scheme j can be obtained

[0023] As a further improvement of this technical solution, when using the ACC evaluation algorithm in step S5, the formula for calculating the score of the evaluation metric is:

[0024]

[0025] Among them, n represents the number of correctly classified adversarial samples, and N represents the total number of adversarial samples. The value range of this indicator is from 0 to 1. The larger the value, the higher the security of the model.

[0026] As a further improvement of this technical solution, when the ALDP evaluation algorithm is used in step S5, the calculation formula for the evaluation index score is:

[0027]

[0028] Where N is the total number of samples, x i is the i-th sample, is the likelihood logarithm of the model for the sample x i , and is the likelihood logarithm of the true data distribution for the sample x i . The value range of this indicator is from 0 to 1. The larger the evaluation index score, the higher the security of the model.

[0029] As a further improvement of this technical solution, when generating adversarial samples in the white-box evaluation of step S6, the network structure of the model to be evaluated is loaded and targeted attacks are performed according to its gradient.

[0030] As a further improvement of this technical solution, the evaluation index normalization formula in step S8 is:

[0031]

[0032] Where min i is the minimum possible value of the evaluation score when the i-th evaluation algorithm evaluates the model, and max i is the maximum possible value of the evaluation score when the corresponding evaluation algorithm evaluates the model.

[0033] Compared with the prior art, the beneficial effects of the present invention are:

[0034] 1. The present invention can perform a standardized evaluation process: by defining a benchmark scheme and dynamic attack parameters, it ensures that the evaluation results are comparable horizontally (across models) and vertically (across versions).

[0035] 2. The present invention can perform a comprehensive multi-dimensional index scoring: the weighted normalization method unifies the scoring ranges of different indexes, intuitively reflects the overall security of the model, and reduces the difficulty of user selection.

[0036] 3. The present invention can generate efficient adversarial samples: dynamically generate adversarial samples based on an iterative formula, avoid manual intervention, the attack intensity is adjustable and covers comprehensively, and the evaluation efficiency is increased by 40%.

[0037] 4. The present invention can perform flexible evaluation: supporting white-box and black-box evaluations, adapting to different scenario requirements (such as trade secret protection), and increasing the evaluation coverage rate by 50%.

[0038] 5. The present invention can enhance security and credibility: through targeted attacks on dynamic attack parameters and model structures, avoiding interference of the evaluation results by pre-strengthened models, and increasing the credibility by 30%. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a flowchart of the method for benchmark evaluation of the security of the artificial intelligence model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0041] In a specific embodiment, as Figure 1 shown, the present invention provides a method for benchmark evaluation of the security of an artificial intelligence model, which specifically includes the following steps:

[0042] The first step, judgment of the benchmark scheme.

[0043] According to the actual application scenario to which the artificial intelligence model to be evaluated belongs, such as scenarios of image classification, target recognition, autonomous driving, etc., it is judged whether there is already a benchmark evaluation scheme. If not, proceed to the second step, otherwise execute the sixth step.

[0044] The second step, establishment of the attack scheme.

[0045] The attack scheme defines the detailed attack process when attacking the benchmark model and the model to be evaluated, which includes the list of attack parameters used in the attack process and the specific parameter values of each attack parameter. In the present invention, the attack parameters include σ (learning rate), ε (upper limit of the perturbation amplitude), and T (maximum number of iterations). The larger the values of the above three attack parameters, the greater the attack intensity.

[0046] The third step, construction of the benchmark evaluation scheme.

[0047] The benchmark evaluation scheme includes the following information:

[0048] Benchmark model: Generally, an industry-recognized artificial intelligence model in this application scenario is selected as the benchmark model. For example, in the field of large models, currently, ChatGPT4 is generally selected as the benchmark model.

[0049] Benchmark dataset: The original dataset used for the security evaluation of the model, which will include the original samples before the attack and the correct result labels corresponding to the original samples.

[0050] Attack scheme: Associated with the attack schemes maintained in Step 1. Each benchmark scheme can be associated with multiple attack schemes, and each attack scheme represents one round of attack on the model to be evaluated. By combining different attack parameters and parameter values, it is possible to avoid certain models from performing targeted defense reinforcement against specific attack parameters.

[0051] Evaluation metrics: The evaluation metrics used for the security evaluation of the model. Each benchmark scheme can be associated with multiple evaluation metrics, so as to comprehensively evaluate the comprehensive security capabilities of the model to be evaluated.

[0052] Step 4: Dynamically generate adversarial samples.

[0053] According to the attack schemes associated with the benchmark schemes in Step 3, call the benchmark model and the benchmark dataset to generate adversarial samples. The specific generation process of the adversarial samples is as follows:

[0054] Traverse each sample x in the benchmark dataset to generate the corresponding adversarial sample x adv , where the generation formula of x adv is as follows:

[0055]

[0056] In the above formula: is the adversarial sample generated in the t-th iteration; is the adversarial sample generated in the (t + 1)-th iteration; θ is the model weight; σ is the step size (learning rate) of each iteration, controlling the amplitude of each update; sign is the sign function, used to obtain the direction of the gradient; y is the correct result label of the sample x; is the partial derivative of the loss function J with respect to ; Clip x,ε is the clipping function, ensuring that the generated adversarial sample is within the ε neighborhood of the original sample.

[0057] After iterating T times through the above formula for generating adversarial samples, a corresponding adversarial sample will be generated for one sample in the benchmark dataset. After traversing each sample x in the benchmark dataset, the adversarial sample set X adv will be obtained through the above formula.

[0058] Since each benchmark evaluation scheme is associated with multiple attack schemes, after replacing the attack parameters in the above formula with the attack parameter values in the corresponding attack schemes, the adversarial sample set

[0059] Step 5: Benchmark model evaluation.

[0060] Using the importlib tool in Python, dynamically load the weight file of the benchmark model, and then use each generated group of adversarial samples x adv as input, call the API interface of the benchmark model to obtain the output result y of the benchmark model adv . After that, based on the true results of the adversarial samples and the model output results, run the evaluation algorithm one by one, and calculate the security evaluation indicators by comparing the actual output results of the benchmark model with the expected correct results of the adversarial samples. For example, when using the ACC (accuracy) evaluation algorithm, the formula for calculating the score of the evaluation indicator is:

[0061]

[0062] where n represents the number of correctly classified samples among all adversarial samples, and N represents the total number of adversarial samples. The value range of this indicator is from 0 to 1, and the larger the value, the higher the security of the model.

[0063] When using the ALDP evaluation algorithm, the formula for calculating the score of the evaluation indicator is:

[0064]

[0065] where N is the total number of samples; x i is the i-th sample; is the likelihood of the model for sample x i ; is the likelihood of the true data distribution for sample x i . The value range of this indicator is from 0 to 1, and the larger the score of the evaluation indicator, the higher the security of the model.

[0066] Through the above steps, the actual score score of the benchmark model evaluated by the i-th evaluation indicator when using the j-th group of adversarial samples will be obtained i,j .

[0067] The scores of each evaluation indicator generated by the benchmark model in this step will be used as the control benchmark for subsequent benchmark evaluations of other algorithm models.

[0068] Step 6: Judgment of evaluation type.

[0069] The present invention supports two types of evaluation: black-box evaluation and white-box evaluation. During white-box evaluation, the network structure of the model to be evaluated can be analyzed for targeted attacks, but the source code file of the model to be evaluated needs to be provided in this process. For cases where the source code file of the model to be evaluated cannot be directly provided due to various business secrets or the large size of the model file, the present invention supports black-box evaluation of it. If white-box evaluation is to be performed, step seven is executed; if black-box evaluation is to be performed, step eight is executed.

[0070] Step Seven: Attack and calculation of the model to be evaluated.

[0071] 1. Generate adversarial samples for attack:

[0072] The overall process of this step is the same as that of step four, and the attack parameters and their specific values also come from the attack plan maintained in step two. However, the algorithm model loaded when generating adversarial samples at this time is the model to be evaluated, rather than the benchmark model, so that targeted attacks can be carried out according to the network structure of the model to be evaluated to ensure that the security issues of the model to be evaluated are discovered to the greatest extent.

[0073] 2. Calculate the evaluation metrics of the model to be evaluated:

[0074] If the evaluation type is white-box evaluation, the overall process of this step is the same as that of step five, but there are the following differences:

[0075] In this step, the weight file of the model to be evaluated is dynamically loaded through the importlib tool, rather than the weight file of the benchmark model;

[0076] The adversarial samples used in this step are the adversarial samples generated in step 1 for attacking and generating adversarial samples above, rather than the adversarial samples generated in step four.

[0077] If the evaluation type is black-box evaluation, the overall process of this step is the same as that of step five, but there are the following differences:

[0078] In this step, since the weight file of the model to be evaluated cannot be obtained, the evaluation metrics related to the weights of the model to be evaluated (such as the metric ALDp) will no longer be calculated, and only the evaluation metrics not requiring model weights (such as the metric ACC) will be calculated;

[0079] The adversarial samples used in this step are the adversarial samples generated in step four.

[0080] Through this step, the actual scores of the model to be evaluated using each evaluation metric when using each group of adversarial samples will also be obtained.

[0081] Step Eight: Calculate the comprehensive score.

[0082] According to the scores of each evaluation index generated by the benchmark model in step five and the scores of each evaluation index generated by the model to be evaluated in step seven, the benchmark evaluation results are generated through the following operations.

[0083] 1. Calculate the normalized scores of each evaluation index for each round of attack. The formula is as follows:

[0084]

[0085] where: min i is the minimum possible value of the evaluation score when the i-th evaluation algorithm evaluates the model; max i is the maximum possible value of the evaluation score when the corresponding evaluation algorithm evaluates the model; score i,j is the actual score of the i-th evaluation algorithm evaluating the model in the j-th round of attack.

[0086] 2. Calculate the weighted scores of the evaluation indexes in all rounds of attack. The formula is as follows:

[0087]

[0088] where: k j is the weight of each round of attack; Score i is the normalized score of the i-th evaluation algorithm.

[0089] For k j The specific calculation formula is as follows:

[0090]

[0091] where: param t is the parameter value of the attack parameter t in this attack scheme; pmax t is the maximum parameter value of the attack parameter t in this attack scheme; pmin t is the minimum parameter value of the attack parameter t in this attack scheme.

[0092] 3. Calculate the weighted score of the final evaluation index. The formula is as follows:

[0093]

[0094] where: p i is the weight of the corresponding evaluation algorithm.

[0095] Repeat the above steps to obtain the weighted score Score base of the final evaluation index of the benchmark model, and the weighted score Score of the final evaluation index of the model to be evaluated.

[0096] 4. Calculate the benchmark score. The formula is as follows:

[0097]

[0098] Where: Score base is the weighted score of the final evaluation metric of the benchmark model calculated through the above steps 1, 2, and 3, and Score is the weighted score of the final evaluation metric of the model under evaluation calculated through the above steps 1, 2, and 3.

[0099] Through the above operations, the security benchmark evaluation score SScore of the model under evaluation is obtained, and its value range is between 0 and 100. This score can intuitively reflect the security of the model under evaluation in its application scenario. For example, a score above 60 indicates that the security of the model under evaluation is qualified, and 100 points represents that the security of the model under evaluation reaches the security level of the benchmark model in this application scenario. At the same time, the model scoring has a theoretical basis for horizontal comparison with other models and vertical comparison with different versions of its own model.

[0100] In summary, compared with the prior art, an artificial intelligence model security benchmark evaluation method provided by the present invention realizes the benchmark evaluation of the security of the algorithm model. By giving a clear definition of the benchmark evaluation, quantitative analysis of the security of the model under evaluation is realized on a theoretical basis, making the evaluation results comparable horizontally with other models and vertically with different versions of its own model. Considering the different attack intensities and the differences of different evaluation metrics in the evaluation process, by giving a comprehensive scoring formula for multi-dimensional metrics, the scoring results are intuitively understandable numerically, greatly reducing the selection difficulty of the general public in the actual application of artificial intelligence models. Through the dynamic generation mechanism of adversarial samples, the manual time cost in the process of generating adversarial samples is greatly reduced, the possibility of the model under evaluation obtaining adversarial samples is avoided, and the security and credibility of the evaluation results are ensured.

[0101] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An artificial intelligence model security benchmark evaluation method, characterized in that, It includes the following steps: S1. Benchmark scheme judgment: Determine whether there is a benchmark evaluation scheme according to the application scenario to which the model to be evaluated belongs; If not, establish an attack scheme and a benchmark evaluation scheme; S2. Attack scheme establishment: Define a list of attack parameters, including the learning rate σ, the upper limit of perturbation amplitude ε, and the maximum number of iterations T. The larger the values of the attack parameters, the greater the attack intensity; S3. Benchmark evaluation scheme construction: It includes a benchmark model, a benchmark dataset, an associated attack scheme, and evaluation metrics. The benchmark model is a model recognized in the industry, and the benchmark dataset contains original samples and their correct labels; S4. Dynamically generate adversarial examples: Generate a set of adversarial examples \(x\) through an iterative formula adv , and the formula is: Among them, is the adversarial sample generated in the t-th iteration; is the adversarial sample generated in the (t + 1)-th iteration; θ is the model weight; σ is the step size (learning rate) of each iteration, controlling the magnitude of each update; sign is the sign function, used to obtain the direction of the gradient; y is the correct result label of the sample x; is the partial derivative of the loss function J with respect to ; Clip x,ε is the clipping function, ensuring that the generated adversarial sample is within the ε-neighborhood of the original sample; S5. Benchmark model evaluation: Load the weights of the benchmark model and calculate the scores score of each evaluation metric i,j , including accuracy (ACC) and average log-likelihood deviation (ALDP); S6. Evaluation type judgment: Support white-box evaluation and black-box evaluation. The source code of the model to be evaluated is required for white-box evaluation, and only adversarial samples are used for black-box evaluation; S7. Attack and calculation of the model to be evaluated: Generate adversarial samples according to the evaluation type and calculate the scores of the evaluation metrics; S8. Comprehensive score calculation: After normalizing the scores of the evaluation metrics, calculate the final benchmark score SScore through a weighted formula. The formula is: Among them, Score is the comprehensive score of the model to be evaluated, and Score base is the score of the baseline model.

2. The artificial intelligence model security benchmark evaluation method according to claim 1, wherein After iterating T times through the formula for generating adversarial samples in step S4, an adversarial sample corresponding to a sample in the benchmark dataset will be generated. After traversing each sample x in the benchmark dataset, an adversarial sample set will be obtained through the above formula; Since each benchmark evaluation scheme is associated with multiple attack schemes, after replacing the attack parameters in the formula for generating adversarial samples in step S4 with the numerical values of the attack parameters in the corresponding attack scheme, the adversarial sample set corresponding to attack scheme j can be obtained 3. The artificial intelligence model security benchmark evaluation method according to claim 1, wherein, When the ACC evaluation algorithm is adopted in step S5, the formula for calculating the score of the evaluation metric is: Where n represents the number of correctly classified samples among all adversarial samples, and N represents the total number of adversarial samples. The value range of this metric is from 0 to 1. The larger the value, the higher the security of the model.

4. The artificial intelligence model security benchmark evaluation method according to claim 1, characterized in that When the ALDP evaluation algorithm is adopted in step S5, the formula for calculating the score of the evaluation metric is: where N is the total number of samples, and x i is the i-th sample, is the log-likelihood of the model for the sample x i , is the log-likelihood of the true data distribution for the sample x i . The value range of this metric is from 0 to 1. The larger the evaluation metric score, the higher the security of the model.

5. The artificial intelligence model security benchmark evaluation method according to claim 1, characterized in that In the white-box evaluation of step S6, when generating adversarial samples, load the network structure of the model to be evaluated and perform targeted attacks according to its gradient.

6. The artificial intelligence model security benchmark evaluation method according to claim 1, wherein The formula for normalizing the evaluation metrics in step S8 is: where min i is the minimum possible value of the evaluation score when the i-th evaluation algorithm evaluates the model, and max i is the maximum possible value of the evaluation score when the corresponding evaluation algorithm evaluates the model.

Citation Information

Patent Citations

  • Robustness evaluation method and device for deep learning model and storage medium

    CN110222831A

  • Artificial intelligence model security automatic evaluation method oriented to general service scene

    CN118627059A

Cited By

  • Artificial intelligence model safety reference evaluation system based on interpretability

    CN122490531A