Large model safety performance evaluation method based on analytic hierarchy process and Shapley value fusion

Through the fusion method of hierarchical analysis and Shapley value, combined with multi-dimensional data and expert experience, the problem of subjective and objective data coordination in large-scale model security assessment is solved, and a comprehensive, reliable assessment and optimization governance of large-scale model security performance is achieved.

CN120337234AInactive Publication Date: 2025-07-18UNIV OF ELECTRONICS SCI & TECH OF CHINA

Patent Information

Application Number
CN202510807531.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing large-scale model security evaluation methods, subjective experience and objective data are difficult to effectively coordinate, resulting in the evaluation results being one-sided or separated from the actual scenario risk, and cannot fully and accurately reflect the security performance of the large-scale model.

Method used

The method of fusion of hierarchical analysis and Shapley value is adopted to collect multi-dimensional data, build a hierarchy, set up a comparison matrix, calculate AHP weights, and calculate marginal contributions in combination with Shapley values, blend subjective and objective weights to achieve interpretable and verifiable quantitative weight allocation for security performance evaluation.

Benefits of technology

It realizes the comprehensiveness and scientificity of the large-scale security performance evaluation, provides an interpretable quantitative weight allocation solution, optimizes security governance strategies, and improves the credibility and accuracy of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337234A_ABST
    Figure CN120337234A_ABST
Patent Text Reader

Abstract

The invention discloses a large model safety performance evaluation method based on an analytic hierarchy process and Shapley value fusion, and belongs to the technical field of artificial intelligence safety. Comprising the following steps: setting a large model safety performance evaluation index, collecting data and carrying out standardization processing; constructing a hierarchical structure of large model safety performance evaluation, and setting a comparison matrix; constructing a feature equation based on the comparison matrix, solving the feature equation, and normalizing feature vectors to obtain the AHP weight of each safety evaluation index; calculating the Shapley value weight of each safety evaluation index in the safety evaluation index vector; the AHP subjective weight and the Shapley objective weight are fused, a model safety score is calculated, and the model safety score is used for evaluating the model safety performance. According to the method, the hierarchical decision framework of the AHP is creatively combined with the cooperative game theory of the Shapley value, and through dynamic fusion of subjective and objective weights, an explainable and verifiable quantitative weight distribution scheme is provided for large model safety performance evaluation, and accurate optimization of a safety governance strategy is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence security, and relates to a security evaluation method for large models. Further, it can be used to quantify the weight distribution of performance indicators of large models in dimensions such as data security, algorithm robustness, and privacy protection, providing an interpretable quantitative basis for model security governance. In particular, it relates to a large model security performance evaluation method based on the fusion of the analytic hierarchy process and the Shapley value. Background Art

[0002] In today's digital age, generative large models have been extremely widely used in many fields due to their powerful capabilities. Machine learning, as the core technology of generative large models, supports large models to complete various complex tasks. However, the vulnerability characteristics of deep learning models have gradually made two typical attack methods, poisoning attacks and adversarial attacks, the main forms threatening the security of large models. Poisoning attacks contaminate the training data set and inject malicious samples during the model training stage, resulting in the model establishing incorrect feature associations; adversarial attacks target the deployed model and induce the model to produce outputs deviating from expectations by constructing adversarial samples that are indistinguishable to the human eye. The successful implementation of these two types of attacks may cause serious consequences such as model output deviation and decision-making system failure, and may even endanger life safety in high-risk scenarios such as autonomous driving and medical diagnosis.

[0003] In view of this, the security performance evaluation of large models has become the core link to ensure the credibility of large models. It is crucial to ensure that large models can operate stably, reliably, and securely in various complex application scenarios. However, the current mainstream evaluation methods have significant defects. The methods of assigning weights based on expert experience such as the analytic hierarchy process mainly rely on the subjective judgment and experience of experts in the evaluation process. Due to the differences in the knowledge backgrounds, cognitive levels, and understanding angles of different experts, this evaluation method is extremely vulnerable to the cognitive biases of evaluators. Eventually, the evaluation results are one-sided and cannot comprehensively and accurately reflect the security performance of large models. The evaluation methods relying on data-driven models such as the entropy weight method can keenly reflect the distribution characteristics of indicator data and mine potential information from the data level. However, in actual applications, large models often operate in specific scenario environments with various constraints. These data-driven evaluation methods often ignore these scenario constraints and simply assign weights based on the data itself. This leads to a disconnection between weight assignment and real risks, resulting in the evaluation results being unable to truly reflect the security risk status faced by large models in actual scenarios. Summary of the Invention

[0004] The object of the present invention is to provide a large model security performance evaluation method based on the fusion of the analytic hierarchy process and the Shapley value. Through the dynamic fusion of subjective and objective weights, it provides an interpretable and verifiable quantitative weight allocation scheme for the large model security performance evaluation, and realizes the precise optimization of security governance strategies. It is used to solve the technical problem that subjective experience and objective data are difficult to effectively cooperate in the large model security evaluation of the prior art.

[0005] To solve the above technical problems, the specific technical solutions of the present invention are as follows:

[0006] A large model security performance evaluation method based on the fusion of the analytic hierarchy process and the Shapley value, the method comprising the following steps:

[0007] Step S1: Set the large model security performance evaluation indicators, collect multi-dimensional data corresponding to each evaluation indicator and perform standardization processing. By introducing different functions, the value range of the data is converted to [0, 1], and at the same time, it is ensured that the larger the data, the better the large model security performance. The security performance of the large model is comprehensively evaluated from multiple dimensions; set the security evaluation index vector .

[0008] Set the security evaluation index vector It is expressed as follows:

[0009] F=[ F 1 ,…, F i ,..., F m ]

[0010] Among them, represents the th security evaluation index. The security evaluation index can be the attack success rate, the accuracy rate of answering sensitive questions, the model attack time consumption, the passing rate of toxicity detection, the robustness of adversarial samples, etc., is the total number of security evaluation indicators, .

[0011] Step S2: Through the analytic hierarchy process (AHP), construct the hierarchical structure of the large model security performance evaluation. The target layer is the comprehensive score of the large model security performance, comprehensively evaluating the security performance of the large model; the criterion layer is different evaluation indicators, such as the attack success rate, the accuracy rate of answering sensitive questions, the model attack time consumption, the passing rate of toxicity detection, the robustness of adversarial samples, etc.; the scheme layer is different large models to be evaluated. According to expert experience, the importance of each security evaluation indicator is compared pairwise. Set the comparison matrix :

[0012]

[0013] Among them, represents the safety evaluation index vector in the th safety evaluation index relative to the th safety evaluation index importance ratio, , satisfying and .

[0014] Step S3: Construct a characteristic equation based on the comparison matrix, solve the characteristic equation using the eigenvector method, and normalize the eigenvector to obtain the AHP weight of each safety evaluation index.

[0015] The characteristic equation is expressed as follows:

[0016]

[0017] Among them, is the eigenvector, and the eigenvector is solved through the characteristic equation, is the maximum eigenvalue.

[0018] After normalizing the eigenvector w, the AHP weight of each safety evaluation index is obtained, and the normalization operation is as follows:

[0019]

[0020] Among them, represents the AHP weight value of the th safety evaluation index , and the AHP weights correspond in order to the safety evaluation index vector F = [ F 1 ,…, F i ,..., F m ] in sequence.

[0021] To ensure that the comparison matrix satisfies logical consistency and to ensure the rationality and credibility of the weight calculation results, it is necessary to calculate the consistency index and the consistency ratio .

[0022]

[0023]

[0024] Among them, is the random consistency index. If the consistency of the comparison matrix is considered acceptable, otherwise, the comparison matrix needs to be readjusted. To readjust the comparison matrix, by analyzing the values of the elements in the comparison matrix, find the obviously unreasonable assignments of some elements, rethink the relative importance between the elements, and reasonably adjust the values of the elements in the comparison matrix. Find the index that has the greatest impact on CR, adjust the importance ratio of this index to other indices, and repeat step S3.

[0025] Step S4: Analyze the importance of the safety assessment indices through the Shapley value, and calculate the marginal contribution of the safety assessment indices to the final result.

[0026] The value function contributed by the safety index set is expressed as follows:

[0027]

[0028] where, , represents the safety indices in the set ;

[0029] For the th safety assessment index , define the Shapley value of the th safety assessment index:

[0030] φ i = ∑ S ⊆ N\{ F i } S !( N - S -1)! N ! [v(S ∪ { F i ) - v(S)]

[0031] where, represents the set containing all safety assessment indices, represents the number of members in the set N, represents the number of members in the set S, [v(S ∪ { F i ) - v(S)] represents the th safety assessment index 's marginal contribution, represents the weight of the marginal contribution of the th safety assessment index , represents the factorial.

[0032] After calculating the Shapley value of each safety assessment index, normalize it to obtain the Shapley value weights of each safety assessment index in the safety assessment index vector :

[0033]

[0034] Among them, represents the th safety assessment index of the Shapley value weight.

[0035] Step S5: Integrate the AHP subjective weight and the Shapley objective weight to obtain the comprehensive weight of each safety assessment index, and calculate the model safety score according to the comprehensive weight. The model safety score is used to evaluate the model safety performance.

[0036] The comprehensive weight is calculated in the following way:

[0037]

[0038] Among them, represents the comprehensive weight of the th safety assessment index, α ∈[0,1] is a regulation parameter. The regulation parameter α is dynamically adjusted within the value range of [0, 1]. By changing the regulation parameter α, the comprehensive weight of each safety assessment index is calculated multiple times. When the variance of all safety assessment indexes reaches the minimum value, the subjective and objective weight coefficients can be balanced, that is, when it is the smallest, the regulation parameter α is the one sought.

[0039] Finally, the comprehensive weight of each safety assessment index can be obtained , and through the weight vector and the safety assessment index vector , the safety score of the model can be calculated:

[0040]

[0041] Among them, represents the safety score of the model, represents the comprehensive weight vector, represents the transpose. According to the safety score of the model, the safety performances of different models can be compared.

[0042] The present invention fully considers multi-dimensional safety indexes, integrates subjective expert experience and objective Shapley values, and synthesizes subjective and objective weight coefficients, realizing the scientificity and credibility of the comprehensive evaluation of the safety performance of large models.

[0043] Compared with the prior art, the present invention has the following beneficial technical effects:

[0044] (1) The present invention realizes the comprehensiveness and scientificity of the safety performance evaluation of large models by collecting multi-dimensional index data.

[0045] (2) The present invention innovatively combines the hierarchical decision-making framework of AHP with the cooperative game theory of Shapley value. Through the dynamic fusion of subjective and objective weights, the comprehensive subjective and objective weight coefficients are obtained, realizing the credibility of the comprehensive evaluation of the security performance of large models, providing an interpretable and verifiable quantitative weight allocation scheme for the security performance evaluation of large models, and achieving the precise optimization of security governance strategies. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 It is the overall flowchart of the method for evaluating the weights of the security performance indicators of large models based on the fusion of the analytic hierarchy process and Shapley value of the present invention.

[0048] Figure 2 It is the structural schematic diagram of the calculation flowchart of the AHP analytic hierarchy process of the present invention. Detailed Embodiments

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0050] The following gives five security evaluation indicators as special cases to further describe the technical solutions of the present invention in detail.

[0051] As Figure 1 shown, the method includes the following steps:

[0052] Step S1: Set the following five evaluation indicators: attack success rate , accuracy rate of answering sensitive questions , attack time consumption of the model , passing rate of toxicity detection , robustness of adversarial samples . Collect data in these five dimensions and perform standardization processing on the data, and set the security evaluation index vector .

[0053] When standardizing data, different data is processed differently, as follows:

[0054] Due to the attack success rate and the passing rate of toxicity detection originally meaning the smaller the better, they need to be converted into positive indicators.

[0055] For and , the value ranges are both [0, 1], and it can be defined as:

[0056]

[0057]

[0058] Among them, represents the standardized attack success rate, the larger it is, the lower the attack success rate and the better the security; represents the standardized passing rate of toxicity detection, the larger it is, the lower the passing rate of toxicity detection and the better the security.

[0059] At the same time, for the attack time-consuming of the model, the value range is [0, ∞], and it can be defined as:

[0060]

[0061] Among them, represents the standardized attack time-consuming. When = 0, = 0; when -> ∞, -> 1. In this way, the value range can be compressed to [0, 1], the larger it is, the longer the attack time-consuming and the better the model security.

[0062] The accuracy rate of answering sensitive questions and the robustness of adversarial samples both have a value range of [0, 1], and the larger the value, the better the security, and they can be directly used.

[0063] The set security evaluation index vector is:

[0064] F= F 1 ,…, F i ,..., F 5 =[ f ASR ' , f SAA , f AT ' , f TDP ' , f AR ]

[0065] Among them, 。

[0066] Step S2: By the Analytic Hierarchy Process (AHP), pairwise comparisons are made on the importance of the five security evaluation indicators based on expert judgments, and a comparison matrix is set up, which is shown as follows:

[0067]

[0068] Among them, represents the security evaluation index vector in the th security evaluation index relative to the th and , i∈ 1,5 , j∈[1,5] 。

[0069] Step S3: Based on the comparison matrix, a characteristic equation is constructed, and the characteristic equation is solved using the eigenvector method. After normalizing the eigenvector, the AHP weight of each security evaluation index is obtained.

[0070] The characteristic equation is shown as follows:

[0071]

[0072] Among them, is the eigenvector, is the maximum eigenvalue.

[0073] After normalizing the eigenvector , the AHP weight of each security evaluation index can be obtained. The AHP weight is the AHP subjective weight. The AHP weight of the th security evaluation index is obtained through the following normalization operation:

[0074]

[0075] where the AHP weights are in order and the index vector F=[ f ASR ' , f SAA , f AT ' , f TDP ' , f AR ] Correspond in sequence.

[0076] To ensure that the comparison matrix meets logical consistency and to ensure the rationality and reliability of the weight calculation results, it is necessary to calculate the consistency index and the consistency ratio .

[0077]

[0078]

[0079] Among them is the random consistency index. If then the consistency of the comparison matrix is considered acceptable; otherwise, it is necessary to re-adjust the comparison matrix. To re-adjust the comparison matrix, by analyzing the values of the elements in the comparison matrix, find the obviously unreasonable assignments of some elements, re-consider the relative importance between the elements, and reasonably adjust the values of the elements in the comparison matrix. Find the index that has the greatest impact on CR, adjust the importance ratio of this index to other indexes, and repeat step S3, as Figure 2 shown.

[0080] Step S4: Conduct an importance analysis of the safety assessment indicators through the Shapley value, and calculate the marginal contribution of the safety assessment indicators to the final result.

[0081] Value function contributed by the safety index set :

[0082]

[0083] Among them , represents the indicators in the set .

[0084] For the th safety assessment indicator , define the Shapley value of the th safety assessment indicator:

[0085] φ i = ∑ S ⊆ N\{ F i } S !( N - S -1)! N ! [v(S ∪ { F i ) - v(S)]

[0086] Among them, represents the set containing these five safety assessment indicators, denotes the number of members in set N, denotes the number of members in set S, [v(S ∪ { F i ) - v(S)] denotes the th security assessment indicator 's marginal contribution, denotes the th security assessment indicator 's weight of the marginal contribution, denotes factorial.

[0087] After calculating the Shapley value of each security assessment indicator, normalize it to obtain the Shapley value weight of each security assessment indicator in the security assessment indicator vector . The Shapley value weight is the Shapley objective weight, and the calculation method is as follows:

[0088]

[0089] Finally, obtain the Shapley value weight of each security assessment indicator in the security assessment indicator vector . .

[0090] Step S5: Integrate the AHP subjective weight and the Shapley objective weight in a weighted average manner to obtain the final comprehensive weight , and calculate the model security score according to the comprehensive weight. The model security score is used to evaluate the model security performance.

[0091]

[0092] where α∈[0,1] is the adjustment parameter.

[0093] Finally, the comprehensive weight of each security assessment indicator can be obtained . Through the weight vector and the security assessment indicator vector , the security score of the model can be calculated as follows:

[0094]

[0095] where denotes the security score of the model, denotes the comprehensive weight vector, denotes transpose. According to the security score of the model, the security performance of different models can be compared.

[0096] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will be aware that, without departing from the spirit and scope of the present invention, various changes or equivalent substitutions can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the scope protected by the present invention.

Claims

1. A large model security performance evaluation method based on the fusion of the analytic hierarchy process and Shapley value, characterized in that The method includes the following steps: Step S1: Set the security performance evaluation indicators of the large model, collect multi-dimensional data corresponding to each evaluation indicator and perform standardization processing, and set the security evaluation indicator vector; Step S2: Construct the hierarchical structure of the large model security performance evaluation by the analytic hierarchy process (AHP), and set the comparison matrix; Step S3: Based on the comparison matrix, construct the characteristic equation, solve the characteristic equation using the eigenvector method, and obtain the AHP weight of each security evaluation indicator after normalizing the eigenvector; Step S4: Conduct importance analysis on the security evaluation indicators by the Shapley value, and calculate the Shapley value weight of each security evaluation indicator in the security evaluation indicator vector; Step S5: Integrate the AHP subjective weight and the Shapley objective weight to obtain the comprehensive weight of each security evaluation indicator, calculate the model security score according to the comprehensive weight, and the model security score is used to evaluate the model security performance.

2. The method for evaluating the security performance of a large model based on the fusion of the analytic hierarchy process and the Shapley value according to claim 1, wherein Safety assessment index vector is expressed as , where represents the th safety assessment index, represents the total number of safety assessment indexes; Set up a comparison matrix Indicates that: Among them, represents the th safety evaluation index relative to the th safety evaluation index , satisfying and .

3. The security performance evaluation method of the large model based on the fusion of the analytic hierarchy process and the Shapley value according to claim 2, characterized in that In step S3, The characteristic equation is expressed as follows: Among them, is the eigenvector, and the eigenvector is solved by the characteristic equation; After normalizing the eigenvector w, the AHP weight of each security evaluation indicator is obtained, and the normalization operation is as follows: Among them, represents the AHP weight value of the th safety evaluation index, and the AHP weights correspond to the safety evaluation index vector in sequence.

4. The security performance evaluation method of the large model based on the fusion of the analytic hierarchy process and the Shapley value according to claim 3, characterized in that, In step S4, Value function contributed by the set of safety indicators It is expressed as follows: Among them, , represents the safety index in the set; For the th safety assessment index , define the Shapley value of the th safety assessment index : Among them, represents the set containing all security assessment indicators, represents the number of members of set N, represents the number of members of set S, represents the th security assessment indicator 's marginal contribution, represents the weight of the marginal contribution of the th security assessment indicator ; represents the factorial; After calculating the Shapley values of each security evaluation index, normalize them to obtain the security evaluation index vector The Shapley value weights of each security evaluation index in Among them, represents the safety assessment index of the Shapley value weight.

5. The security performance evaluation method of the large model based on the fusion of the analytic hierarchy process and the Shapley value according to claim 4, wherein In step S5, The comprehensive weight is calculated in the following manner: Among them, represents the comprehensive weight of the th safety assessment index, represents the adjustment parameter; the adjustment parameter α is dynamically adjusted within the value range of [0, 1]. By changing the adjustment parameter α, the comprehensive weight of each safety assessment index is calculated multiple times. When the variance of all safety assessment indexes reaches the minimum value, the subjective and objective weight coefficients can be balanced, that is, when it is the smallest, the adjustment parameter α is the one sought; Finally, the comprehensive weight of each security evaluation index is obtained , through the weight vector and the security evaluation index vector , the security score of the model can be calculated as follows: Among them, represents the security performance of the model, represents the comprehensive weight vector, represents the transpose.

Citation Information

Patent Citations

  • Risk assessment method based on Shapley value and interaction index

    CN106447044A

  • Large model evaluation method, device, equipment, system and program product

    CN120106210A

  • Method for evaluating building roof photovoltaic power quality based on AHP and critic-entropy

    US20240183916A1

Cited By

  • Artificial general intelligent safety assessment method and system based on AHP and genetic algorithm

    CN121658898A