AI large model security evaluation method and device, electronic equipment and program product
By assigning the most suitable evaluation model to the safety evaluation of large AI models, and utilizing routing models and uncertainty calibration mechanisms, the problems of low evaluation efficiency, insufficient accuracy, and high cost in existing technologies are solved, achieving efficient, accurate, and economical evaluation results.
Patent Information
- Application Number
- CN202511336717.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing AI large-scale model security evaluation methods cannot balance evaluation efficiency, accuracy, and cost. In particular, in multi-model integration evaluation, there is a problem of linear growth in API call costs and computing resource consumption. Furthermore, fixed model combination strategies are difficult to adapt to the risk levels and data characteristics of different tasks.
By assigning the most suitable evaluation AI model to the response text to be tested for different risk dimensions, and using a routing model to dynamically select the most suitable evaluation AI model, combined with machine learning and uncertainty calibration mechanisms, the evaluation process is optimized to reduce costs and improve accuracy.
It achieves a balance between efficiency, accuracy, and cost in the evaluation process, dynamically matches risk dimensions with model capabilities, reduces evaluation costs, and improves evaluation accuracy and consistency.
Smart Images

Figure CN120832676A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an AI large model security evaluation method and device, electronic equipment and computer program product. BACKGROUND
[0002] The security evaluation of the AI large model refers to the security evaluation of the reply content output by the AI large model based on the risk problem. The processing flow mainly includes: inputting the risk problem text into the AI large model to be tested to obtain the reply text output by the AI large model to be tested; analyzing the reply text by experts manually or by an evaluation AI large model to output a judgment result of whether the reply text is a risk text; determining the risk pass rate of the AI large model to be tested by counting the number of risk texts contained in all reply texts, so as to measure the security degree of the AI large model to be tested.
[0003] At present, the security evaluation of the AI large model is mainly completed by experts manually or by an evaluation AI large model. Although the manual evaluation has high accuracy, the processing efficiency is low. Although the AI large model evaluation has high processing efficiency, the accuracy is insufficient. In view of this situation, the prior art proposes a multi-model integrated evaluation method, that is, the output results of multiple evaluation AI large models are combined to improve the accuracy of the security evaluation. This method usually adopts voting, weighted average or other aggregation strategies to comprehensively evaluate the scores or judgment results of different evaluation AI large models, so as to fully utilize the advantages of each evaluation AI large model and reduce the prediction deviation and errors of a single model.
[0004] However, since this method needs to call the API interfaces of multiple AI large models, the API calling cost and the computing resource consumption will linearly increase with the number of models and the amount of evaluation data, which finally leads to a substantial increase in the evaluation cost. As can be seen, the existing AI large model security evaluation method cannot balance the evaluation efficiency, the evaluation accuracy and the evaluation cost. SUMMARY
[0005] Therefore, the embodiments of the present application provide an AI large model security evaluation method, device, electronic equipment and computer program product, which can balance the evaluation efficiency, the evaluation accuracy and the evaluation cost by directionally assigning the most suitable evaluation AI large model to the to-be-tested reply text of different risk dimensions.
[0006] The first aspect of the embodiments of the present application provides an AI large model security evaluation method, comprising: obtaining a to-be-tested reply text output by an AI large model to be tested based on a risk problem text; from multiple evaluation AI large models, selecting a first evaluation AI large model adapted to a first risk dimension to which the to-be-tested reply text belongs; The first prediction risk label of the reply text to be tested is output by predicting the reply text to be tested through the first evaluation AI large model.
[0007] In the technical solution of the embodiment of the application, first, the reply text to be tested output by the AI large model to be tested based on the risk problem text is obtained, then, from a plurality of evaluation AI large models, a first evaluation AI large model that is adapted to the first risk dimension to which the reply text to be tested belongs is selected, and finally, the first prediction risk label of the reply text to be tested is output by predicting the reply text to be tested through the first evaluation AI large model. The above process uses the first evaluation AI large model to complete automatic evaluation, which can obtain higher evaluation efficiency. The above process can realize directional matching of risk dimensions and model capabilities by directionally assigning the most suitable evaluation AI large model to the reply text to be tested of different risk dimensions, so that the evaluation AI large model that performs better in a specific risk dimension is specially used to process the evaluation task of the corresponding risk dimension, thereby effectively improving the evaluation accuracy. The above process only needs to call a single evaluation AI large model, which can avoid the problem that the API calling cost and the computing resource consumption linearly increase with the number of models caused by the simultaneous running of multiple models, thereby greatly reducing the evaluation cost. Therefore, the technical solution of the embodiment of the application can balance the evaluation efficiency, the evaluation accuracy and the evaluation cost.
[0008] In one implementation manner of the embodiment of the application, selecting, from a plurality of evaluation AI large models, a first evaluation AI large model that is adapted to the first risk dimension to which the reply text to be tested belongs comprises: The reply text to be tested is input into the trained routing model for processing, and the routing model outputs the adaptation degree weight of each evaluation AI large model and the reply text to be tested; wherein the routing model is a machine learning model that respectively calculates the adaptation degree weight of each evaluation AI large model and the reply text to be tested based on historical evaluation data of each evaluation AI large model on the first risk dimension; From the plurality of evaluation AI large models, the evaluation AI large model with the highest adaptation degree weight of the reply text to be tested is selected as the first evaluation AI large model.
[0009] In one implementation manner of the embodiment of the application, the training data set of the routing model is constructed in the following manner: The sample text dataset, the expert label dataset, the model prediction result dataset, and the model prediction uncertainty score dataset are obtained, wherein the sample text dataset includes multiple sample reply texts covering multiple preset risk dimensions, the expert label dataset includes true risk labels of some of the sample reply texts, the model prediction result dataset includes predicted risk labels of each evaluation AI large model for each sample reply text, and the model prediction uncertainty score dataset includes uncertainty scores of each evaluation AI large model for each sample reply text. The multiple preset risk dimensions are divided into multiple different risk levels, different risk levels correspond to different risk coefficients, and each sample reply text is determined to correspond to a risk level according to a respective preset risk dimension. The interface calling cost and the expert labeling cost of each evaluation AI large model are respectively counted. The training dataset is constructed according to the sample text dataset, the expert label dataset, the model prediction result dataset, the model prediction uncertainty score dataset, and the interface calling cost and the expert labeling cost of each evaluation AI large model.
[0010] In an implementation manner of the embodiment of the application, the second risk dimension is any risk dimension in the multiple preset risk dimensions, the second evaluation AI large model is any evaluation AI large model in the multiple evaluation AI large models, and the target score range interval is any score range interval in the multiple preset score range intervals. Before the training dataset is constructed according to the sample text dataset, the expert label dataset, the model prediction result dataset, the model prediction uncertainty score dataset, and the interface calling cost and the expert labeling cost of each evaluation AI large model, the method further includes: The uncertainty score set of the second evaluation AI large model for predicting the sample reply text belonging to the second risk dimension is obtained. Each uncertainty score in the uncertainty score set is arranged in descending order and is respectively divided into the multiple preset score range intervals. In combination with the expert label dataset, the actual error rate of the second evaluation AI large model for predicting under the condition corresponding to the second risk dimension and the target score range interval is calculated. According to the actual error rate and the average value of each uncertainty score in the target score range interval, the score correction value of the target score range interval is determined. If the score correction value exceeds a preset tolerance threshold, each uncertainty score in the target score range interval is respectively corrected according to the score correction value, and the step of arranging each uncertainty score in the uncertainty score set in descending order and respectively dividing each uncertainty score into the multiple preset score range intervals is returned to be executed until the score correction value does not exceed the tolerance threshold.
[0011] In an implementation form of the embodiment of the application, the objective function of the routing model is used to minimize the total cost of labeling under the constraint of the overall error, the objective function introduces a risk coefficient as an error control coefficient, and the parameters optimized by the objective function include network weights and an uncertainty score threshold.
[0012] In an implementation form of the embodiment of the application, the training process of the routing model includes: randomly initializing network weights and initializing the uncertainty score threshold to the median of each uncertainty score contained in the model prediction uncertainty score data set; using an alternating optimization method, fixing the uncertainty score threshold, updating the network weights by the gradient descent method to minimize the objective function, and recalculating the uncertainty score threshold based on the updated network weights to constrain the overall error until the convergence condition is met, to obtain the final network weights and uncertainty score threshold.
[0013] In an implementation form of the embodiment of the application, after the first evaluation AI large model is used to predict the to-be-tested reply text and output the first predicted risk label of the to-be-tested reply text, the method further includes: determining a target uncertainty score of the prediction of the to-be-tested reply text by the first evaluation AI large model; if the target uncertainty score does not exceed the uncertainty score threshold, determining the first predicted risk label as the risk prediction result of the to-be-tested reply text; if the target uncertainty score exceeds the uncertainty score threshold, obtaining a second predicted risk label of the prediction of the to-be-tested reply text by an expert as the risk prediction result of the to-be-tested reply text.
[0014] The second aspect of the embodiment of the application provides an AI large model safety evaluation device, which includes: a reply text acquisition module configured to acquire a to-be-tested reply text output by a to-be-tested AI large model based on a risk problem text; an evaluation model selection module configured to select a first evaluation AI large model adapted to a first risk dimension to which the to-be-tested reply text belongs from a plurality of evaluation AI large models; a risk prediction module configured to predict the to-be-tested reply text by the first evaluation AI large model and output a first predicted risk label of the to-be-tested reply text.
[0015] The third aspect of the embodiment of the application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the AI large model safety evaluation method provided by the first aspect of the embodiment of the application when executing the computer program.
[0016] The fourth aspect of the embodiments of the present application provides a computer program product, which, when running on an electronic device, causes the electronic device to perform the AI large model security evaluation method provided by the first aspect of the embodiments of the present application.
[0017] The fifth aspect of the embodiments of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the AI large model security evaluation method provided by the first aspect of the embodiments of the present application.
[0018] It can be understood that the beneficial effects of the above-mentioned second aspect to fifth aspect can be referred to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart of an AI large model security evaluation method provided by the embodiments of the present application; Figure 2 is an operation principle schematic diagram of the AI large model security evaluation method provided by the embodiments of the present application in an actual application scenario; Figure 3 is a structural schematic diagram of an AI large model security evaluation device provided by the embodiments of the present application; Figure 4 is a schematic diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0020] In the following description, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, persons skilled in the art should understand that the embodiments of the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary details. In addition, in the description of the present application and the appended claims, the terms "first", "second", "third" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0021] The safety evaluation process of the AI large model is of great significance for the normative use of the AI large model. At present, the safety evaluation mainly includes two ways of being completed by experts manually and through an evaluation AI large model. The first way relies on experts to manually score the model output one by one. When facing large-scale data sets, manual scoring is inefficient and costly, and the demand for expert resources becomes a significant bottleneck. The second way uses a single evaluation AI large model to realize automatic scoring. Although it can effectively reduce costs, it has the problem of insufficient prediction accuracy, especially in high-risk scenarios, where misjudgment can lead to serious consequences. In addition to the above two ways, the existing technology also proposes a multi-model integrated evaluation method. By integrating the scores or judgments of different evaluation AI large models, the advantages of each evaluation AI large model can be fully utilized, and the prediction bias and errors of a single model can be reduced, thereby improving the accuracy of safety evaluation. However, this method needs to call the API interfaces of multiple AI large models, and the API calling fees and computing resource consumption will linearly increase with the number of models and the amount of evaluation data, ultimately leading to a substantial increase in evaluation costs. Moreover, when using multiple AI large models for safety evaluation, this method usually only uses simple voting or weighted averaging aggregation strategies, which cannot effectively eliminate the value bias of different models in the risk dimension, but may lead to systematic errors in the evaluation results in some risk dimensions, thereby affecting the consistency and fairness of the evaluation results. In addition, this method usually uses fixed model combinations and aggregation strategies, which are difficult to dynamically adjust according to the risk level or data characteristics of specific tasks. Therefore, the existing AI large model safety evaluation method cannot balance the evaluation efficiency, evaluation accuracy, and evaluation cost.
[0022] To solve the above problems, the embodiments of the present application propose an AI large model safety evaluation method, device, electronic equipment and computer program product, which can balance the evaluation efficiency, evaluation accuracy and evaluation cost by directionally assigning the most suitable evaluation AI large model for the to-be-tested reply text in different risk dimensions. For more specific technical implementation details of the embodiments of the present application, please refer to the various method embodiments described below.
[0023] It should be understood that the subject performing each method embodiment presented in the present application can be various types of electronic devices, such as a mobile phone, a tablet computer, a desktop computer, a wearable device, a medical device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a large-screen television, and the like, and the present embodiments do not limit the specific type of electronic device.
[0024] Referring to Figure 1 , a method for evaluating the safety of an AI large model is provided, which includes: 101. Obtain a to-be-tested reply text output by a to-be-tested AI large model based on a risk problem text; First, obtain a to-be-tested reply text output by a to-be-tested AI large model based on a risk problem text. The to-be-tested AI large model refers to any AI large model that needs to be evaluated for safety, the risk problem text refers to a problem description text that contains risk content, and the to-be-tested reply text refers to a reply text output by the to-be-tested AI large model based on the risk problem text to evaluate the risk. Each risk problem text has a corresponding to-be-tested reply text.
[0025] The source of the risk problem text can be a problem text set covering multiple preset risk dimensions. For example, in the context of generating an AI safety guide file, the safety evaluation needs to cover up to 31 safety indicators, each indicator corresponding to a risk dimension, and the problem text set needs to include multiple risk problem texts covering 31 risk dimensions, i.e., the number of risk problem texts for each risk dimension is at least one. By setting the above problem text set, the safety of the to-be-tested AI large model in 31 risk dimensions can be comprehensively evaluated.
[0026] 102. From a plurality of evaluation AI large models, select a first evaluation AI large model that is adapted to the first risk dimension to which the to-be-tested reply text belongs; After obtaining the to-be-tested reply text, since the risk problem text corresponding to the to-be-tested reply text is known, and the risk dimension corresponding to the risk problem text is also known, the risk dimension to which the to-be-tested reply text belongs can be determined, denoted as the first risk dimension.
[0027] In the technical solutions of the embodiments of the present application, a candidate evaluation model set is pre-configured, which includes a plurality of AI large models of different types of characteristics for performing safety evaluation, which can be referred to as evaluation AI large models. As an example, the candidate evaluation model set can be represented as , which represents k different evaluation AI large models, which are influenced by training data, algorithm design and optimization objectives during the training process, and may exhibit inherent value bias in certain risk dimensions, ultimately leading to different risk dimensions that different evaluation AI large models are good at handling. For example, a certain evaluation AI large model A may perform well in detecting politically sensitive content, while another evaluation AI large model B may be better at identifying false information or malicious content.
[0028] From the plurality of evaluation AI large models, an evaluation AI large model adapted to the first risk dimension to which the reply text to be evaluated belongs can be selected, denoted as a first evaluation AI large model, which is good at handling the evaluation task of the first risk dimension. In actual operation, the respective adapted risk dimensions of each evaluation AI large model can be determined and stored according to the processing results of each evaluation AI large model on historical evaluation tasks of different risk dimensions. For example, evaluation AI large model A is adapted to risk dimensions D1, D3 and D5, evaluation AI large model B is adapted to risk dimensions D1, D2 and D4, and so on. Each evaluation AI large model can be adapted to one or more risk dimensions, and different evaluation AI large models can be adapted to the same or different risk dimensions. If the number of evaluation AI large models adapted to the first risk dimension is more than one, one of them can be selected as the first evaluation AI large model in a set manner (e.g., randomly or in order). By directing the allocation of the most suitable evaluation AI large model to the reply text to be evaluated of different risk dimensions, the directional matching of risk dimensions and model capabilities can be achieved, allowing evaluation AI large models that perform better in specific risk dimensions to handle evaluation tasks of corresponding risk dimensions, thereby effectively improving the evaluation accuracy.
[0029] In an implementation manner of an embodiment of the present application, from the plurality of evaluation AI large models, an evaluation AI large model adapted to the first risk dimension to which the reply text to be evaluated belongs is selected, including: (1) inputting the reply text to be evaluated into a trained routing model for processing, and outputting the adaptation degree weight of each evaluation AI large model and the reply text to be evaluated through the routing model; wherein the routing model is a machine learning model that calculates the adaptation degree weight of each evaluation AI large model and the reply text to be evaluated based on historical evaluation data of each evaluation AI large model on the first risk dimension; (2) From the plurality of evaluation AI large models, select an evaluation AI large model with the highest adaptation weight of the to-be-tested reply text as the first evaluation AI large model.
[0030] In order to more accurately select the first evaluation AI large model adapted to the current evaluation task from the plurality of evaluation AI large models, the embodiment of the application introduces a risk-weighted multi-model dynamic routing mechanism. The core of the mechanism is to train a routing model, assign the most suitable evaluation AI large model for each to-be-tested reply text through intelligent routing decision, and significantly optimize the overall evaluation cost under the premise of meeting the prediction accuracy requirement. The routing model is a machine learning model based on historical evaluation data of each evaluation AI large model on the first risk dimension, which respectively calculates the adaptation weight of each evaluation AI large model and the to-be-tested reply text. After inputting the to-be-tested reply text into the routing model for processing, the adaptation weight of each evaluation AI large model and the to-be-tested reply text can be output. Finally, one of the evaluation AI large models with the highest adaptation weight of the to-be-tested reply text can be selected from the plurality of evaluation AI large models as the first evaluation AI large model. Next, the specific training process of the routing model is described.
[0031] In an implementation manner of the embodiment of the application, the training data set of the routing model is constructed by the following method: (1) Obtain a sample text data set, an expert label data set, a model prediction result data set, and a model prediction uncertainty score data set; wherein the sample text data set includes a plurality of sample reply texts covering a plurality of preset risk dimensions, the expert label data set includes real risk labels of a part of the sample reply texts, the model prediction result data set includes predicted risk labels of each evaluation AI large model for each sample reply text, and the model prediction uncertainty score data set includes uncertainty scores of each evaluation AI large model for each sample reply text. The plurality of preset risk dimensions are divided into a plurality of different risk levels, different risk levels correspond to different risk coefficients, and each sample reply text determines its corresponding risk level according to its own preset risk dimension; (2) Statistically count the interface call cost and the expert labeling cost of each evaluation AI large model; (3) Construct the training data set according to the sample text data set, the expert label data set, the model prediction result data set, the model prediction uncertainty score data set, and the interface call cost and the expert labeling cost of each evaluation AI large model.
[0032] The training of a machine learning model cannot be accomplished without a high-quality training dataset. In order to construct a training dataset for a routing model, a sample text dataset, an expert label dataset, a model prediction result dataset, and a model prediction uncertainty score dataset are first obtained. The sample text dataset includes a plurality of sample reply texts covering a plurality of preset risk dimensions output by an AI large model to be tested based on a risk problem dataset. For example, the sample text dataset can be represented as , which contains n sample reply texts, n The 31 risk dimensions are covered by the sample reply texts, and each risk dimension has at least one sample reply text, i.e. n ≥ 31, and in a normal case n the number of sample reply texts can reach several thousand or even tens of thousands.
[0033] The expert label dataset includes true risk labels of a portion of the sample reply texts, which are obtained by expert manual labeling and are considered to have 100% accuracy. The label dataset can be represented as , which contains m true risk labels of the sample reply texts, m ≤ n The true risk labels are used to calibrate the error of the routing model, and the label values are 0 or 1, representing no risk or risk, respectively. A small portion of the sample reply texts have expert-labeled true risk labels, which can be used to train the routing model. Subsequently, the trained routing model is used to judge the reply texts that have not been labeled by experts, to determine which reply texts need to be labeled by experts (high prediction cost but high prediction accuracy) and which reply texts only need to be labeled by an evaluation AI large model (lower prediction cost and prediction accuracy than experts), thereby achieving the effect of saving prediction cost while ensuring prediction accuracy.
[0034] The model prediction result dataset includes the predicted risk labels of each evaluation AI large model for each sample reply text. For example, for the candidate evaluation model set , the model prediction result dataset can be represented as , is the predicted risk label of the evaluation AI large model for the sample reply text , and the label value is 0 or 1.
[0035] The model prediction uncertainty score dataset includes the uncertainty scores of each evaluation AI large model for each sample reply text, which can be represented as , is the uncertainty score of the evaluation AI large model for the sample reply text The uncertainty score of the prediction, which ranges from 0 to 1. Each evaluation AI large model outputs an uncertainty score corresponding to the prediction result when it predicts the risk label of a sample reply text. The closer the uncertainty score is to 0, the more accurate the model's prediction result is, otherwise, the less accurate the model's prediction result is.
[0036] The plurality of preset risk dimensions are divided into a plurality of different risk levels, different risk levels correspond to different risk coefficients, and the higher the risk level, the greater the corresponding risk coefficient. As an example, the 31 risk dimensions can be divided into a set of high-risk levels K H , a set of medium-risk levels K M , and a set of low-risk levels K L , let the risk coefficient corresponding to the high-risk level be , the risk coefficient corresponding to the medium-risk level be , and the risk coefficient corresponding to the low-risk level be . For example, the risk dimensions "violation of core values" and "contain discriminatory content" can be divided into the set of high-risk levels K H , the risk dimension "commercial illegal violation" can be divided into the set of medium-risk levels K M , and the risk dimensions "infringement of others' legal rights and interests" and "unable to meet the requirements of a specific service type" can be divided into the set of low-risk levels K L .
[0037] In the data preprocessing stage, each sample reply text can determine its corresponding risk level according to the respective preset risk dimension it belongs to, so as to group n sample reply texts according to risk levels, which can be represented as By dividing each risk dimension into the corresponding risk level, the routing model can not only select the appropriate evaluation AI large model according to the risk dimension, but also select the appropriate evaluation AI large model according to the risk level.
[0038] The interface call cost and expert labeling cost of each evaluation AI large model for each sample to be tested are counted respectively. The interface call cost of a certain evaluation AI large model can be represented as , and the expert labeling cost can be represented as The interface calling cost can be determined by querying the official API price of the evaluation AI large model. Generally, the official API document will give the price per million tokens, so the token number of the average test sample can be counted to convert the price of using the corresponding large model API per test sample as the interface calling cost. The expert labeling cost can use a preset cost value, for example, 0.6 yuan per test sample.
[0039] According to the sample text data set, the expert label data set, the model prediction result data set, the model prediction uncertainty score data set, and the interface calling cost and the expert labeling cost of each evaluation AI large model, a training data set can be constructed. Subsequently, using the basic training method of the machine learning model, the neural network is trained based on the training data set, and the above routing model can be obtained.
[0040] In the model prediction uncertainty score data set described above, the uncertainty scores of each evaluation AI large model for predicting each sample reply text may have the problem of "confidence mismatching with actual error", for example, overconfidence or lack of confidence. To solve this problem, the embodiment of the present application introduces a risk-dimension uncertainty calibration mechanism, which can calibrate the model prediction uncertainty score data set according to the risk level. The specific processing process of the uncertainty calibration mechanism is described below.
[0041] In an implementation manner of the embodiment of the present application, let the second risk dimension be any risk dimension in the plurality of preset risk dimensions, let the second evaluation AI large model be any evaluation AI large model in the plurality of evaluation AI large models, and let the target score range interval be any score range interval in the plurality of preset score range intervals; before constructing the training data set according to the sample text data set, the expert label data set, the model prediction result data set, the model prediction uncertainty score data set, and the interface calling cost and the expert labeling cost of each evaluation AI large model, the method further includes: (1) obtaining a set of uncertainty scores of the second evaluation AI large model for predicting the sample reply text belonging to the second risk dimension; (2) arranging each uncertainty score in the uncertainty score set from small to large, and dividing them into the plurality of preset score range intervals; (3) combining the expert label data set, calculating the actual error rate of the second evaluation AI large model for predicting under the condition corresponding to the second risk dimension and the target score range interval; (4) determining a score correction value of the target score range interval according to the actual error rate and the average value of each uncertainty score in the target score range interval; (5) If the score correction value exceeds the preset tolerance threshold, each uncertainty score in the target score range is corrected according to the score correction value, and the process returns to execute the step of arranging each uncertainty score in the uncertainty score set from small to large and dividing them into multiple preset score ranges, until the score correction value does not exceed the tolerance threshold.
[0042] When executing the uncertainty calibration mechanism, group calibration is performed, and the uncertainty score data of each risk dimension and each evaluation AI model is processed separately. Specifically, assuming that the second risk dimension For any of the multiple pre-set risk dimensions, the second evaluation AI model j Target score range for any of the multiple AI evaluation models For any score range in multiple preset score ranges, first obtain the second evaluation AI large model j For the second risk dimension Then, the uncertainty scores in the uncertainty score set are arranged from small to large and divided into multiple preset score ranges. The lengths of different score ranges can be the same or different, and there is no intersection between different score ranges. For example, 10 score ranges of equal length can be divided and recorded as , target score range Any one of the 10 score ranges.
[0043] Next, we combine the expert label dataset described above to calculate the second evaluation AI model j In the second risk dimension and target score range In the case of , the actual error rate of the prediction is mainly used to statistically evaluate the percentage of inconsistencies between the predicted labels given by the AI large model and the expert labels. The specific calculation formula is as follows:
[0044] in, Indicates the second evaluation AI large model j In the second risk dimension and target score range The actual error rate of prediction under the condition of , is the risk loss function.
[0045] Then, based on the actual error rate and the target score range The average of the uncertainty scores in the target score range is determined The score correction value is calculated as follows:
[0046] in, Indicates the target score range The score correction value of Indicates the target score range The average value of each uncertainty score in , that is, the score correction value is equal to the difference between the actual error rate and the average value of the uncertainty score, and its value can be positive or negative.
[0047] Set a tolerance threshold in advance (For example, it can be 0.05), determine the score correction value calculated above Whether the tolerance threshold is exceeded If it exceeds, the value will be adjusted according to the score , for the target score range Each uncertainty fraction in is corrected separately, that is, for the sample , update its uncertainty score When making corrections, it is necessary to ensure that the corrected uncertainty scores are within the value range [0, 1]. If they are out of range, they are truncated. For example, if the corrected uncertainty score is less than 0, it is taken as 0, and if the corrected uncertainty score exceeds 1, it is taken as 1. After calculating the uncertainty scores in the uncertainty score set, the process returns to the step of arranging the uncertainty scores in the uncertainty score set from small to large and dividing them into multiple preset score ranges, and then enters the next round of convergence judgment process, and repeats this process until the calculated score correction value is ≤ When , the target score range is output Individual uncertainty scores after calibration.
[0048] It is understood that each score range can be divided into the target score range. The uncertainty score is calibrated in the same way to obtain the calibrated second evaluation AI model. j The second risk dimension The uncertainty score of the sample response text is predicted. Each risk dimension can be divided into the second risk dimension The uncertainty score is calibrated in the same way to obtain the calibrated second evaluation AI model. jThe uncertainty score of the sample response text for each risk dimension is predicted. Each evaluation AI model can be compared with the second evaluation AI model. j The uncertainty scores are calibrated in the same manner, resulting in the uncertainty scores predicted by each calibrated AI model for each sample response text for each risk dimension. This completes the calibration of the entire model prediction uncertainty score dataset. Considering that different AI models may have different value biases for different risk dimensions, this uncertainty calibration mechanism calibrates each risk dimension separately, resulting in better calibration results than a unified calibration across all risk dimensions.
[0049] After completing the above uncertainty score calibration, the training process of the routing model begins. The routing model can dynamically calculate and allocate the optimal evaluation AI model based on the characteristics of the risk dimension of the input response text to be tested.
[0050] In one implementation of an embodiment of the present application, the objective function of the routing model is used to minimize the total labeling cost while constraining the overall error. The objective function introduces a risk coefficient as an error control coefficient, and the parameters optimized by the objective function include network weight and uncertainty score threshold.
[0051] The routing model can use lightweight neural networks when training , whose input is the sample reply text The feature vector of k dimensional vectors, representing k Evaluation of AI large models and The objective function is defined to minimize the total annotation cost while constraining the overall error. The total annotation cost includes the interface call cost and expert annotation cost described above. This objective function introduces the risk factor described above as an error control factor to strengthen error control of key items. The parameters optimized during model training include network weights and uncertainty score thresholds. As an example, an objective function is defined as follows:
[0052] in, represents the network weight, represents the uncertainty score threshold, Represents a large AI model for evaluation j and The fitness weight, represents an estimate of the uncertainty score threshold, express sigmoid function, is the error control coefficient, is the risk loss function, and the constraint condition is that the overall error ≤ .
[0053] The fitness weight is obtained by softmax Function calculation, its calculation expression is as follows:
[0054] in, Indicates evaluation of AI large models j For samples The fitness score of each evaluation AI model is randomly initialized, and normalization is used to ensure that the sum of the fitness weights of each evaluation AI model is 1.
[0055] In the above objective function, The risk factor described above 、 or , by introducing As an error control coefficient, the routing model can select an appropriate evaluation AI model according to the risk level of the input reply text. For example, for a reply text with a high risk level, its risk coefficient is Larger, that is, to increase the weight of the prediction result risk in the objective function, so that the evaluation AI model with higher prediction accuracy will be selected first to ensure the accuracy of the evaluation results. For example, for the reply text with low risk level, its risk coefficient is Smaller, that is, reducing the weight of the prediction result risk in the objective function. In this way, the evaluation AI large model with lower interface call cost and expert annotation cost will be given priority to reduce the evaluation cost and improve resource utilization efficiency.
[0056] In one implementation of the embodiment of the present application, the training process of the routing model includes: (1) Randomly initialize the network weights and initialize the uncertainty score threshold to the median of the uncertainty scores contained in the model prediction uncertainty score dataset; (2) Using an alternating optimization approach, the uncertainty score threshold is fixed, the network weights are updated by gradient descent to minimize the objective function, and the uncertainty score threshold is recalculated based on the updated network weights to constrain the overall error until the convergence conditions are met, and the final network weights and uncertainty score threshold are obtained.
[0057] During the training of the routing model, the network weights are first initialized and uncertainty score threshold , network weight Using random initialization, the uncertainty score threshold initialized as the median of all uncertainty scores contained in the model prediction uncertainty score dataset, i.e., initialized as Then, in an alternating optimization manner, the uncertainty score threshold is fixed, the network weight is updated by gradient descent method to minimize the objective function, and then the uncertainty score threshold is recalculated according to the updated network weight to constrain the overall error. It is determined whether the network weight and the uncertainty score threshold satisfy the convergence condition. If the convergence condition is not satisfied, the uncertainty score threshold is fixed, and the network weight is updated by gradient descent method to minimize the objective function until the convergence condition is satisfied, and the final network weight and the uncertainty score threshold are obtained.
[0058] The trained routing model and uncertainty score threshold can be applied to a new AI large model safety evaluation task. In specific operation, the feature vector of the to-be-tested reply text is extracted, and then the feature vector is input into the routing model. The routing model calculates the adaptation weight of each evaluation AI large model to the to-be-tested reply text, and selects the evaluation AI large model with the highest adaptation weight as the optimal model . The related calculation formula is , and the optimal model is the first evaluation AI large model described above.
[0059] 103. The to-be-tested reply text is predicted by the first evaluation AI large model, and the first prediction risk label of the to-be-tested reply text is output.
[0060] After the routing model is used to select the first evaluation AI large model most suitable for the current evaluation task, the to-be-tested reply text is input into the first evaluation AI large model, and the to-be-tested reply text is predicted by the first evaluation AI large model to output the prediction risk label of the to-be-tested reply text, denoted as the first prediction risk label , , the label value of which is 0 or 1, indicating that the to-be-tested reply text has no risk / presence of risk.
[0061] In an implementation manner of the embodiment of the present application, after the to-be-tested reply text is predicted by the first evaluation AI large model to output the first prediction risk label of the to-be-tested reply text, the method further includes: (1) determining a target uncertainty score of the first evaluation AI large model for predicting the to-be-tested reply text; (2) if the target uncertainty score does not exceed the uncertainty score threshold, determining the first prediction risk label as the risk prediction result of the to-be-tested reply text; (3) if the target uncertainty score exceeds the uncertainty score threshold, obtaining a second prediction risk label of the expert for predicting the to-be-tested reply text as the risk prediction result of the to-be-tested reply text.
[0062] In order to further improve the accuracy of risk prediction, the embodiment of the application introduces the judgment mechanism of the uncertainty score threshold. First, the target uncertainty score of the first evaluation AI large model for predicting the to-be-tested reply text is determined , and it is judged whether it exceeds the final uncertainty score threshold ; if the uncertainty score threshold is not exceeded, it indicates that the reliability of the prediction result of the first evaluation AI large model is higher, and therefore the first prediction risk label is determined as the risk prediction result of the to-be-tested reply text; if the uncertainty score threshold is exceeded, it indicates that the reliability of the prediction result of the first evaluation AI large model is insufficient, at which time the expert manually predicts the to-be-tested reply text, obtains the second prediction risk label provided by the expert , and determines the second prediction risk label as the risk prediction result of the to-be-tested reply text. The corresponding processing process is shown in the following formula:
[0063] wherein, represents the risk prediction result of the to-be-tested reply text.
[0064] It can be understood that each to-be-tested reply text can obtain a corresponding risk prediction result label in the same way as described above. Assuming that the number of to-be-tested reply texts output by the to-be-tested AI large model based on the risk problem text is n , a label data set composed of n risk prediction result labels can be obtained , and the label data set satisfies the specified error rate and confidence.
[0065] Using the above label data set , the number of risk texts contained in all to-be-tested reply texts can be counted, so as to evaluate the risk pass rate of the to-be-tested AI large model, measure the safety degree of the to-be-tested AI large model, and finally complete the safety evaluation process of the to-be-tested AI large model.
[0066] In the technical solution of the embodiment of the application, first, the to-be-tested reply text output by the to-be-tested AI large model based on the risk problem text is acquired, then a first evaluation AI large model adapted to the first risk dimension to which the to-be-tested reply text belongs is selected from a plurality of evaluation AI large models, and finally the to-be-tested reply text is predicted by the first evaluation AI large model to output a first predicted risk label of the to-be-tested reply text. The above process uses the first evaluation AI large model to complete automatic evaluation, which can obtain higher evaluation efficiency; the above process can realize directional matching of risk dimensions and model capabilities by directionally assigning the most suitable evaluation AI large model to the to-be-tested reply text of different risk dimensions, so that the evaluation AI large model that performs better in a specific risk dimension is used to process the evaluation task of the corresponding risk dimension, thereby effectively improving the evaluation accuracy; the above process only needs to call a single evaluation AI large model, which can avoid the problem that the API calling cost and the computing resource consumption linearly increase with the number of models caused by the simultaneous running of multiple models, thereby greatly reducing the evaluation cost. Therefore, the technical solution of the embodiment of the application can balance the evaluation efficiency, the evaluation accuracy and the evaluation cost.
[0067] As an example, Figure 2 is an operational principle schematic diagram of the AI large model security evaluation method provided by the embodiment of the application in an actual application scenario. Figure 2 The core of is to build a risk-weighted multi-model dynamic routing system. The system assigns appropriate evaluation AI large models or artificial experts to each to-be-evaluated reply text through intelligent routing decision, significantly optimizes the overall evaluation cost under the premise of meeting the preset accuracy requirement. In Figure 2 , first, enter the routing model training phase: acquire the required sample data sets for preprocessing, calibrate the uncertainty scores of the sample data sets according to the uncertainty calibration mechanism described above, and construct the training data set of the routing model; in the training process of the routing model, initialize the network weights and the uncertainty score threshold , use the alternating optimization method, fix the uncertainty score threshold first, update the network weights by the gradient descent method to minimize the objective function, and then recalculate the uncertainty score threshold according to the updated network weights The whole error is constrained, and the above steps are repeated until a convergence condition is met. After the routing model is trained, the model deployment application stage is entered: after the to-be-tested reply text output by the to-be-tested AI large model is input into the routing model, the routing model calculates the adaptation weight of each evaluation AI large model and the to-be-tested reply text, and selects the evaluation AI large model with the highest adaptation weight as the optimal model; the optimal model is used to predict the to-be-tested reply text, if the uncertainty score of the prediction does not exceed the uncertainty score threshold, the prediction risk label output by the optimal model is retained as the final result label; if the uncertainty score of the prediction exceeds the uncertainty score threshold, the expert prediction label is called as the final result label. Through the risk-weighted multi-model dynamic routing mechanism, the above process realizes the safety evaluation of multiple models in cooperation, can significantly improve the evaluation accuracy, optimizes the use of expert resources, greatly saves the cost of manual annotation, and ensures that the evaluation result meets the error rate and confidence requirements specified by the user.
[0068] In summary, the AI large model safety evaluation method provided by the embodiments of the present application has the following advantages: (1) by introducing a dynamic routing mechanism, the most suitable evaluation model is assigned for evaluation tasks of different risk dimensions, only a single optimal model or a small number of human experts are selected for evaluation, and multiple evaluation models are not called at the same time, which can greatly reduce the evaluation cost; (2) abandoning simple voting or weighted average aggregation strategies, the evaluation is divided by directional matching of risk dimensions and model capabilities, allowing evaluation models that perform better in specific risk dimensions to handle evaluation tasks of corresponding risk dimensions, reducing conflicts caused by inherent value bias of models from the source, avoiding interference between different model biases when aggregating, and thus ensuring that the evaluation results of each risk dimension are more consistent and fair; (3) having flexible dynamic adjustment capability, the model selection and resource allocation strategy can be dynamically adjusted according to the risk level and data characteristics of the evaluation task, for high-risk evaluation tasks, evaluation models with high reliability and low uncertainty are preferentially selected, and the uncertainty score threshold is reduced to ensure prediction accuracy; for low-risk evaluation tasks, low-cost evaluation models are preferentially selected to improve resource utilization efficiency; this dynamic adaptation mechanism based on risk level can solve the contradiction between resource waste and insufficient prediction accuracy of fixed model combinations and aggregation strategies.
[0069] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0070] The above mainly describes an AI large model safety evaluation method, and a AI large model safety evaluation device will be described below.
[0071] Please refer to Figure 3, and shows an AI large model security evaluation device provided in the embodiment of the application, comprising: The reply text acquisition module 301 is configured to acquire a to-be-tested reply text output by the to-be-tested AI large model based on the risk problem text; The evaluation model selection module 302 is configured to select, from a plurality of evaluation AI large models, a first evaluation AI large model that is adapted to a first risk dimension to which the to-be-tested reply text belongs. The risk prediction module 303 is configured to perform prediction on the to-be-tested reply text by the first evaluation AI large model, and output a first predicted risk label of the to-be-tested reply text.
[0072] In an implementation manner of the embodiment of the application, the evaluation model selection module comprises: The routing model processing unit is configured to input the to-be-tested reply text into a trained routing model for processing, and output, by the routing model, an adaptation degree weight of each evaluation AI large model and the to-be-tested reply text; wherein the routing model is a machine learning model that respectively calculates the adaptation degree weight of each evaluation AI large model and the to-be-tested reply text based on historical evaluation data of each evaluation AI large model on the first risk dimension. The evaluation model selection unit is configured to select, from the plurality of evaluation AI large models, an evaluation AI large model with the highest adaptation degree weight of the to-be-tested reply text as the first evaluation AI large model.
[0073] In an implementation manner of the embodiment of the application, the AI large model security evaluation device further comprises: The data set acquisition module is configured to acquire a sample text data set, an expert label data set, a model prediction result data set, and a model prediction uncertainty score data set; wherein the sample text data set comprises a plurality of sample reply texts covering a plurality of preset risk dimensions, the expert label data set comprises real risk labels of a part of the sample reply texts, the model prediction result data set comprises predicted risk labels of each sample reply text by each evaluation AI large model, and the model prediction uncertainty score data set comprises an uncertainty score of each sample reply text by each evaluation AI large model; the plurality of preset risk dimensions are divided into a plurality of different risk levels, different risk levels correspond to different risk coefficients, and each sample reply text determines a corresponding risk level according to a respective preset risk dimension. The cost statistics module is configured to respectively count an interface call cost and an expert labeling cost of each evaluation AI large model. The data set construction module is configured to construct a training data set according to the sample text data set, the expert label data set, the model prediction result data set, the model prediction uncertainty score data set, and the interface call cost and the expert labeling cost of each evaluation AI large model.
[0074] In an implementation form of the embodiment of the application, the second risk dimension is any risk dimension in the plurality of preset risk dimensions, the second evaluation AI large model is any evaluation AI large model in the plurality of evaluation AI large models, and the target score range interval is any score range interval in the plurality of preset score range intervals; the AI large model security evaluation apparatus further comprises: an uncertainty score set obtaining module configured to obtain a set of uncertainty scores of the second evaluation AI large model in predicting the sample reply texts belonging to the second risk dimension; an uncertainty score arrangement module configured to arrange each uncertainty score in the set of uncertainty scores from small to large and divide the uncertainty scores into the plurality of preset score range intervals; an actual error rate calculating module configured to calculate, in combination with the expert label data set, an actual error rate of the second evaluation AI large model in predicting under the condition of the second risk dimension and the target score range interval; a score correction value determining module configured to determine a score correction value of the target score range interval according to the actual error rate and an average of each uncertainty score in the target score range interval; an uncertainty score correcting module configured to, if the score correction value exceeds a preset tolerance threshold, correct each uncertainty score in the target score range interval according to the score correction value, and return to perform the step of arranging each uncertainty score in the set of uncertainty scores from small to large and dividing the uncertainty scores into the plurality of preset score range intervals until the score correction value does not exceed the tolerance threshold.
[0075] In an implementation form of the embodiment of the application, the objective function of the routing model is used to minimize the total labeling cost under the constraint of the overall error, the objective function introduces a risk coefficient as an error control coefficient, and the parameters optimized by the objective function include network weights and an uncertainty score threshold.
[0076] In an implementation form of the embodiment of the application, the AI large model security evaluation apparatus further comprises: a parameter randomization module configured to randomly initialize the network weights and initialize the uncertainty score threshold to the median of each uncertainty score contained in the model prediction uncertainty score data set; a parameter optimization module configured to fix the uncertainty score threshold, update the network weights by the gradient descent method to minimize the objective function in an alternating optimization manner, and re-calculate the uncertainty score threshold according to the updated network weights to constrain the overall error until a convergence condition is met, to obtain the final network weights and uncertainty score threshold.
[0077] In an implementation form of the AI large model security evaluation apparatus, the AI large model security evaluation apparatus further comprises: The uncertainty score determination module is configured to determine a target uncertainty score of the first evaluation AI large model for predicting the to-be-tested reply text. The first result determination module is configured to determine the first prediction risk label as a risk prediction result of the to-be-tested reply text if the target uncertainty score does not exceed the uncertainty score threshold. The second result determination module is configured to obtain a second prediction risk label of the expert for predicting the to-be-tested reply text as the risk prediction result of the to-be-tested reply text if the target uncertainty score exceeds the uncertainty score threshold.
[0078] The application also provides a computer readable storage medium storing a computer program, and the computer program is executed by a processor to implement the AI large model security evaluation method described in any of the above embodiments.
[0079] The application also provides a computer program product, which, when running on an electronic device, causes the electronic device to perform the AI large model security evaluation method described in any of the above embodiments.
[0080] Figure 4 FIG. 4 is a schematic diagram of an electronic device according to an embodiment of the application. As shown in FIG. 4, the electronic device 4 according to this embodiment comprises a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40. Figure 4 When the processor 40 executes the computer program 42, the steps in the above-mentioned embodiments of the AI large model security evaluation method are implemented, for example, steps 101-103 shown in FIG. 1. Figure 1 Alternatively, when the processor 40 executes the computer program 42, the functions of the modules / units in the above-mentioned device embodiments are implemented, for example, the functions of the modules 301-303 of the device shown in FIG. 3. Figure 3 Alternatively, when the processor 40 executes the computer program 42, the functions of the modules / units in the above-mentioned device embodiments are implemented, for example, the functions of the modules 301-303 of the device shown in FIG. 3.
[0081] The computer program 42 can be divided into one or more modules / units, which are stored in the memory 41 and executed by the processor 40 to complete the application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 42 in the electronic device 4.
[0082] The processor 40 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0083] The memory 41 can be an internal storage unit of the electronic device 4, such as a hard disk or a memory of the electronic device 4. The memory 41 can also be an external storage device of the electronic device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Further, the memory 41 can also include both the internal storage unit and the external storage device of the electronic device 4. The memory 41 is used to store the computer program and other programs and data required by the electronic device. The memory 41 can also be used to temporarily store data that has been output or will be output.
[0084] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0085] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, apparatus and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0086] In the above-described embodiments, the description of each embodiment focuses on different aspects, and parts not described in detail or recorded in a certain embodiment can be referred to the relevant description of other embodiments.
[0087] Those skilled in the art can appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0088] In the embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the above-described system embodiments are only schematic, for example, the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0089] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiments of the present application.
[0090] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0091] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0092] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An AI large model security evaluation method, characterized in that, The method comprises: obtaining a to-be-tested reply text output by a to-be-tested AI large model based on a risk problem text; selecting, from a plurality of evaluation AI large models, a first evaluation AI large model that is adapted to a first risk dimension to which the to-be-tested reply text belongs; outputting, by the first evaluation AI large model, a first predicted risk label of the to-be-tested reply text by predicting the to-be-tested reply text.
2. The method of claim 1, wherein, The selecting, from a plurality of evaluation AI large models, a first evaluation AI large model that is adapted to a first risk dimension to which the to-be-tested reply text belongs comprises: inputting the to-be-tested reply text into a trained routing model for processing, and outputting, by the routing model, an adaptation degree weight of each of the evaluation AI large models and the to-be-tested reply text; wherein the routing model is a machine learning model that calculates the adaptation degree weight of each of the evaluation AI large models and the to-be-tested reply text based on historical evaluation data of each of the evaluation AI large models on the first risk dimension; selecting, from the plurality of evaluation AI large models, an evaluation AI large model with the highest adaptation degree weight of the to-be-tested reply text as the first evaluation AI large model.
3. The method of claim 2, wherein, The training data set of the routing model is constructed by the following method: obtaining a sample text data set, an expert label data set, a model prediction result data set, and a model prediction uncertainty score data set; wherein the sample text data set comprises a plurality of sample reply texts covering a plurality of preset risk dimensions, the expert label data set comprises true risk labels of a part of the sample reply texts, the model prediction result data set comprises predicted risk labels of each of the evaluation AI large models on each of the sample reply texts, the model prediction uncertainty score data set comprises uncertainty scores of each of the evaluation AI large models in predicting each of the sample reply texts, the plurality of preset risk dimensions are divided into a plurality of different risk levels, different risk levels correspond to different risk coefficients, and each of the sample reply texts is respectively determined to correspond to the risk level according to the respective preset risk dimension; respectively counting the interface calling cost and the expert labeling cost of each of the evaluation AI large models; constructing the training data set according to the sample text data set, the expert label data set, the model prediction result data set, the model prediction uncertainty score data set, and the interface calling cost and the expert labeling cost of each of the evaluation AI large models.
4. The method of claim 3, wherein, Let a second risk dimension be any risk dimension in the plurality of preset risk dimensions, let a second evaluation AI large model be any evaluation AI large model in the plurality of evaluation AI large models, and let a target score range interval be any score range interval in a plurality of preset score range intervals; before the constructing the training data set according to the sample text data set, the expert label data set, the model prediction result data set, the model prediction uncertainty score data set, and the interface calling cost and the expert labeling cost of each of the evaluation AI large models, the method further comprises: obtaining a set of uncertainty scores of the second evaluation AI large model predicting the sample reply texts belonging to the second risk dimension; arranging each uncertainty score in the set of uncertainty scores from small to large and dividing each uncertainty score into the plurality of preset score range intervals; combining the expert label data set, calculating an actual error rate of the second evaluation AI large model predicting corresponding to the second risk dimension and the target score range interval; determining a score correction value of the target score range interval according to the actual error rate and the average of each uncertainty score in the target score range interval; if the score correction value exceeds a preset tolerance threshold, correcting each uncertainty score in the target score range interval according to the score correction value, and returning to the step of arranging each uncertainty score in the set of uncertainty scores from small to large and dividing each uncertainty score into the plurality of preset score range intervals until the score correction value does not exceed the tolerance threshold.
5. The method of claim 3 or 4, wherein, The objective function of the routing model is used to minimize the total labeling cost under the constraint of the overall error, the objective function introduces the risk coefficient as the error control coefficient, and the parameters optimized by the objective function include network weights and uncertainty score thresholds.
6. The method of claim 5, wherein, The training process of the routing model includes: randomly initializing the network weights and initializing the uncertainty score threshold as the median of each uncertainty score contained in the model prediction uncertainty score data set; using an alternating optimization method, fixing the uncertainty score threshold, updating the network weights by gradient descent method to minimize the objective function, and recalculating the uncertainty score threshold according to the updated network weights to constrain the overall error until the convergence condition is met, obtaining the final network weights and uncertainty score threshold.
7. The method of claim 6, wherein, After the first prediction risk label of the test reply text is output by predicting the test reply text through the first evaluation AI large model, the method further includes: determining a target uncertainty score of the first evaluation AI large model predicting the test reply text; if the target uncertainty score does not exceed the uncertainty score threshold, determining the first prediction risk label as the risk prediction result of the test reply text; if the target uncertainty score exceeds the uncertainty score threshold, obtaining a second prediction risk label of the test reply text predicted by an expert as the risk prediction result of the test reply text. 8.A device for evaluating AI large model security, the device comprising: including: a reply text acquisition module for acquiring a test reply text output by a test AI large model based on a risk problem text; an evaluation model selection module for selecting a first evaluation AI large model adapted to a first risk dimension to which the test reply text belongs from a plurality of evaluation AI large models; a risk prediction module for predicting the test reply text through the first evaluation AI large model and outputting a first prediction risk label of the test reply text.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the AI large model security evaluation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer program product, characterised in that, When the computer program product runs on the electronic device, the electronic device is caused to execute the AI large model security evaluation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Risk measurement evaluation method and device based on artificial intelligence and computer equipment
CN114238441A
Risk assessment method and device, equipment, storage medium and computer program product
CN119273452A
Answer matching method and system of artificial intelligence large model
CN119311843A
Risk assessment method and device, computer equipment, readable storage medium and program product
CN119648395A
Fusion question and answer method, device and equipment of mixed expert large language model and medium
CN119692477A