General assessment method for safety performance of large language model

By constructing a knowledge graph and dynamically generating evaluation data using retrieval enhancement technology, security evaluation of large language models is solved, and the problem of lack of general methods for security performance evaluation of large language models in the existing technology is solved, and a comprehensive evaluation and enhancement of model security performance is achieved.

CN119989354APending Publication Date: 2025-05-13WUXI URBAN BLOCKCHAIN ADVANCED RES CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411855909.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing large language models use different quality of text data during the pre-training stage, resulting in the risk of harmful content and personal information leakage, and are vulnerable to jailbreak attacks, affecting security and moral risks.

Method used

Provide a general evaluation method, by collecting and classifying data sets, building a knowledge graph, using retrieval enhancement technology to dynamically generate evaluation data, evaluate large language models to be detected, and judge its security performance.

Benefits of technology

A comprehensive evaluation of the security performance of large language models is achieved, the risk of harmful content and personal information leakage is reduced, the security and robustness of the model are enhanced, and it is suitable for different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989354A_ABST
    Figure CN119989354A_ABST
Patent Text Reader

Abstract

The invention discloses a general assessment method for safety performance of a large language model. Comprising the following steps: S1, collecting a data set; s2, collecting a large language model; s3, constructing a knowledge graph; s4, performing fine adjustment on the large language model; s5, making a threshold value and a scoring rule; s6, generating evaluation data; s7, evaluating the large language model; s8, analyzing the performance of the large language model; the evaluation method is not limited to a certain environment, parameter quantity of a large model and functions of the large model, any large model can be comprehensively evaluated, that is, question and answer output of the to-be-detected large language model is evaluated through a safety evaluation model, field evaluation data is dynamically generated, and the evaluation efficiency is greatly improved. The specific field evaluation data is used for outputting question and answer data of the to-be-tested large language model, the field evaluation data is independent of a training data set of the to-be-tested large language model, the quality of the used data set is evaluated, cheating is difficult to perform during evaluation, and the robustness of an evaluation result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information security technology, and more specifically, particularly relates to a general evaluation method for the security performance of a large language model. Background Art

[0002] With the rapid development of machine learning, artificial intelligence has become one of the most popular topics. Large language models (LLMs) have excellent performance in many fields (such as search engines, translation, text generation, etc.) due to their profound language understanding, more human-like text generation, context awareness, and powerful problem-solving capabilities. Many companies are attracted by the excellent performance of large language models and have begun to build and develop their own large language models. These large language models can generate answers close to human levels to questions raised by users, and can further use technical means such as prompt engineering, professional knowledge bases, and third-party tools to show the potential for wide application in key fields such as medicine, finance, and law.

[0003] However, the base models used by these applications or artificial intelligence systems developed based on large language models are usually pre-trained using massive amounts of text data on the Internet. The data sets that these basic models rely on during the pre-training phase often have varying quality. Therefore, some text content that does not meet the standards is inevitably mixed into the training data used. Therefore, most companies' applications developed using pre-trained LLMs will generate risks of harmful content or leak personal identity information.

[0004] Moreover, many studies have proven that large language models are vulnerable to jailbreak attacks. Jailbreak attacks in large language models refer to techniques that bypass model content restrictions and security mechanisms, causing the model to generate sensitive, inappropriate, or harmful content beyond its design intent. The principle of large model jailbreak attacks is to exploit the structure or language logic of the model to make it output content that is strictly restricted at ordinary times. This not only destroys the usability of the large model but also causes privacy leaks. With the increasing application of AI today, this attack has brought new challenges to the security and ethical risks of artificial intelligence. Summary of the invention

[0005] In view of the problems existing in the prior art, the purpose of the present invention is to provide a general evaluation method for the security performance of large language models. By utilizing the excellent natural language understanding and text generation capabilities of large language models, the method can understand the meaning of data in static data sets, extract and split the data and store it in the form of a knowledge graph in a graph database, and then set different parameters according to the needs for different large language models to be tested, dynamically generate evaluation data using retrieval enhancement technology, and then use the evaluation data to evaluate the model to be tested.

[0006] To achieve the above object, the present invention provides the following technical solution: a general evaluation method for the security performance of a large language model, comprising the following steps:

[0007] S1. Collecting datasets: The collected datasets are divided into three parts. One part is a general dataset, which is used to build a knowledge graph to generate general evaluation data. Another part is a fine-tuning dataset, which is used to fine-tune the large language model to make it a security assessment model. The last part is a dedicated dataset, which is a dataset of specific knowledge in a certain field. It is used to generate field evaluation data and is reserved for users.

[0008] S2. Collect large language models: Collect three large language models, one of which is used to build the knowledge graph, the other is a detection model used to judge the large language model to be detected. The detection model is fine-tuned using the second part of the data set, and the last model is the large language model to be detected.

[0009] S3. Build a knowledge graph: Use a large language model to build a knowledge graph for the collected general data set and special data set, and generate general evaluation data and domain evaluation data;

[0010] S4. Fine-tune the large language model: Use the collected fine-tuning dataset to fine-tune the detection model so that the large language model forms a security assessment model;

[0011] S5. Formulate thresholds and scoring rules: Formulate thresholds and scoring rules based on the characteristics of the large language model and the needs of the usage scenario;

[0012] S6. Generate evaluation data: Use the general data set to build a good knowledge graph, and use retrieval enhancement technology to generate general evaluation data according to the characteristics of the large language model and the needs of the usage scenario;

[0013] S7. Evaluate the large language model: Use the evaluation data to evaluate the large language model to be tested, so that the large language model to be tested outputs corresponding answers to the evaluation data questions;

[0014] S8. Analyze the performance of the large language model: Use the security assessment model to evaluate the answers output by the large language model to be tested, and score them according to the threshold and scoring rules to determine the security performance of the large language model to be tested.

[0015] Specifically, the specific steps of S1 are as follows:

[0016] S11. Before building a knowledge graph using general evaluation data, you should collect as many data sets as possible to build the knowledge graph, and collect the data sets by categories;

[0017] S12. Before fine-tuning the model for security assessment, fine-tune the large language model using a fine-tuning dataset, where the classification of the fine-tuning dataset is the same as S11, but not intersecting;

[0018] S13. Before the detection model starts to evaluate the large language model to be detected, if the user has special needs for evaluation in certain specific vertical fields, the user is required to provide the corresponding field evaluation data when building the knowledge graph, and the field evaluation data is retained by the user.

[0019] Specifically, the specific implementation method of step S2 is:

[0020] S21. The large language model for building the knowledge graph using a general data set should come from a reliable open source website. The number of parameters of the large language model should be at least 7B, and 14B is optimal.

[0021] S22. The security assessment model generated using the fine-tuning dataset should also come from a reliable open source website, or be trained by yourself. The number of parameters should be one of 6B, 7B, 13B, and 14B, with 14B performing best.

[0022] S23. The data set of domain-specific knowledge is provided to users and used to generate specific domain evaluation data. It is retained by users, and the models to be tested are also owned and provided by users.

[0023] Specifically, the knowledge graph selects a suitable graph database, then segments the data set to extract entities, and then constructs the relationship between entities; and the construction steps of the knowledge graph are as follows:

[0024] Entity recognition and relationship extraction: Use natural language processing techniques to identify entities and extract relationships between entities from general and specialized datasets;

[0025] Transform and generate data: extract entities and relationships from general data sets and special data sets, and generate knowledge graphs from entities and relationships to form general evaluation data and domain evaluation data;

[0026] Graph storage and representation: The generated general evaluation data and domain evaluation data are stored in a graph database, usually using RDF or graph database format.

[0027] Specifically, the model of the security assessment is to fine-tune the detection model through the collected fine-tuning data set. Taking the model with 14B parameters as an example, it is necessary to train at least 3 epochs, and make appropriate adjustments according to the parameters of the detection model, the size of the fine-tuning data set and its own algorithm;

[0028] The detection model is fine-tuned using the fine-tuning dataset through back-propagation and gradient descent algorithms. Fine-tuning usually uses a relatively small learning rate to avoid destroying the general knowledge that has been learned. During the fine-tuning process, the detection model adjusts the weights according to the requirements of the target task, so that it can better understand and handle the task.

[0029] Specifically, the calculation of the threshold and scoring rule formulated in S5 is as follows:

[0030] Because different large language models are used in different scenarios, different thresholds and scoring criteria are required to evaluate large models. Certain sensitive types of content require stricter review, the threshold of the evaluation criteria will become relatively high, and the weight of the weighted score will also increase accordingly;

[0031]

[0032] Among them, P1, P2, P3, P4, P5, and P6 are the scores of six major categories of evaluation data. The highest score of each item is 100 points. The scores are greater than the pre-selected threshold to be considered qualified, and the minimum value of the threshold is 70 points. α, β, γ, δ, θ, and μ are the weights corresponding to the six major categories of evaluation data. The sum of α, β, γ, δ, θ, and μ is equal to one, and α, β, γ, δ, θ, and μ are respectively greater than zero. S is the comprehensive score with a full score of 100 points.

[0033] Specifically, the general evaluation data in S6 utilizes a large language model and retrieval enhancement technology to generate a corresponding amount of evaluation data from a general data set, and the general data set generates the general evaluation data according to a threshold and a scoring weight.

[0034] Specifically, the evaluation data in S7 is input to the large language model to be tested, and the large language model to be tested outputs a corresponding answer to each question of the evaluation data, and then the questions of the evaluation data and the corresponding answers are structured and stored in the form of QA pairs.

[0035] Specifically, the detection model in S8 evaluates the answer output by the large language model to be detected:

[0036] The stored QA pairs are input into the security assessment model, which will judge whether each QA pair is safe and reasonable and give corresponding analysis. Finally, a comprehensive score of the large language model to be tested is obtained based on the judgment result and the score calculation formula in S51. If the score is greater than the threshold, the large model to be tested is considered to be safe, otherwise.

[0037] Technical effects and advantages of the present invention:

[0038] The present invention is the first universal large-scale model safety performance evaluation method.

[0039] The evaluation method of the present invention is not limited to a certain environment, the number of parameters of a large model, and the function of a large model, and can comprehensively evaluate any large model, that is, the question-answering output of the large language model to be tested is evaluated through the model of security evaluation, and the domain evaluation data used is dynamically generated, and the large language model to be tested is output with the specific domain evaluation data provided by the user, and the domain evaluation data is independent of the training data set of the large language model to be tested. The quality of the data set used for evaluation is also difficult to cheat, which ensures the robustness of the evaluation results;

[0040] The present invention utilizes the excellent natural language understanding and text generation capabilities of the large language model to understand the meaning of the data in the static data set, extracts and splits the data and stores it in the form of a knowledge graph in a graph database, and then sets different parameters according to the needs for different large language models to be tested, dynamically generates evaluation data using retrieval enhancement technology, and then uses the evaluation data to evaluate the model to be tested. It is suitable for different scenarios and solves the current lack of a universal and unified method for evaluating the security performance of large language models.

[0041] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic diagram of the steps provided by the present invention;

[0043] Figure 2 It is a schematic diagram of a process of collecting a data set provided by the present invention;

[0044] Figure 3 It is a schematic diagram of the large language model selection process provided by the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0046] like Figures 1 to 3 As shown, the general evaluation method for the security performance of a large language model provided by an embodiment of the present invention includes the following steps:

[0047] S1. Collecting datasets: The collected datasets are divided into three parts. One part is a general dataset, which is used to build a knowledge graph to generate general evaluation data. Another part is a fine-tuning dataset, which is used to fine-tune the large language model to make it a security assessment model. The last part is a dedicated dataset, which is a dataset of specific knowledge in a certain field. It is used to generate field evaluation data and is reserved for users.

[0048] S2. Collect large language models: Collect three large language models, one of which is used to build the knowledge graph, the other is a detection model used to judge the large language model to be detected. The detection model is fine-tuned using the second part of the data set, and the last model is the large language model to be detected.

[0049] S3. Build a knowledge graph: Use a large language model to build a knowledge graph for the collected general data set and special data set, and generate general evaluation data and domain evaluation data;

[0050] S4. Fine-tune the large language model: Use the collected fine-tuning dataset to fine-tune the detection model so that the large language model forms a security assessment model;

[0051] S5. Formulate thresholds and scoring rules: Formulate thresholds and scoring rules based on the characteristics of the large language model and the needs of the usage scenario;

[0052] S6. Generate evaluation data: Use the general data set to build a good knowledge graph, and use retrieval enhancement technology to generate general evaluation data according to the characteristics of the large language model and the needs of the usage scenario;

[0053] S7. Evaluate the large language model: Use the evaluation data to evaluate the large language model to be tested, so that the large language model to be tested outputs corresponding answers to the evaluation data questions;

[0054] S8. Analyze the performance of the large language model: Use the security assessment model to evaluate the answers output by the large language model to be tested, and score them according to the threshold and scoring rules to determine the security performance of the large language model to be tested.

[0055] In this embodiment, preferably, the specific steps of S1 are as follows:

[0056] S11. Before building a knowledge graph using general evaluation data, you should collect as many data sets as possible to build the knowledge graph, and collect the data sets by categories;

[0057] S12. Before fine-tuning the model for security assessment, fine-tune the large language model using a fine-tuning dataset, where the classification of the fine-tuning dataset is the same as S11, but not intersecting;

[0058] S13. Before the detection model starts to evaluate the large language model to be detected, if the user has special needs for evaluation in certain specific vertical fields, the user is required to provide the corresponding field evaluation data when building the knowledge graph, and the field evaluation data is retained by the user;

[0059] It should be noted that by building a knowledge graph for general evaluation data, collecting a large amount of data for classification, and fine-tuning the large language model by fine-tuning the data set, the large language model is convenient for testing the large language model to be tested, and for evaluating and processing domain evaluation data for special fields.

[0060] In this embodiment, preferably, the specific implementation method of step S2 is:

[0061] S21. The large language model for building the knowledge graph using a general data set should come from a reliable open source website. The number of parameters of the large language model should be at least 7B, and 14B is optimal.

[0062] S22. The security assessment model generated using the fine-tuning dataset should also come from a reliable open source website, or be trained by yourself. The number of parameters should be one of 6B, 7B, 13B, and 14B, with 14B performing best.

[0063] S23. The data set of domain-specific knowledge is provided by the user and used to generate specific domain evaluation data. It is retained by the user, and the model to be tested is also owned and provided by the user.

[0064] It should be noted that a knowledge graph is built for the data set through a large language model, and a model for security assessment is established to improve the detection of the model to be tested.

[0065] In this embodiment, preferably, the knowledge graph selects a suitable graph database, then segments the data set to extract entities, and then constructs the relationship between entities; and the steps of constructing the knowledge graph are as follows:

[0066] Entity recognition and relationship extraction: Use natural language processing techniques to identify entities and extract relationships between entities from general and specialized datasets;

[0067] Transform and generate data: extract entities and relationships from general data sets and special data sets, and generate knowledge graphs from entities and relationships to form general evaluation data and domain evaluation data;

[0068] Graph storage and representation: Store the generated general evaluation data and domain evaluation data into a graph database, usually in RDF or graph database format;

[0069] It should be noted that by selecting a suitable graph database, entities and relationships in the data set can be extracted, which makes it easier to build a knowledge graph for the entities and relationships in the entity data set and store the knowledge graph.

[0070] In this embodiment, preferably, the model of the security assessment is to fine-tune the detection model through the collected fine-tuning data set, wherein taking the model with 14B parameters as an example, it is necessary to train at least 3 epochs, and make appropriate adjustments according to the parameters of the detection model, the size of the fine-tuning data set and its own algorithm;

[0071] Use the fine-tuning dataset to fine-tune the detection model through back propagation and gradient descent algorithms. Fine-tuning usually uses a relatively small learning rate to avoid destroying the general knowledge that has been learned. During the fine-tuning process, the detection model adjusts the weights according to the requirements of the target task, so that it can better understand and handle the task.

[0072] It should be noted that the security assessment model is fine-tuned by fine-tuning the dataset, and the detection model is fine-tuned through back propagation and gradient descent algorithms, so that the security assessment model can better understand and handle the task.

[0073] In this embodiment, preferably, the calculation of the threshold and scoring rule formulated in S5 is as follows:

[0074] Because different large language models are used in different scenarios, different thresholds and scoring criteria are required to evaluate large models. Certain sensitive types of content require stricter review, the threshold of the evaluation criteria will become relatively high, and the weight of the weighted score will also increase accordingly;

[0075]

[0076] Among them, P1, P2, P3, P4, P5, and P6 are the scores of the six major categories of evaluation data. The highest score of each item is 100 points. The score is greater than the pre-selected threshold to be considered qualified, and the minimum value of the threshold is 70 points. α, β, γ, δ, θ, and μ are the weights corresponding to the six major categories of evaluation data. The sum of α, β, γ, δ, θ, and μ is equal to one, and α, β, γ, δ, θ, and μ are all greater than zero. S is the comprehensive score, and the full score is 100 points.

[0077] It should be noted that evaluating the large model through different thresholds and scoring criteria will facilitate stricter review of certain sensitive types of content, and the threshold of the evaluation criteria will become relatively high, and the weight of the weighted score will also increase accordingly.

[0078] In this embodiment, preferably, the general evaluation data in S6 uses a large language model and retrieval enhancement technology to generate a corresponding number of evaluation data from the general data set, and the general data set generates the general evaluation data according to the threshold and the scoring weight;

[0079] It should be noted that, by using a large language model and retrieval enhancement technology, a corresponding amount of evaluation data is generated from a general data set, so as to improve the accuracy of the evaluation data and realize detection processing of the large language model to be detected.

[0080] In this embodiment, preferably, the evaluation data in S7 is input to the large language model to be detected, and the large language model to be detected outputs a corresponding answer to each question of the evaluation data, and then the question of the evaluation data and the corresponding answer are structured and stored in the form of a QA pair;

[0081] It should be noted that by storing the questions and corresponding answers of the assessment data in a structured manner in the form of QA pairs, it is convenient to improve the storage efficiency and security between questions and answers, and it is convenient to review and test the questions and answers.

[0082] In this embodiment, preferably, the detection model in S8 evaluates the answer output by the large language model to be detected:

[0083] The stored QA pairs are input into the security assessment model. The security assessment model will judge whether each QA pair is safe and reasonable and give corresponding analysis. Finally, a comprehensive score of the large language model to be tested is obtained according to the judgment result and the score calculation formula in S51. If the score is greater than the threshold, the large model to be tested is considered to be safe, otherwise;

[0084] It should be noted that the model that passes the security assessment will judge whether each QA pair is safe and reasonable and give corresponding analysis, which is convenient for the comprehensive scoring of the large language model to be tested, and the security of the large model to be tested is determined based on the threshold and score.

[0085] The specific operation process of this application:

[0086] Step 1: Collect data sets: Divide the collected data sets into three parts. One part is a general data set, which is used to build a knowledge graph to generate general evaluation data. Another part is a fine-tuning data set, which is used to fine-tune the large language model to make it a security assessment model. The last part is a dedicated data set, which is a data set of specific knowledge in a certain field. It is used to generate field evaluation data and is reserved for users.

[0087] Step 2: Collect large language models: Collect three large language models, one of which is used to build the knowledge graph, and the other is a detection model used to judge the large language model to be detected. The detection model is fine-tuned using the second part of the data set, and the last model is the large language model to be detected;

[0088] Step 3: Build a knowledge graph: Use the large language model to build a knowledge graph for the collected general data set and special data set, and generate general evaluation data and domain evaluation data;

[0089] Step 4: Fine-tune the large language model: Use the collected fine-tuning dataset to fine-tune the detection model so that the large language model forms a security assessment model;

[0090] Step 5: Formulate thresholds and scoring rules: Formulate thresholds and scoring rules based on the characteristics of the large language model and the needs of the usage scenario;

[0091] Step 6: Generate evaluation data: Use the general data set to build a good knowledge graph, and use retrieval enhancement technology to generate general evaluation data according to the characteristics of the large language model and the needs of the usage scenario;

[0092] Step 7: Evaluate the large language model: Use the evaluation data to evaluate the large language model to be tested, so that the large language model to be tested outputs corresponding answers to the evaluation data questions;

[0093] Step 8. Analyze the performance of the large language model: Use the security assessment model to evaluate the answers output by the large language model to be tested, and score them according to the threshold and scoring rules to determine the security performance of the large language model to be tested.

[0094] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A general evaluation method for the security performance of large language models, characterized by: The steps include: S1. Collecting datasets: The collected datasets are divided into three parts. One part is a general dataset, which is used to build a knowledge graph to generate general evaluation data. Another part is a fine-tuning dataset, which is used to fine-tune the large language model to make it a security assessment model. The last part is a dedicated dataset, which is a dataset of specific knowledge in a certain field. It is used to generate field evaluation data and is reserved for users. S2. Collect large language models: Collect three large language models, one of which is used to build the knowledge graph, the other is a detection model used to judge the large language model to be detected. The detection model is fine-tuned using the second part of the data set, and the last model is the large language model to be detected. S3. Build a knowledge graph: Use a large language model to build a knowledge graph for the collected general data set and special data set, and generate general evaluation data and domain evaluation data; S4. Fine-tune the large language model: Use the collected fine-tuning dataset to fine-tune the detection model so that the large language model forms a security assessment model; S5. Formulate thresholds and scoring rules: Formulate thresholds and scoring rules based on the characteristics of the large language model and the needs of the usage scenario; S6. Generate evaluation data: Use the general data set to build a good knowledge graph, and use retrieval enhancement technology to generate general evaluation data according to the characteristics of the large language model and the needs of the usage scenario; S7. Evaluate the large language model: Use the evaluation data to evaluate the large language model to be tested, so that the large language model to be tested outputs corresponding answers to the evaluation data questions; S8. Analyze the performance of the large language model: Use the security assessment model to evaluate the answers output by the large language model to be tested, and score them according to the threshold and scoring rules to determine the security performance of the large language model to be tested.

2. The general evaluation method for the security performance of a large language model according to claim 1, characterized in that: The specific steps of S1 are as follows: S11. Before building a knowledge graph using general evaluation data, you should collect as many data sets as possible to build the knowledge graph, and collect the data sets by categories; S12. Before fine-tuning the model for security assessment, fine-tune the large language model using a fine-tuning dataset, where the classification of the fine-tuning dataset is the same as S11, but not intersecting; S13. Before the detection model starts to evaluate the large language model to be detected, if the user has special needs for evaluation in certain specific vertical fields, the user is required to provide the corresponding field evaluation data when building the knowledge graph, and the field evaluation data is retained by the user.

3. The general evaluation method for the security performance of a large language model according to claim 1, characterized in that: The specific implementation method of step S2 is: S21. The large language model for building the knowledge graph using a general data set should come from a reliable open source website. The number of parameters of the large language model should be at least 7B, and 14B is optimal. S22. The security assessment model generated using the fine-tuning dataset should also come from a reliable open source website, or be trained by yourself. The number of parameters should be one of 6B, 7B, 13B, and 14B, with 14B performing best. S23. The data set of domain-specific knowledge is provided to users and used to generate specific domain evaluation data. It is retained by users, and the models to be tested are also owned and provided by users.

4. The general evaluation method for the security performance of a large language model according to claim 1, characterized in that: The knowledge graph selects a suitable graph database, then segments the data set to extract entities, and then constructs the relationship between entities; and the construction steps of the knowledge graph are as follows: Entity recognition and relationship extraction: Use natural language processing techniques to identify entities and extract relationships between entities from general and specialized datasets; Transform and generate data: extract entities and relationships from general data sets and special data sets, and generate knowledge graphs from entities and relationships to form general evaluation data and domain evaluation data; Graph storage and representation: The generated general evaluation data and domain evaluation data are stored in a graph database, usually using RDF or graph database format.

5. The general evaluation method for the security performance of a large language model according to claim 1, characterized in that: The model of the security assessment is to fine-tune the detection model through the collected fine-tuning data set. Taking the model with 14B parameters as an example, it is necessary to train at least 3 epochs and make appropriate adjustments according to the parameters of the detection model, the size of the fine-tuning data set and its own algorithm; The detection model is fine-tuned using the fine-tuning dataset through back-propagation and gradient descent algorithms. Fine-tuning usually uses a relatively small learning rate to avoid destroying the general knowledge that has been learned. During the fine-tuning process, the detection model adjusts the weights according to the requirements of the target task, so that it can better understand and handle the task.

6. The general evaluation method for the security performance of a large language model according to claim 1, characterized in that: The calculation of the threshold and scoring rules formulated in S5 is as follows: Because different large language models are used in different scenarios, different thresholds and scoring criteria are required to evaluate large models. Certain sensitive types of content require stricter review, the threshold of the evaluation criteria will become relatively high, and the weight of the weighted score will also increase accordingly; Among them, P1, P2, P3, P4, P5, and P6 are the scores of six major categories of evaluation data. The highest score of each item is 100 points. The scores are greater than the pre-selected threshold to be considered qualified, and the minimum value of the threshold is 70 points. α, β, γ, δ, θ, and μ are the weights corresponding to the six major categories of evaluation data. The sum of α, β, γ, δ, θ, and μ is equal to one, and α, β, γ, δ, θ, and μ are respectively greater than zero. S is the comprehensive score with a full score of 100 points.

7. The general evaluation method for the security performance of a large language model according to claim 6, characterized in that: The general evaluation data in S6 utilizes a large language model and retrieval enhancement technology to generate a corresponding amount of evaluation data from a general data set, and the general data set generates general evaluation data according to a threshold and a scoring weight.

8. The general evaluation method for the security performance of a large language model according to claim 1, characterized in that: The evaluation data in S7 is input to the large language model to be tested, and the large language model to be tested outputs a corresponding answer to each question of the evaluation data, and then the question of the evaluation data and the corresponding answer are structured and stored in the form of QA pairs.

9. The general evaluation method for the security performance of a large language model according to claim 7, characterized in that: The detection model in S8 evaluates the answer output by the large language model to be detected: The stored QA pairs are input into the security assessment model, which will judge whether each QA pair is safe and reasonable and give corresponding analysis. Finally, a comprehensive score of the large language model to be tested is obtained based on the judgment result and the score calculation formula in S51. If the score is greater than the threshold, the large model to be tested is considered to be safe, otherwise.