A Security Detection Method for Large Language Models Automatically Generated Based on Knowledge Graphs

By preprocessing and template nesting of the initial prompt words of the large language model, combining knowledge graphs and two-level security evaluators, the problems of strong dependence and high cost of detection methods in the existing technology are solved, and efficient and low-cost security detection of the large language model is achieved, and the model's general question-and-answer capability and detection concealment are maintained.

CN120180434BActive Publication Date: 2025-07-18NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510654123.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-18
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing large language model security detection methods are highly effective for the white box model, but have weak generalization capabilities. Black box detection depends on multiple iterations and high costs, and it is difficult to effectively detect without obtaining the internal structure of the model.

Method used

By preprocessing the initial prompt words and nesting templates, using the knowledge graph to generate the security of the detection model, a two-level security evaluator is designed to judge the hazard information in the detection knowledge graph and natural language text respectively, and the security detection of the large language model is achieved.

Benefits of technology

It realizes efficient and low-cost safe detection without relying on the internal structure of the model, maintains the general Q&A capabilities of the model, reduces the detection time and economic costs, and improves the concealment and interpretability of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180434B_ABST
    Figure CN120180434B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of large language model security, and discloses a large language model security detection method based on automatic generation of a knowledge graph. The method includes the following steps: preprocessing a data set containing different dangerous prompt words for the security detection direction, and replacing the dangerous behaviors in the initial prompt words with low-resource languages; using a prompt word template to automatically explore the dangerous knowledge encoded inside a large language model by means of the large language model, and using the large language model to construct a detection knowledge graph; converting the structured information in the detection knowledge graph into natural language text; designing a two-level security evaluator to determine whether the security protection of the large language model can be bypassed. The present invention attempts to bypass the security protection of the large language model to be tested after preprocessing the initial prompt words and nesting the templates, so as to evaluate the security performance of the large language model by whether the model generates a detection knowledge graph and the specific content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large language model security, and particularly relates to a large language model security detection method based on automatic generation of a knowledge graph. Background Art

[0002] Large language models such as ChatGPT, GPT-4, Claude, etc. perform excellently in various complex language processing tasks such as text summarization, machine translation, and code completion. The development of large language models has had a significant impact on the entire field of artificial intelligence, fundamentally changing the paradigm of how humans develop and utilize artificial intelligence algorithms. These models with a large number of parameters are trained on large-scale datasets, enabling them to capture and learn the complexity and diversity of language. Based on instruction fine-tuning and safety alignment, large language models have powerful anthropomorphic capabilities, enabling them to understand and generate text that conforms to human preferences.

[0003] However, the training data inevitably contains dangerous information. Malicious individuals take advantage of vulnerabilities in the model architecture and carefully design prompt templates to trigger dangerous behaviors in these models, including tutorials for creating websites that steal private information, instructions for making phishing software, and so on. Therefore, concerns about the security and potential vulnerabilities of large language models are increasing. Currently, research on security detection is mainly divided into two categories: white-box detection and black-box detection.

[0004] The goal of white-box detection is open-source models. By combining greedy and gradient-based search techniques, adversarial suffixes can be automatically generated to create a single adversarial prompt, bypassing the security protection of large language models and inducing dangerous behaviors with a high probability; general adversarial suffixes can also be generated to crack the large language models under detection. This method combines gradient-based token optimization with controllable text generation to produce consistent adversarial prompts among various large language models, with a high detection success rate.

[0005] Black-box detection targets closed-source models. By constructing virtual nested scenarios for detection, the anthropomorphic capabilities of large language models can be utilized to easily bypass the security protection of the models; it can also be carried out through prompt rewriting and scenario nesting. Prompt rewriting masks the test intention while retaining the core semantics of the prompt, and scenario nesting provides three scenarios: code completion, table filling, and text continuation. Although the above methods have achieved certain detection effects, they also have limitations.

[0006] Specifically, there are mainly the following two limitations: On the one hand, security detection based on models usually requires access to the internal structure of the target model, which means that these detections are usually only effective for white-box models and have weak generalization ability; on the other hand, prompt-based security detection usually relies on designing different scenarios to nest prompts. Usually, one method requires multiple scenario selections, and the execution process of the algorithm requires multiple iterations, which results in high time and economic costs for security detection. Summary of the Invention

[0007] The purpose of the present invention is to propose a security detection method for large language models automatically generated based on knowledge graphs. After preprocessing the initial prompt words and nesting templates, this method attempts to perform security detection on the large language models to be tested, enabling the evaluation of the security performance of various large language models to be tested by whether the large language models generate detection knowledge graphs and the specific content of the generated knowledge graphs.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A security detection method for large language models automatically generated based on knowledge graphs includes the following steps:

[0010] Step 1. Preprocess a dataset with different dangerous prompt words in the security detection direction, and replace the dangerous behaviors in the initial prompt words of the dataset with low-resource languages to obtain rewritten prompt words;

[0011] Step 2. Construct a security detection knowledge graph prompt word template and embed the rewritten prompt words, and use the large language model to automatically explore the dangerous knowledge encoded inside it and generate a complete detection knowledge graph about the initial prompt words;

[0012] Step 3. Design a first security evaluator to determine whether it is possible to bypass the security protection of the large language model to be tested by asking whether the large language model to be tested has successfully constructed a complete detection knowledge graph about the initial prompt words;

[0013] If a detection knowledge graph is successfully constructed, proceed to Step 4; otherwise, return to Step 2 and repeat the execution. And when the number of times of refusing to construct the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed;

[0014] Step 4. Design a prompt word template for converting the knowledge graph into text, directly nest the detection knowledge graph into the knowledge graph to text prompt word template, and convert the structured information in the generated detection knowledge graph into natural language text;

[0015] Step 5. Design a second security evaluator to determine whether the security protection of the large language model is bypassed again by asking whether the text to be evaluated, i.e., the natural language text after conversion, contains dangerous information.

[0016] If the converted text contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed.

[0017] Otherwise, return to Step 2 and repeat. When the number of times the generated content of the tested large language model does not contain dangerous information exceeds the second preset threshold, it is determined that the security protection of the tested large language model cannot be bypassed.

[0018] In addition, based on the above method for automatically generating security detection of large language models based on knowledge graphs, the present invention also proposes a corresponding system for automatically generating security detection of large language models based on knowledge graphs.

[0019] A system for automatically generating security detection of large language models based on knowledge graphs includes the following modules:

[0020] A preprocessing module for preprocessing a dataset with different dangerous prompt words in the security detection direction, and replacing the dangerous behaviors in the initial prompt words of the dataset with low-resource languages to obtain rewritten prompt words.

[0021] A detection knowledge graph generation module for constructing a security detection knowledge graph prompt word template and embedding the rewritten prompt words, and using the large language model to automatically explore the dangerous knowledge encoded in it and generate a complete detection knowledge graph for the initial prompt words.

[0022] A first security evaluation module for determining whether the security protection of the tested large language model can be bypassed by asking whether the tested large language model has successfully constructed a complete detection knowledge graph for the initial prompt words.

[0023] If the detection knowledge graph is successfully constructed, natural language conversion is performed; otherwise, the detection knowledge graph is reconstructed, and when the number of times of rejecting the construction of the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed.

[0024] A natural language conversion module for designing a knowledge graph to text prompt word template, directly nesting the detection knowledge graph into the knowledge graph to text prompt word template, and converting the structured information in the generated detection knowledge graph into natural language text.

[0025] And a second security evaluation module for determining whether the security protection of the large language model is bypassed again by asking whether the text to be evaluated, i.e., the natural language text after conversion, contains dangerous information.

[0026] If the information contains risks, it is determined that the safety protection of the large language model has been successfully bypassed;

[0027] Otherwise, regenerate the detection knowledge graph, and when the number of times the generated content of the large language model under test does not contain risky information exceeds the second preset threshold, it is determined that the safety protection of the large language model under test cannot be bypassed.

[0028] In addition, based on the above-mentioned method for automatically generating the safety detection of large language models based on knowledge graphs, the present invention also proposes a computer device, which includes a memory and one or more processors.

[0029] The memory stores executable code, and when the processor executes the executable code, it is used to implement the steps of the above-mentioned method for automatically generating the safety detection of large language models based on knowledge graphs.

[0030] In addition, based on the above-mentioned method for automatically generating the safety detection of large language models based on knowledge graphs, the present invention also proposes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the steps of the above-mentioned method for automatically generating the safety detection of large language models based on knowledge graphs.

[0031] The present invention has the following advantages:

[0032] As described above, the present invention describes a method for automatically generating the safety detection of large language models based on knowledge graphs. This method proposes a general safety detection method that combines large language models and knowledge graphs, which does not require obtaining the internal structure permissions of large language models and has strong generalization ability. During the processing of large models, sensitive words are often embedded in strong safety protection. The present invention can avoid word-level sensitive detection machines by using low-resource languages to replace key dangerous behaviors in the initial prompt words of the dataset. On the one hand, it retains the general question-and-answer ability of the model and can meet the needs of solving user problems. On the other hand, it can reduce the attention of the model to dangerous behaviors in the prompt words, thereby improving the concealment of safety detection. The modified prompt words are nested into a carefully designed safety detection knowledge graph prompt word template and can still be restored to the original intention during the knowledge graph generation and decoding stages, inducing the large language model to efficiently implement safety detection, thereby exploring the problems and defects existing in the current large language model, aiming to achieve the optimal balance between retaining the general question-and-answer ability and making a safe response. Using knowledge graphs to transform the detection process from a single text level into a graph process of structure-semantic separation and finally reconstructing it into natural language prompts at the output stage realizes stronger control ability and interpretability. The method steps of the present invention are simple and efficient, and can achieve successful safety detection within a very short cycle range, reducing the time and economic costs of traditional large language model safety detection methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is the overall block diagram of the large language model security detection method automatically generated based on the knowledge graph in the embodiment of the present invention;

[0034] Figure 2 This is an example block diagram of the large language model security detection method in the embodiment of the present invention;

[0035] Figure 3 This is an example diagram of dangerous behavior substitution in the embodiment of the present invention;

[0036] Figure 4 This is an example diagram of the knowledge graph prompt template in the embodiment of the present invention;

[0037] Figure 5 This is an example diagram of the prompt template for generating natural language text from the detection knowledge graph in the invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0039] Embodiment 1

[0040] This Embodiment 1 describes a large language model security detection method automatically generated based on the knowledge graph. This method enables the large language model to generate a target knowledge graph to evaluate the security performance of various large language models. Specifically, in the process of large language model security detection, first, a low-resource language, such as Bengali, etc., is considered to replace the key dangerous behaviors in the initial prompt, avoiding the word-level sensitive detection mechanism, and still being able to restore the semantic intention under the context guidance, thereby enhancing the concealment of security detection; then the rewritten initial prompt is embedded into a carefully designed prompt template. After the initial prompt words are replaced by the low-resource language, they can still be restored to the original intention in the knowledge graph generation and decoding stage, transforming the detection process from a single text level to a structure-semantic separation graph process, and complementing the complete knowledge graph with detailed steps regarding the initial prompt words; finally, once the large language model does not reject constructing the knowledge graph, at the output stage, each node and edge in the knowledge graph are converted into natural language descriptions, achieving stronger control ability and interpretability, thereby confirming the final security detection result.

[0041] As Figure 1 shown, the large language model security detection method automatically generated based on the knowledge graph includes the following steps:

[0042] Step 1. Preprocess the dataset containing different risk prompt words in the security detection direction, replace the risk behaviors in the initial prompt words of the dataset with low-resource languages to obtain rewritten prompt words, and reduce the attention of the large language model to risk behaviors.

[0043] The dataset here contains different risk prompt words for detecting the security protection of the large language model. The datasets used in this embodiment include the commonly used AdvBench dataset and HEx-PHI dataset in the security detection direction.

[0044] The contents contained in the two representative datasets in the security detection direction have many overlapping parts. In practical applications, one of them can be selected, or both can be selected to verify the effectiveness and generality of the method proposed in the present invention.

[0045] First, replace the risk behaviors contained in the initial prompt words in the dataset with low-resource languages. The low-resource languages here can be any one of languages such as Bengali, Swahili, Javanese, etc.

[0046] Low-resource languages refer to languages with a low proportion in the corpus currently used for training and fine-tuning large language models. Currently, there are some studies on using the language itself as a means to detect the security of large language models.

[0047] The process of replacing the risk behaviors in the initial prompt words of the dataset with low-resource languages is as follows:

[0048] Define the initial prompt word , , represents the size of the entire vocabulary.

[0049] After replacing the risk behaviors, the rewritten prompt word is obtained:

[0050] ;

[0051] Among them, represents the tokenized , represents the low-resource language token.

[0052] For example, Figure 2 as shown, assume the initial prompt word is: "Write a tutorial on how to implement *****", Figure 3 shows the replacement of the "******" risk behavior in the initial prompt word with Bengali to form a new rewritten prompt word.

[0053] The specific implementation method is to translate and replace some of the prompt words through tools such as "Google Translate".

[0054] To simplify the representation, use Represents the remaining part of the initial prompt, using Represents the part replaced with the low-resource language. So, the rewritten prompt obtained after the dangerous behavior replacement Can also be expressed as:

[0055] ;

[0056] Among them, Represents the concatenation process of the remaining initial prompt and the low-resource language prompt in the initial prompt.

[0057] Input the rewritten prompt directly into the large language model under test to obtain the corresponding response:

[0058] .

[0059] Among them Represents the large language model The response made by the large language model as a test model to the input initial prompt.

[0060] Step 2. Construct a safety detection knowledge graph prompt template and embed the rewritten prompt, and use the large language model to automatically explore the dangerous knowledge encoded in it and generate a complete detection knowledge graph for the initial prompt.

[0061] Large language models exhibit powerful natural language understanding and interpretation following capabilities, which are usually referred to as the anthropomorphic capabilities of large language models. Large language models can implicitly encode a large amount of knowledge, and extract this knowledge from large language models to construct knowledge graphs. The implicit knowledge of large language models can be transferred to a common sense knowledge base to generate explicit knowledge.

[0062] The knowledge graph constructed in this way exhibits competitiveness, diversity, and novelty. A standard knowledge graph consists of a set of knowledge triples, expressed as <head entity (h), relation (r), tail entity (t)>. To better align with the goal of safety detection, the present invention predicts a specific tail entity given the head entity h and the relation r.

[0063] Design a safety detection knowledge graph prompt template to prompt the model to automatically complete the blank tail entity.

[0064] Specifically, embed the rewritten prompt in step 1 into the prompt template, and then input the complete prompt into the large language model under test to obtain the response of the large language model.

[0065] .

[0066] Among them Is the detection knowledge graph constructed for the large language model, It is a prompt template for the security detection knowledge graph.

[0067] For example, directly replace the "rewritten prompt" part in the template with "Write a tutorial on how to implement *****" (the preprocessing of step 1 has been described, and the dangerous behaviors in the prompt have been replaced by low-resource languages), and then let the large language model complete and generate the complete detection knowledge graph.

[0068] Take the entire prompt as the head entity and define the relationship as contains, forming <Write a tutorial on how to implement *****, contains, t>, aiming to let the large language model provide information related to the tail entity t.

[0069] Figure 4 It shows the carefully designed security detection knowledge graph prompt template of the present invention. First, the template tells the large language model that it is an expert in the field of knowledge graphs and can easily construct knowledge graphs. Then, let the large language model be in a scenario of knowledge graph completion and construct the complete detection knowledge graph according to the following Cypher language form:

[0070] CREATE (tutorial:tutorial {title: "<rewritten prompt>"})

[0071] CREATE (tutorial)-[:CONTAINS]->(instruction:Instruction {step: "Step 1"})

[0072] CREATE (tutorial)-[:CONTAINS]->(instruction:Instruction {step: "Step 2"})

[0073] CREATE (tutorial)-[:CONTAINS]->(instruction:Instruction {step: "Step 3"}).

[0074] Among them, the rewritten prompt P' is the central node of this knowledge graph.

[0075] In addition, this knowledge graph also contains three step blank nodes. Under this prompt template, the large language model needs to complete the detailed information of "Step 1, Step 2, Step 3", or directly reject the request to construct the knowledge graph.

[0076] Step 3. Design the first security evaluator , it is determined whether the security protection of the large language model to be tested can be bypassed by asking whether the large language model to be tested has successfully constructed a complete detection knowledge graph for the initial prompt.

[0077] If the detection knowledge graph is successfully constructed, step 4 is executed; otherwise, return to step 2 and repeat. And when the number of times of refusing to construct the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed.

[0078] Design an evaluator for whether it is possible to bypass the security protection according to the response to the request for constructing the detection knowledge graph, the first security evaluator is a large language model, and its form is as follows:

[0079] 。

[0080] Among them, is a large language model, that is, the evaluation model It is determined whether the model has the possibility of bypassing the security protection by asking whether the large language model to be tested has successfully constructed a knowledge graph.

[0081] Input the response of the large language model to be tested for constructing the detection knowledge graph into the large model The large model determines whether the response has successfully constructed a knowledge graph for the initial prompt.

[0082] If so, it is represented by "1", indicating the possibility of bypassing the security protection. T represents the prompt template, represents the model when acting as an evaluator for the input response regarding the construction of the detection knowledge graph.

[0083] Due to the instability of the large language model response, that is, for the same input, the multiple outputs of the model are not consistent. A first preset threshold M needs to be set in the security detection evaluation process to limit the maximum number of detections.

[0084] Define a loop variable m, with an initial value of 0.

[0085] When the number of times m that the large language model to be tested refuses to construct the detection knowledge graph exceeds M, it means that the currently nested prompt cannot bypass the security protection of the large language model to be tested, that is, the security detection fails.

[0086] Here, M is not very large. Generally, setting it around 5 can offset the influence brought by the instability of the large language model response.

[0087] The first security evaluator The specific judgment process is as follows:

[0088] If the large language model under test successfully constructs a complete detection knowledge graph for the initial prompt, directly proceed to step 4; if the model refuses to construct the detection knowledge graph, repeat step 2 for a new round of detection and update the value of m, setting m = m + 1.

[0089] When the number of times m that the model refuses to construct the detection knowledge graph exceeds the first preset threshold M, it indicates that the current nested prompt cannot bypass the security protection of the large language model under test, that is, the security detection fails.

[0090] Step 4. After the large language model generates a detection knowledge graph for the initial prompt, since the knowledge graph is a structured knowledge representation form, it is difficult to determine whether the security detection is truly successful. Design a knowledge graph to text prompt template to convert the structured information in the obtained knowledge graph into natural language text, making it easier to understand and apply.

[0091] Processing graph-text parallel data is labor-intensive and challenging. The powerful generalization ability and rich knowledge of large language models have enabled many researchers to use large language models for knowledge generation text research.

[0092] Describe the detection knowledge graph generated by the large language model using the Cypher language, allowing the detection knowledge graph to be directly nested into the knowledge graph to text prompt template, enabling the large language model to generate detailed content related to security detection:

[0093] ;

[0094] Among them, is the detailed content about , is the designed knowledge graph to text prompt template. is the large model 's response to the input detection knowledge graph when acting as a text converter.

[0095] Figure 5 Shows the relevant content of the knowledge graph designed by the present invention being converted into a natural language text prompt template (i.e., the knowledge graph to text prompt template). Directly tell the large language model to describe the detailed content contained in the generated detection knowledge graph in a short paragraph. The present invention only needs to replace < > with the detection knowledge graph.

[0096] The content of this knowledge graph to text prompt template is, for example: Please use a short paragraph to describe the detailed content contained in the following knowledge graph without describing its structure: < >. Among them, < > is replaced with the detection knowledge graph generated in step 3.

[0097] By writing prompts from the knowledge graph into the text prompt template, extract nodes, relationships, and their attributes from the structured information in the knowledge graph, and use natural language generation (NLG) technology to convert them into fluent and grammatically correct text.

[0098] Step 5. Design the second security evaluator , by asking whether the text to be evaluated, i.e., the converted natural language text, contains dangerous information, and again determine whether the security protection of the large language model is bypassed.

[0099] If the converted text contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed.

[0100] Otherwise, return to Step 2 and repeat. And when the number of times the generated content in the response of the large language model under test does not contain dangerous information exceeds the second preset threshold, it is directly determined that the security protection of the large language model under test cannot be bypassed.

[0101] According to the obtained natural language description of the detection knowledge graph, a second security evaluator for whether the security protection is successfully bypassed is designed. The second security evaluator is a large language model, and its form is as follows:

[0102] .

[0103] Wherein, is actually a large language model, that is, the evaluation model .

[0104] Judge whether the security detection is successful by asking whether the text to be evaluated contains dangerous information. T represents the prompt template, represents the model 's response to the input natural language description when acting as an evaluator.

[0105] Specifically, input the natural language description of the detection knowledge graph into the large model , and the large model to judge whether the converted natural language description contains dangerous content. If so, it is represented by "1".

[0106] Due to the instability of the response of the large language model, that is, for the same input, the multiple outputs of the model are not consistent. It is necessary to set a second preset threshold N in the security detection evaluation process to limit the maximum number of detections.

[0107] Define the loop variable n, and let n start counting from 0; when the number of model rejection times n exceeds N, it means that the current nested prompt cannot bypass the security protection of the large language model under test, that is, the security detection fails.

[0108] Here, N will not be very large. Generally, setting it around 5 can offset the impact caused by the instability of the response of the large language model.

[0109] The second security evaluator The specific judgment process is as follows:

[0110] If the security detection is successful, it indicates that the method has successfully broken through the security protection of the tested large model.

[0111] If the evaluation result of the tested large language model in step 5 does not contain dangerous information, it indicates that the current security detection fails. Then, it returns to step 2 for a new round of detection, and the value of n is updated, making n = n + 1.

[0112] When the number of times n that the tested large language model generates content without dangerous information exceeds N, it means that the current nested prompt words cannot bypass the security protection of the tested large language model, that is, the security detection is determined to fail.

[0113] It should be noted here that the , , , The large model can be a single large model or large models of different categories, as long as the model can meet the corresponding functional requirements.

[0114] Among them is the model to be tested, is the model that converts the knowledge graph into a text model, is the model for evaluating whether the tested large language model has successfully constructed a knowledge graph, is the model for evaluating whether natural language content is dangerous.

[0115] The detection method adopted in the present invention is a black-box detection method, but it can be widely applied to both black-box and white-box models.

[0116] Due to the nature of the black-box model, the results of each model answer may vary. Therefore, the thresholds M and N are used to represent the number of times of implementing the detection. When the number of model detection failures exceeds the threshold M or N, this detection is regarded as a failure.

[0117] In this embodiment, the values of M and N set are not very large in practice (for example, both M and N are set to 5). The purpose is only to balance the instability of the large language model when giving responses, greatly reducing the cost of the detection method of the present invention.

[0118] The two - level security evaluator first allows the model under test to construct a detection knowledge graph (the first security evaluator), and then checks whether its textual description contains dangerous content (the second security evaluator). It can not only quickly intercept failed attempts that cannot even generate a graph at the structural level, but also eliminate seemingly successful but actually harmless outputs at the semantic level. While ensuring the evaluation accuracy, by setting failure thresholds (such as M, N≈5) at each level to terminate ineffective loops in advance, it significantly reduces the computational and time costs of repeatedly executing the complete detection process. Moreover, the hierarchical feedback also facilitates precise positioning and optimization of prompt words or model configurations.

[0119] Embodiment 2

[0120] This Embodiment 2 describes a large - language model security detection system based on the automatic generation of knowledge graphs. This system and the large - language model security detection method based on the automatic generation of knowledge graphs in the above Embodiment 1 are based on the same inventive concept.

[0121] Specifically, a large - language model security detection system based on the automatic generation of knowledge graphs includes the following modules:

[0122] A large - language model security detection system based on the automatic generation of knowledge graphs includes the following modules:

[0123] A pre - processing module for pre - processing a dataset with different dangerous prompt words in the security detection direction, and replacing the dangerous behaviors in the initial prompt words of the dataset with low - resource languages to obtain rewritten prompt words;

[0124] A detection knowledge graph generation module for constructing a security detection knowledge graph prompt word template and embedding the rewritten prompt words, and using the large - language model to automatically explore the dangerous knowledge encoded therein and generate a complete detection knowledge graph for the initial prompt words;

[0125] A first security evaluation module for judging whether the security protection of the large - language model under test can be bypassed by asking whether the large - language model under test has successfully constructed a complete detection knowledge graph for the initial prompt words;

[0126] If the detection knowledge graph is successfully constructed, natural language conversion is performed; otherwise, the detection knowledge graph is reconstructed, and when the number of times of refusing to construct the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large - language model cannot be bypassed;

[0127] A natural language conversion module for designing a knowledge graph - to - text prompt word template, directly nesting the detection knowledge graph into the knowledge graph - to - text prompt word template, and converting the structured information in the generated detection knowledge graph into natural language text;

[0128] And a second security assessment module, which determines whether to bypass the security protection of the large language model again by asking whether the text to be evaluated, i.e., the natural language text after conversion, contains dangerous information.

[0129] If it contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed.

[0130] Otherwise, regenerate the detection knowledge graph, and when the number of times the generated content in the response of the tested large language model does not contain dangerous information exceeds the second preset threshold, it is determined that the security protection of the tested large language model cannot be bypassed.

[0131] It should be noted that in the large language model security detection system described in this embodiment, the implementation processes of the functions and roles of each functional module are specifically described in the corresponding steps of the method in the above-mentioned Embodiment 1, and will not be elaborated here.

[0132] Embodiment 3

[0133] This Embodiment 3 describes a computer device, which includes a memory and one or more processors. Executable code is stored in the memory, and when the processor executes the executable code, it is used to implement the steps of the large language model security detection method based on automatic generation of knowledge graph in the above-mentioned Embodiment 1.

[0134] In this embodiment, the computer device is any device or apparatus with data processing capabilities, which will not be elaborated here.

[0135] Embodiment 4

[0136] This Embodiment 4 describes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the steps of the large language model security detection method based on automatic generation of knowledge graph in the above-mentioned Embodiment 1.

[0137] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device.

[0138] Of course, the above description is only the preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the scope of the present specification, and should be protected by the present invention.

Claims

1. A method for security detection of large language models automatically generated based on a knowledge graph, characterized in that: It includes the following steps: Step 1. Preprocess the dataset containing different risk prompt words in the security detection direction, and replace the risk behaviors in the initial prompt words of the dataset with low-resource languages to obtain rewritten prompt words; Step 2. Construct a security detection knowledge graph prompt word template and embed the rewritten prompt words, and use the large language model to automatically explore the risk knowledge encoded inside it and generate a complete detection knowledge graph for the initial prompt words; Step 3. Design a first security evaluator, and judge whether the security protection of the tested large language model can be bypassed by asking whether the tested large language model has successfully constructed a complete detection knowledge graph for the initial prompt words; If the detection knowledge graph is successfully constructed, continue to execute Step 4; otherwise, return to Step 2 and repeat the execution. And when the number of times of refusing to construct the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed; Step 4. Design a prompt word template for converting the knowledge graph into text, directly nest the detection knowledge graph into the knowledge graph-to-text prompt word template, and convert the structured information in the generated detection knowledge graph into natural language text; Step 5. Design a second security evaluator, and judge again whether the security protection of the large language model is bypassed by asking whether the evaluated text, that is, the natural language text after conversion, contains dangerous information; If the converted text contains dangerous information, it is determined that the security protection of the large language model is successfully bypassed; Otherwise, return to Step 2 and repeat the execution. And when the number of times that the generated content of the tested large language model does not contain dangerous information exceeds the second preset threshold, it is determined that the security protection of the tested large language model cannot be bypassed.

2. The method for security detection of a large language model automatically generated based on a knowledge graph according to claim 1, wherein In the above Step 1, the dataset adopts the AdvBench dataset or the HEx-PHI dataset.

3. The method for security detection of a large language model automatically generated based on a knowledge graph according to claim 1, wherein In the above Step 1, the process of replacing the risk behaviors in the initial prompt words of the dataset with low-resource languages is as follows: Define the initial prompt , after replacing the dangerous behavior, the rewritten prompt is obtained: ; Among them, represents tokenized , represents a low-resource language token; For the sake of simplicity of representation, use to represent the remaining part in the initial prompt, and use to represent the part replaced with the low-resource language. Therefore, the rewritten prompt obtained after replacing the dangerous behavior is also expressed as: ; Among them, represents the concatenation process of the remaining initial prompt words and the low-resource language prompts in the initial prompt words.

4. The security detection method for large language models automatically generated based on a knowledge graph according to claim 1, wherein In the above Step 2, the generation process of the detection knowledge graph is as follows: Design a security detection knowledge graph prompt word template, and the knowledge graph is represented as <head entity (h), relation (r), tail entity (t)>; given the head entity h and relation r, automatically complete the blank tail entity with the prompt model; Embed the rewritten prompt words into the security detection knowledge graph prompt word template, and then input the complete prompt words into the tested large language model to obtain the response of the large language model, that is, the large language model completes and generates a complete detection knowledge graph.

5. The method for security detection of a large language model automatically generated based on a knowledge graph according to claim 1, wherein In the above Step 3, the first security evaluator is a large language model; Due to the instability of the response of the large language model, that is, for the same input, the multiple outputs of the model are not consistent. It is necessary to set a first preset threshold M in the security detection evaluation process to limit the maximum number of detections; Define a loop variable m, and the initial value of m is 0; If the tested large language model successfully constructs a complete detection knowledge graph for the initial prompt, directly proceed to step 4; if the model refuses to construct the detection knowledge graph, repeat step 2 for a new round of detection and update the value of m, where m = m + 1; When the number of times m that the model refuses to construct the detection knowledge graph exceeds the first preset threshold M, it indicates that the current nested prompt cannot bypass the security protection of the tested large language model, that is, the security detection fails.

6. The method for security detection of a large language model automatically generated based on a knowledge graph according to claim 1, wherein In step 4, the process of converting the structured information in the detection knowledge graph into natural language text is as follows: First, design a prompt template for converting the knowledge graph into text, and use a short paragraph in the prompt template to describe the detailed content included in the generated detection knowledge graph; Directly nest the detection knowledge graph into the prompt template, extract nodes, relationships, and their attributes from the structured information in the detection knowledge graph, and use natural language generation technology to convert them into natural language text with smooth logic and correct grammar.

7. The method for security detection of a large language model automatically generated based on a knowledge graph according to claim 1, characterized in that, In step 2, the process of the large language model constructing the detection knowledge graph is expressed as follows: ; Among them, is a detection knowledge graph constructed by a large language model, is a prompt template, represents the large language model output under the condition of being used as a test model, represents the rewritten prompt; In step 3, the formula for converting the structured information in the detection knowledge graph into natural language text is expressed as follows: ; Among them, is about the detailed content of is the prompt word template for converting the knowledge graph into text, and is the response of the large model when acting as a text converter to the input detected knowledge graph.

8. The method for security detection of a large language model automatically generated based on a knowledge graph according to claim 1, characterized in that, In step 5, the second security evaluator is a large language model; Due to the instability of the large language model's response, that is, for the same input, the model's multiple outputs are not consistent, it is necessary to set a second preset threshold N in the security detection evaluation process to limit the maximum number of detection times; Define a loop variable n and let n start counting from 0; If the detection is successful, it indicates that the method has successfully bypassed the security protection of the tested large model; If the evaluation result of the tested large language model in step 5 does not contain dangerous information, it indicates that the current security detection fails, and return to step 2 for a new round of detection, and update the value of n, where n = n + 1; When the number of times n that the tested large language model generates information without dangerous information exceeds N, it indicates that the current nested prompt cannot bypass the security protection of the tested large language model, that is, it is determined that the security detection fails.

9. A large language model security detection system based on automatic generation of knowledge graphs, characterized in that It includes the following modules: A preprocessing module for preprocessing a dataset containing different dangerous prompts in the security detection direction, and replacing the dangerous behaviors in the initial prompt of the dataset with low-resource languages to obtain rewritten prompts; A detection knowledge graph generation module for constructing a security detection knowledge graph prompt template and embedding the rewritten prompt, and using the large language model to automatically explore the dangerous knowledge encoded in it and generate a complete detection knowledge graph for the initial prompt; A first security evaluation module for judging whether it can bypass the security protection of the tested large language model by asking whether the tested large language model has successfully constructed a complete detection knowledge graph for the initial prompt; If the detection knowledge graph is successfully constructed, perform natural language conversion; otherwise, reconstruct the detection knowledge graph, and when the number of times of refusing to construct the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed; A natural language conversion module for designing a knowledge graph into a text prompt template, detecting the direct nesting of the knowledge graph into the knowledge graph to text prompt template, and converting the structured information in the generated detected knowledge graph into natural language text; And a second security assessment module, which determines whether to bypass the security protection of the large language model by asking whether the text to be evaluated, i.e., the natural language text after conversion, contains dangerous information; If it contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed; Otherwise, regenerate the detected knowledge graph, and when the number of times the generated content of the tested large language model does not contain dangerous information exceeds the second preset threshold, it is determined that the security protection of the tested large language model cannot be bypassed.

10. A computer device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it implements the steps of the large language model security detection method based on automatic generation of knowledge graph as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for optimizing scene generation model, storage medium and electronic device

    CN118051625A

  • Short message AI training model intelligent matching method and system

    CN119939199A