Large language model security detection method based on automatic generation of knowledge graph

Through the automated generation method based on knowledge graph, the existing large language model security detection methods depend on white box model, weak generalization ability and high detection cost are solved, and efficient, hidden and interpretable security detection of large language models is achieved.

CN120180434AActive Publication Date: 2025-06-20NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510654123.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing large language model security detection methods have limitations, including dependence on white box models, weak generalization capabilities, and high detection time and cost.

Method used

The automated generation method based on knowledge graph is used to perform safety detection of large language models. By preprocessing the initial prompt words, a safety detection knowledge graph prompt word template is constructed, and a detection knowledge graph is generated using the large language model to evaluate the security performance of the model.

Benefits of technology

It realizes that the security detection of large language models is not limited by internal structure permissions, has strong generalization capabilities, reduces detection time and economic costs, and improves the concealment and interpretability of security detection while retaining the model's Q&A capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180434A_ABST
    Figure CN120180434A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of big language model security, and discloses a big language model security detection method based on automatic generation of a knowledge graph. The method comprises the following steps: preprocessing a data set containing different danger prompt words in a security detection direction, and replacing dangerous behaviors in initial prompt words with a low-resource language; the method comprises the following steps: automatically exploring dangerous knowledge encoded in a cue word template by using a large language model through the cue word template, and constructing a detection knowledge graph by using the large language model; converting structured information in the detection knowledge graph into a natural language text; and designing a two-stage security evaluator to judge whether security protection of a large language model can be bypassed. According to the method, the initial cue word is subjected to preprocessing and template nesting, and then the security protection of the tested large language model is tried to be bypassed, so that the security performance of the large language model is evaluated by judging whether the model generates the detection knowledge graph and the specific content or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large language model security, and particularly relates to a large language model security detection method based on automatic generation of a knowledge graph. Background Art

[0002] Large language models such as ChatGPT, GPT-4, Claude, etc. perform excellently in various complex language processing tasks such as text summarization, machine translation, and code completion. The development of large language models has had a significant impact on the entire field of artificial intelligence, fundamentally changing the paradigm of how humans develop and utilize artificial intelligence algorithms. These models with a large number of parameters are trained on large-scale datasets, enabling them to capture and learn the complexity and diversity of language. Based on instruction fine-tuning and safety alignment, large language models have powerful anthropomorphic capabilities, enabling them to understand and generate text that conforms to human preferences.

[0003] However, the training data inevitably contains dangerous information. Malicious individuals take advantage of vulnerabilities in the model architecture and carefully design prompt templates to trigger dangerous behaviors in these models, including creating tutorials for websites that steal private information, instructions for making phishing software, and so on. Therefore, concerns about the security and potential vulnerabilities of large language models are increasing. Currently, research on security detection is mainly divided into two categories: white-box detection and black-box detection.

[0004] The goal of white-box detection is open-source models. By combining greedy and gradient-based search techniques, adversarial suffixes can be automatically generated to create a single adversarial prompt, bypassing the security protection of large language models and inducing dangerous behaviors with a high probability; or general adversarial suffixes can be generated to crack the large language models under detection. This method combines gradient-based token optimization with controllable text generation to produce consistent adversarial prompts among various large language models, with a high detection success rate.

[0005] Black-box detection targets closed-source models. By constructing virtual nested scenarios for detection, the anthropomorphic capabilities of large language models can be utilized to easily bypass the security protection of the models; or it can be carried out through prompt rewriting and scenario nesting. Prompt rewriting masks the test intention while retaining the core semantics of the prompt, and scenario nesting provides three scenarios: code completion, table filling, and text continuation. Although the above methods have achieved certain detection effects, they also have limitations.

[0006] Specifically, there are mainly the following two limitations: on the one hand, model-based security detection usually requires access to the internal structure of the target model, which means that these detections are usually only effective for white-box models and have weak generalization ability; on the other hand, prompt-based security detection usually relies on designing different scenarios to nest prompts. Usually, one method requires multiple scenario selections, and the execution process of the algorithm requires multiple iterations, which results in high time and economic costs for security detection. Summary of the Invention

[0007] The purpose of the present invention is to propose a security detection method for large language models automatically generated based on knowledge graphs. After preprocessing and template nesting of the initial prompt words, this method attempts to perform security detection on the large language models to be tested, enabling the evaluation of the security performance of various large language models to be tested by whether the large language models generate detection knowledge graphs and the specific content of the generated knowledge graphs.

[0008] To achieve the above object, the present invention adopts the following technical solutions: A security detection method for large language models automatically generated based on knowledge graphs, comprising the following steps: Step 1. Preprocess a dataset whose security detection direction contains different dangerous prompt words, and replace the dangerous behaviors in the initial prompt words of the dataset with low-resource languages to obtain rewritten prompt words; Step 2. Construct a security detection knowledge graph prompt word template and embed the rewritten prompt words, and use the large language model to automatically explore the dangerous knowledge encoded therein and generate a complete detection knowledge graph about the initial prompt words; Step 3. Design a first security evaluator, and judge whether the security protection of the large language model to be tested can be bypassed by asking whether the large language model to be tested has successfully constructed a complete detection knowledge graph about the initial prompt words; If a detection knowledge graph is successfully constructed, continue to execute Step 4; otherwise, return to Step 2 and repeat the execution, and when the number of times of rejecting to construct the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed; Step 4. Design a prompt word template for converting the knowledge graph into text, directly nest the detection knowledge graph into the knowledge graph-to-text prompt word template, and convert the structured information in the generated detection knowledge graph into natural language text; Step 5. Design a second security evaluator, and judge again whether the security protection of the large language model is bypassed by asking whether the evaluated text, that is, the natural language text after conversion, contains dangerous information; If the converted text contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed; Otherwise, return to step 2 and repeat the execution. When the number of times the generated content by the large language model under test does not contain dangerous information exceeds the second preset threshold, it is determined that the security protection of the large language model under test cannot be bypassed.

[0009] In addition, based on the above method for automatically generating a security detection of a large language model based on a knowledge graph, the present invention also proposes a corresponding security detection system for automatically generating a large language model based on a knowledge graph.

[0010] A security detection system for automatically generating a large language model based on a knowledge graph includes the following modules: A preprocessing module, configured to preprocess a data set whose security detection direction includes different dangerous prompt words, and replace the dangerous behaviors in the initial prompt words of the data set with low-resource languages to obtain rewritten prompt words; A detection knowledge graph generation module, configured to construct a security detection knowledge graph prompt word template and embed the rewritten prompt words, and use the large language model to automatically explore the dangerous knowledge encoded therein and generate a complete detection knowledge graph about the initial prompt words; A first security evaluation module, configured to determine whether the security protection of the large language model under test can be bypassed by asking whether the large language model under test has successfully constructed a complete detection knowledge graph about the initial prompt words; If the detection knowledge graph is successfully constructed, perform natural language conversion; otherwise, reconstruct the detection knowledge graph, and when the number of times of rejecting to construct the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed; A natural language conversion module, configured to design a knowledge graph to text prompt word template, directly nest the detection knowledge graph into the knowledge graph to text prompt word template, and convert the structured information in the generated detection knowledge graph into natural language text; And a second security evaluation module, which determines whether the security protection of the large language model is bypassed again by asking whether the evaluated text, that is, the natural language text after conversion, contains dangerous information; If it contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed; Otherwise, regenerate the detection knowledge graph, and when the number of times the generated content by the large language model under test does not contain dangerous information exceeds the second preset threshold, it is determined that the security protection of the large language model under test cannot be bypassed.

[0011] In addition, based on the above method for automatically generating a security detection of a large language model based on a knowledge graph, the present invention also proposes a computer device, which includes a memory and one or more processors.

[0012] The memory stores executable code, and when the processor executes the executable code, it is used to implement the steps of the above-mentioned large language model security detection method automatically generated based on the knowledge graph.

[0013] In addition, based on the above-mentioned large language model security detection method automatically generated based on the knowledge graph, the present invention also proposes a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the steps of the above-mentioned large language model security detection method automatically generated based on the knowledge graph.

[0014] The present invention has the following advantages: As described above, the present invention describes a large language model security detection method automatically generated based on the knowledge graph. This method proposes a general security detection method that combines a large language model and a knowledge graph, does not require obtaining the internal structure permission of the large language model, and has strong generalization ability. During the processing of the large model, sensitive words are often embedded in strong security protection. The present invention can avoid the word-level sensitive detection mechanism by using low-resource languages to replace the key dangerous behaviors in the initial prompt words of the dataset. On the one hand, it retains the general question-and-answer ability of the model and can meet the needs of solving user problems. On the other hand, it can reduce the attention of the model to the dangerous behaviors in the prompt words, thereby enhancing the concealment of security detection. The modified prompt words are nested into a carefully designed security detection knowledge graph prompt word template and can still be restored to the original intention during the knowledge graph generation and decoding stages, inducing the large language model to efficiently implement security detection, thereby exploring the problems and defects existing in the current large language model, aiming to achieve the best balance between retaining the general question-and-answer ability and making a security response for the model. Using the knowledge graph to transform the detection process from a single text level into a graph process of structure-semantic separation and finally reconstructing it into a natural language prompt at the output stage to achieve stronger control ability and interpretability. The method steps of the present invention are simple and efficient, can achieve successful security detection within a very short cycle range, and reduce the time and economic costs of traditional large language model security detection methods. Description of the Drawings

[0015] Figure 1 It is the overall block diagram of the large language model security detection method automatically generated based on the knowledge graph in the embodiment of the present invention; Figure 2 It is the example block diagram of the large language model security detection method in the embodiment of the present invention; Figure 3 It is the example diagram of dangerous behavior substitution in the embodiment of the present invention; Figure 4 It is the example diagram of the knowledge graph prompt template in the embodiment of the present invention; Figure 5This is an example diagram for detecting knowledge graphs to natural language text generation prompt templates in the invention. Detailed implementation manners

[0016] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners: Example 1 This Example 1 describes a large language model security detection method based on automated generation of knowledge graphs. This method enables the large language model to generate target knowledge graphs to evaluate the security performance of various large language models. Specifically, during the large language model security detection process, first, low-resource languages, such as Bengali, are considered to replace critical dangerous behaviors in the initial prompt, avoiding word-level sensitive detection mechanisms while still being able to restore semantic intentions under context guidance, thereby enhancing the concealment of security detection; then the rewritten initial prompt is embedded into a carefully designed prompt template. After the initial prompt words are replaced by low-resource languages, they can still be restored to the original intention during the knowledge graph generation and decoding stages, transforming the detection process from a single text level to a structure-semantic separated graph process, and completing the detailed steps regarding the initial prompt words in the complete knowledge graph; finally, once the large language model does not reject constructing the knowledge graph, at the output stage, each node and edge in the knowledge graph are converted into natural language descriptions, achieving stronger control capabilities and interpretability, thereby confirming the final security detection results.

[0017] As Figure 1 shown, the large language model security detection method based on automated generation of knowledge graphs includes the following steps: Step 1. Preprocess the dataset containing different dangerous prompt words in the security detection direction, replace the dangerous behaviors in the initial prompt words of the dataset with low-resource languages to obtain rewritten prompt words, and reduce the attention of the large language model to dangerous behaviors.

[0018] The dataset here contains different dangerous prompt words for detecting the security protection of large language models. The datasets used in this example include the commonly used AdvBench dataset and HEx-PHI dataset in the security detection direction.

[0019] The contents included in these two representative datasets in the security detection direction have many overlapping parts. In practical applications, one of them can be selected, or both can be selected to verify the effectiveness and generality of the method proposed in the present invention.

[0020] First, replace the dangerous behaviors included in the initial prompt words in the dataset with low-resource languages. The low-resource languages here can be any one of languages such as Bengali, Swahili, Javanese, etc.

[0021] Low-resource languages refer to those languages that account for a low proportion in the corpora currently used for training and fine-tuning large language models. Currently, there are some studies on using the language itself as a means to detect the security of large language models.

[0022] The process of replacing dangerous behaviors in the initial prompt of the dataset with low-resource languages is as follows: Define the initial prompt , , represents the size of the entire vocabulary.

[0023] After replacing dangerous behaviors, the rewritten prompt is obtained: ; Among them, represents tokenized , represents the low-resource language token.

[0024] For example, Figure 2 as shown, assume the initial prompt is: "Write a tutorial on how to implement *****", Figure 3 shows replacing the dangerous behavior "******" in the initial prompt with Bengali to form a new rewritten prompt.

[0025] The specific implementation method is to translate and replace some of the prompts through tools such as "Google Translate".

[0026] To simplify the representation, use to represent the remaining part of the initial prompt, and use to represent the part replaced with the low-resource language. So the rewritten prompt is also expressed as: ; Among them, represents the concatenation process of the remaining initial prompt and the low-resource language prompt in the initial prompt.

[0027] Input the rewritten prompt directly into the large language model to be tested to obtain the corresponding response: .

[0028] Among them represents the large language model responding to the input initial prompt as a test model.

[0029] Step 2. Construct a safety detection knowledge graph prompt template and embed the rewritten prompt, and use the large language model to automatically explore the dangerous knowledge encoded in it and generate a complete detection knowledge graph for the initial prompt.

[0030] Large language models demonstrate powerful natural language understanding and the ability to follow explanations, which is commonly referred to as the anthropomorphic ability of large language models. Large language models can implicitly encode a large amount of knowledge, and extracting this knowledge from large language models to construct a knowledge graph. The implicit knowledge of large language models can be transferred to a common sense knowledge base to generate explicit knowledge.

[0031] The knowledge graph constructed in this way exhibits competitiveness, diversity, and novelty. A standard knowledge graph consists of a set of knowledge triples, represented as <head entity (h), relation (r), tail entity (t)>. To better align with the goal of security detection, the present invention predicts a specific tail entity given the head entity h and relation r.

[0032] Design a security detection knowledge graph prompt template to prompt the model to automatically complete the blank tail entity.

[0033] Specifically, embed the rewritten prompt in step 1 into the prompt template, and then input the complete prompt into the large language model under test to obtain the response of the large language model.

[0034] 。

[0035] where is the detection knowledge graph constructed for the large language model, is the security detection knowledge graph prompt template.

[0036] For example, directly replace the "rewritten prompt" part in the template with "Write a tutorial on how to implement *****" (the preprocessing of step 1 has been described, and the dangerous behaviors in the prompt have been replaced by low-resource languages), and then let the large language model complete the generation of the full detection knowledge graph.

[0037] Take the entire prompt as the head entity and define the relation as contains, forming <Write a tutorial on how to implement *****, contains, t>, aiming to let the large language model provide information related to the tail entity t.

[0038] Figure 4 Shows the carefully designed security detection knowledge graph prompt template of the present invention. First, the template tells the large language model that it is an expert in the field of knowledge graphs and can easily construct a knowledge graph. Then, let the large language model be in a scenario of knowledge graph completion and construct the complete detection knowledge graph according to the following Cypher language form: CREATE (tutorial:tutorial {title: "<rewritten prompt>"}) CREATE (tutorial)-[:CONTAINS]->(instruction:Instruction {step: "Step 1"}) CREATE (tutorial)-[:CONTAINS]->(instruction:Instruction {step: "Step 2"}) CREATE (tutorial)-[:CONTAINS]->(instruction:Instruction {step: "Step 3"})。

[0039] Among them, the rewritten prompt P' is used as the central node of this knowledge graph.

[0040] In addition, this knowledge graph also contains three step blank nodes. Under this prompt template, the large language model needs to complete the detailed information of "Step 1, Step 2, Step 3", or directly reject the request to construct the knowledge graph.

[0041] Step 3. Design the first security evaluator , by asking the tested large language model whether it has successfully constructed a complete detection knowledge graph for the initial prompt, to determine whether it can bypass the security protection of the tested large language model.

[0042] If the detection knowledge graph is successfully constructed, execute Step 4; otherwise, return to Step 2 and repeat. And when the number of times of rejecting the construction of the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed.

[0043] Design an evaluator for whether it is possible to bypass the security protection according to the response to the request for constructing the detection knowledge graph. The first security evaluator is a large language model, and its form is as follows: 。

[0044] Among them, is a large language model, that is, the evaluation model Judge whether the model has the possibility of bypassing the security protection by asking the tested large language model whether it has successfully constructed the knowledge graph.

[0045] Input the response of the tested large language model to the construction of the detection knowledge graph into the large model , and the large model judges whether the response has successfully constructed a knowledge graph for the initial prompt.

[0046] If so, it is represented by "1", indicating the possibility of bypassing the security protection. T represents the prompt template, Representation model When acting as an evaluator, the response to the input regarding the construction of the detection knowledge graph.

[0047] Due to the instability of the responses of large language models, that is, for the same input, the multiple outputs of the model are not consistent. It is necessary to set a first preset threshold M during the security detection and evaluation process to limit the maximum number of detection times.

[0048] Define a loop variable m with an initial value of 0.

[0049] When the number of times m that the tested large language model refuses to construct the detection knowledge graph exceeds M, it means that the currently nested prompt words cannot bypass the security protection of the tested large language model, that is, the security detection fails.

[0050] Here, M is not very large. Generally, setting it around 5 can offset the impact brought by the response instability of large language models.

[0051] The first security evaluator The specific judgment process is as follows: If the tested large language model successfully constructs a complete detection knowledge graph for the initial prompt words, directly go to step 4; if the model refuses to construct the detection knowledge graph, repeat step 2 for a new round of detection, and update the value of m, making m = m + 1.

[0052] When the number of times m that the model refuses to construct the detection knowledge graph exceeds the first preset threshold M, it means that the currently nested prompt words cannot bypass the security protection of the tested large language model, that is, the security detection fails.

[0053] Step 4. After the large language model generates the detection knowledge graph for the initial prompt words, since the knowledge graph is a structured form of knowledge representation, it is difficult to determine whether the security detection is truly successful. Design a knowledge graph to text prompt word template to convert the structured information in the obtained knowledge graph into natural language text, making it easier to understand and apply.

[0054] Processing graph-text parallel data is labor-intensive and challenging. The powerful generalization ability and rich knowledge of large language models have enabled many researchers to use large language models for the research of knowledge generation text.

[0055] Use the Cypher language to describe the detection knowledge graph generated by the large language model, allowing the detection knowledge graph to be directly nested into the knowledge graph to text prompt word template, enabling the large language model to generate detailed content related to security detection: ; Among them, is about The detailed content of is the knowledge graph of the design into a text prompt template. It is the response of the large model when acting as a text converter to the detected knowledge graph of the input.

[0056] Figure 5 It shows the relevant content of the knowledge graph designed by the present invention being transformed into a natural language text prompt template (i.e., the knowledge graph to text prompt template), directly telling the large language model to use a short paragraph to describe the detailed content contained in the generated detected knowledge graph. The present invention only needs to replace < > with the detected knowledge graph.

[0057] The content of this knowledge graph to text prompt template is, for example: Please use a short paragraph to describe the detailed content contained in the following knowledge graph, but do not describe its structure: < >. Among them, < > is replaced with the detected knowledge graph generated in step 3.

[0058] By writing the prompts in the knowledge graph to text prompt template, extract the nodes, relationships and their attributes from the structured information in the knowledge graph, and use natural language generation NLG technology to convert them into smooth and grammatically correct text.

[0059] Step 5. Design the second security evaluator , by asking whether the text to be evaluated, that is, the natural language text after conversion, contains dangerous information, and judge again whether the security protection of the large language model is bypassed.

[0060] If the converted text contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed.

[0061] Otherwise, return to step 2 and repeat the execution. And when the number of times that the generated content of the large language model under test does not contain dangerous information exceeds the second preset threshold, it is directly determined that the security protection of the large language model under test cannot be bypassed.

[0062] According to the obtained natural language description of the detected knowledge graph, a second security evaluator for judging whether the security protection is successfully bypassed is designed. The second security evaluator is a large language model, and its form is as follows: .

[0063] Among them, is actually a large language model, that is, the evaluation model .

[0064] Judge whether the security detection is successful by asking whether the text to be evaluated contains dangerous information. T represents the prompt template, represents the model The response made to the input natural language description when acting as an evaluator.

[0065] Specifically, input the natural language description about the detection knowledge graph into the large model in which the large model to determine whether the converted natural language description contains dangerous content. If so, it is represented by "1".

[0066] Due to the instability of the response of the large language model, that is, for the same input, the multiple outputs of the model are not consistent. It is necessary to set a second preset threshold N during the security detection and evaluation process to limit the maximum number of detection times.

[0067] Define a loop variable n and let n start counting from 0; when the number of model rejection times n exceeds N, it means that the current nested prompt words cannot bypass the security protection of the tested large language model, that is, the security detection fails.

[0068] Here, N is not very large. Generally, setting it around 5 can offset the influence brought by the instability of the response of the large language model.

[0069] The second security evaluator The specific judgment process is as follows: If the security detection is successful, it indicates that the method has successfully broken through the security protection of the tested large model.

[0070] If the evaluation result of the tested large language model in step 5 is that it does not contain dangerous information, it means that the current security detection fails. Go back to step 2 for a new round of detection, and update the value of n, making n = n + 1.

[0071] When the number of times n that the tested large language model generates content without dangerous information exceeds N, it means that the current nested prompt words cannot bypass the security protection of the tested large language model, that is, it is determined that the security detection fails.

[0072] It should be noted here that the , , , large model can be a single large model or different types of large models, as long as the model can meet the corresponding functional requirements.

[0073] Among them is the tested model, is the model that converts the knowledge graph into text, is the model for evaluating whether the tested large language model has successfully constructed the knowledge graph, is the model for evaluating whether the natural language content is dangerous.

[0074] The detection method adopted in the present invention is a black-box detection method, but it can be widely applied to both black-box and white-box models.

[0075] Due to the nature of the black-box model, the results of each answer of the model may vary. Therefore, the thresholds M and N are used to represent the number of times of performing the detection. When the number of failed detections of the model exceeds the threshold M or N, this detection is regarded as a failure.

[0076] In this embodiment, the values of M and N set are not very large in practice (for example, both M and N are set to 5). The purpose is only to balance the instability of the large language model when giving responses, so that the cost of the detection method of the present invention is greatly reduced.

[0077] The two-level security evaluator first allows the model under test to construct a detection knowledge graph (the first security evaluator), and then checks whether its textual description contains dangerous content (the second security evaluator). It can not only quickly intercept failed attempts that cannot even generate a graph at the structural level, but also prevent seemingly successful but actually harmless outputs at the semantic level. While ensuring the accuracy of the evaluation, by setting failure thresholds (such as M, N ≈ 5) at each level to terminate ineffective loops in advance, the computational and time costs of repeatedly executing the complete detection process are greatly reduced, and hierarchical feedback also facilitates accurate positioning and optimization of prompt words or model configurations.

[0078] Embodiment 2 This Embodiment 2 describes a large language model security detection system based on the automatic generation of a knowledge graph. This system is based on the same inventive concept as the large language model security detection method based on the automatic generation of a knowledge graph in Embodiment 1 above.

[0079] Specifically, a large language model security detection system based on the automatic generation of a knowledge graph includes the following modules: A large language model security detection system based on the automatic generation of a knowledge graph includes the following modules: A preprocessing module for preprocessing a data set containing different dangerous prompt words in the security detection direction, and replacing the dangerous behaviors in the initial prompt words of the data set with low-resource languages to obtain rewritten prompt words; A detection knowledge graph generation module for constructing a security detection knowledge graph prompt word template and embedding the rewritten prompt words, and using the large language model to automatically explore the dangerous knowledge encoded therein and generate a complete detection knowledge graph for the initial prompt words; A first security evaluation module for judging whether the security protection of the large language model under test can be bypassed by asking whether the large language model under test has successfully constructed a complete detection knowledge graph for the initial prompt words; If the detection knowledge graph is successfully constructed, natural language conversion is performed; otherwise, the detection knowledge graph is reconstructed, and when the number of times of rejecting the construction of the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed; A natural language conversion module, configured to design a knowledge graph to a text prompt template, directly nest the detection knowledge graph into the knowledge graph to the text prompt template, and convert the structured information in the generated detection knowledge graph into natural language text; And a second security evaluation module, which determines whether the security protection of the large language model is bypassed again by asking whether the text to be evaluated, that is, the natural language text after conversion, contains dangerous information; If it contains dangerous information, it is determined that the security protection of the large language model is successfully bypassed; Otherwise, the detection knowledge graph is regenerated, and when the number of times that the generated content of the response of the tested large language model does not contain dangerous information exceeds the second preset threshold, it is determined that the security protection of the tested large language model cannot be bypassed.

[0080] It should be noted that in the large language model security detection system described in this embodiment, the implementation processes of the functions and roles of each functional module are specifically detailed in the implementation processes of the corresponding steps in the method in Embodiment 1 above, and will not be repeated here.

[0081] Embodiment 3 This Embodiment 3 describes a computer device, which includes a memory and one or more processors. Executable code is stored in the memory, and when the processor executes the executable code, it is used to implement the steps of the large language model security detection method based on knowledge graph automation generation in Embodiment 1 above.

[0082] The computer device in this embodiment is any device or apparatus with data processing capabilities, which will not be elaborated here.

[0083] Embodiment 4 This Embodiment 4 describes a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it is used to implement the steps of the large language model security detection method based on knowledge graph automation generation in Embodiment 1 above.

[0084] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device.

[0085] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. A large language model security detection method based on automatic generation of knowledge graph, characterized in that: The steps include: Step 1. Preprocess the data set containing different danger prompt words in the safety detection direction, replace the dangerous behaviors in the initial prompt words of the data set with low-resource languages, and obtain the rewritten prompt words; Step 2. Construct a safety detection knowledge graph prompt word template and embed the rewritten prompt word, use the large language model to automatically explore the dangerous knowledge encoded in it and generate a complete detection knowledge graph about the initial prompt word; Step 3. Design the first security evaluator to determine whether the security protection of the large language model being tested can be bypassed by asking whether the large language model being tested successfully builds a complete detection knowledge graph about the initial prompt word; If the detection knowledge graph is successfully constructed, continue to step 4; otherwise, return to step 2 and repeat the execution, and when the number of rejections of constructing the detection knowledge graph exceeds the first preset threshold, it is determined that the security protection of the large language model cannot be bypassed; Step 4. Design a prompt word template for converting the knowledge graph into text, directly embed the detection knowledge graph into the knowledge graph-to-text prompt word template, and convert the structured information in the generated detection knowledge graph into natural language text; Step 5. Design a second security evaluator to determine whether the security protection of the large language model is bypassed by asking whether the evaluated text, i.e., the converted natural language text, contains dangerous information; If the converted text contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed; Otherwise, return to step 2 and repeat the process. When the number of times that the large language model under test responds that the generated content does not contain dangerous information exceeds a second preset threshold, it is determined that the security protection of the large language model under test cannot be bypassed.

2. The large language model security detection method based on automatic generation of knowledge graph according to claim 1 is characterized in that: In step 1, the data set uses the AdvBench data set or the HEx-PHI data set.

3. The large language model security detection method based on automatic generation of knowledge graph according to claim 1 is characterized in that: In step 1, the process of replacing the dangerous behavior in the initial prompt words of the data set with low-resource language is as follows: Define the initial prompt word , after the dangerous behavior is replaced, the rewrite prompt word is obtained: ; in, Represents tokenization , Represents low-resource language tokens; To simplify the representation, use Represents the remaining part of the initial prompt word, using Represents the part replaced with low-resource language, so the rewrite prompt word obtained after the dangerous behavior is replaced Also expressed as: ; in, Represents the concatenation process of the remaining initial prompt words and low-resource language prompts in the initial prompt words.

4. The large language model security detection method based on automatic generation of knowledge graph according to claim 1 is characterized in that: In step 2, the generation process of the detection knowledge graph is as follows: Design a security detection knowledge graph prompt word template, where the knowledge graph is represented as <head entity (h), relationship (r), tail entity (t)>; given the head entity h and relationship r, use the prompt model to automatically complete the blank tail entity; The rewritten prompt word is embedded into the prompt word template of the security detection knowledge graph, and then the complete prompt word is input into the large language model being tested to obtain the response of the large language model, that is, the large language model completes and generates a complete detection knowledge graph.

5. The large language model security detection method based on automatic generation of knowledge graph according to claim 1 is characterized in that: In step 3, the first security evaluator is a large language model; Due to the instability of the response of the large language model, that is, when facing the same input, the multiple outputs of the model are not consistent, it is necessary to set a first preset threshold M in the security detection and evaluation process to limit the maximum number of detections; Define loop variable m, the initial value of m is 0; If the large language model being tested successfully builds a complete detection knowledge graph about the initial prompt word, go directly to step 4; if the model refuses to build a detection knowledge graph, repeat step 2 for a new round of detection and update the value of m, setting m=m+1; When the number of times m that the model refuses to build a detection knowledge graph exceeds the first preset threshold M, it means that the current nested prompt word cannot bypass the security protection of the large language model being tested, that is, the security detection fails.

6. The large language model security detection method based on automatic generation of knowledge graph according to claim 1 is characterized in that: In step 4, the process of converting the structured information in the detection knowledge graph into natural language text is as follows: First, a prompt word template is designed to convert the knowledge graph into text. In the prompt word template, a short paragraph is used to describe the detailed content of the generated detection knowledge graph. The detection knowledge graph is directly embedded into the prompt word template, and the nodes, relationships and their attributes are extracted from the structured information in the detection knowledge graph, and converted into logically fluent and grammatically correct natural language text using natural language generation technology.

7. The large language model security detection method based on automatic generation of knowledge graph according to claim 1 is characterized in that: In step 2, the process of constructing the detection knowledge graph by the large language model is expressed as follows: ; in, It is a detection knowledge graph constructed by a large language model. It is a prompt word template. Representing a large language model As the output of the test model, Indicates rewriting of the prompt word; In step 3, the formula for converting the structured information in the detection knowledge graph into natural language text is expressed as follows: ; in, About Details of It is a prompt word template for converting knowledge graph into text. It's a big model The response of the detection knowledge graph to the input when acting as a text converter.

8. The large language model security detection method based on automatic generation of knowledge graph according to claim 1 is characterized in that: In step 5, the second security evaluator is a large language model; Due to the instability of the response of the large language model, that is, when faced with the same input, the model's multiple outputs are inconsistent, it is necessary to set a second preset threshold N in the security detection and evaluation process to limit the maximum number of detections; Define loop variable n and let n start counting from 0; If the detection is successful, it means that the method successfully bypasses the security protection of the large model being tested; If the evaluation result of the tested large language model in step 5 is that it does not contain dangerous information, it means that this security test has failed, and it returns to step 2 for a new round of testing, and updates the value of n, setting n=n+1; When the number of times the tested large language model generates information that does not contain dangerous information exceeds N, it means that the currently nested prompt words cannot bypass the security protection of the tested large language model, and it is determined that the security detection has failed.

9. A large language model security detection system based on automatic generation of knowledge graph, characterized in that: Includes the following modules: A preprocessing module is used to preprocess the data set containing different danger prompt words in the safety detection direction, and replace the dangerous behaviors in the initial prompt words of the data set with low-resource languages ​​to obtain rewritten prompt words; The detection knowledge graph generation module is used to construct a safety detection knowledge graph prompt word template and embed the rewritten prompt words, using the large language model to automatically explore the dangerous knowledge encoded in it and generate a complete detection knowledge graph about the initial prompt words; The first security assessment module is used to determine whether the security protection of the tested large language model can be bypassed by asking whether the tested large language model successfully builds a complete detection knowledge graph about the initial prompt word; If the detection knowledge graph is successfully constructed, natural language conversion is performed; otherwise, the detection knowledge graph is reconstructed, and when the number of rejections for constructing the detection knowledge graph exceeds a first preset threshold, it is determined that the security protection of the large language model cannot be bypassed; The natural language conversion module is used to design the knowledge graph to text prompt word template, directly embed the detection knowledge graph into the knowledge graph to text prompt word template, and convert the structured information in the generated detection knowledge graph into natural language text; The second security assessment module determines whether the security protection of the large language model is bypassed by asking whether the evaluated text, i.e., the converted natural language text, contains dangerous information; If it contains dangerous information, it is determined that the security protection of the large language model has been successfully bypassed; Otherwise, the detection knowledge graph is regenerated, and when the number of times that the large language model being tested responds that the generated content does not contain dangerous information exceeds a second preset threshold, it is determined that the security protection of the large language model being tested cannot be bypassed.

10. A computer device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, the steps of the large language model security detection method automatically generated based on the knowledge graph are implemented as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Tobacco enterprise intelligent information question and answer method based on knowledge graph and large language model

    CN117216227A

  • Method and device for optimizing scene generation model, storage medium and electronic device

    CN118051625A

  • Safety evaluation method based on large language model and related device

    CN119357021A

  • Short message AI training model intelligent matching method and system

    CN119939199A