Large model safety protection method based on thinking chain

By applying zero-sample thinking chain technology to generate security suffixes in the big model and automatically analyzing user input logic, the problem that existing technology is difficult to effectively resist jailbreak attacks is solved, and efficient and flexible security protection is achieved.

CN119989408APending Publication Date: 2025-05-13HARBIN INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510062744.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing large-model security protection technologies are difficult to effectively resist jailbreak attacks, and the traditional security suffix method is poor in flexibility and consumes high computing resources.

Method used

Using technology based on zero-sample thinking chain, the big model's own language understanding ability is used to generate security suffixes, and guide the big model to automatically analyze the logic and intentions behind user input, thereby enhancing security protection capabilities.

Benefits of technology

The optimal security protection effect is achieved under a variety of jailbreak attack methods, reducing computing costs and resource consumption, and improving the security and flexibility of the large model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989408A_ABST
    Figure CN119989408A_ABST
Patent Text Reader

Abstract

The invention relates to a thinking chain-based large model security protection method, which is suitable for enhancing the defense capability of various large language models and does not need extra post-training overhead. The invention relates to the technical field of large model security protection, and ensures that a safe reply is generated through prompt word enhancement of a large language model security defense system; dealing with a jail break attack based on a security defense suffix of a zero sample thinking chain; and the security of the large language model is evaluated by calculating the success rate of prison break attacks. The large language model safety protection method based on the thinking chain comprises two parts, namely a safety system cue word and a zero sample thinking chain. According to the method, extra calculation cost is not introduced, the inference capability of the large language model is fully utilized to resist jailbreak attacks, the safety protection capability of the large language model is greatly enhanced, and stable operation and safe use of the large model in different application scenes are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of large model safety protection, which is a large model safety protection method based on thought chain. Background Art

[0002] Large model security refers to a series of technologies used to ensure the security and reliability of large language models and prevent the output of harmful content during their development, deployment, and application. The essence of large model security is a trust issue. Due to the black box nature of large models, the system has uncontrollable output results, resulting in unexpected behaviors. Such as generating false information and privacy leaks. The root cause of security risks in large language models is that harmful data is mixed in the pre-training process, and it is impossible to remove harmful content in massive pre-training data through manual screening or automated solutions. Therefore, security protection technology is needed to enhance the robustness of large language models.

[0003] At present, the commonly used security protection technology is mainly based on alignment. Alignment technology mainly uses instruction fine-tuning and reinforcement learning based on human feedback to make the output of the large model closer to human preferences and values. This process can greatly improve the security of the model, but it also consumes a lot of computing resources and manpower costs. And because of the diversity and complexity of jailbreak attacks, conventional alignment technology is not enough to resist jailbreak attacks, and additional defense methods are needed to prevent large models from outputting harmful content. A simple and efficient defense method is to add security suffixes to user input to guide the large model to output safe and controllable content, but most of the traditional security suffix methods are based on manual design and have poor flexibility. The present invention uses the thinking chain technology to allow the large model to automatically generate security suffixes, which improves flexibility. At the same time, it is also better than the traditional security suffix method in terms of security. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a large model security protection method based on thinking chain. In order to cope with complex jailbreak attacks and reduce the cost of additional training, the present invention utilizes the powerful language understanding ability of the large model itself, uses zero-sample thinking chain technology to generate security suffixes and then guides the large model to automatically analyze the logic and intention behind the user input. This method has achieved the best security protection effect under a variety of jailbreak attack methods, and the increase in computing cost and the impact on general capabilities are relatively small.

[0005] The present invention provides the following technical solutions:

[0006] A large model security protection method based on thought chain, the method comprising the following steps:

[0007] Step 1: Enhance the prompt words of the large language model security defense system to ensure that safe and expected responses are generated;

[0008] Step 2: Apply security defense suffixes based on zero-sample thinking chains to counter various jailbreak attacks against large language models;

[0009] Step 3: Measure and improve the security of the large language model by evaluating the success rate of jailbreak attacks.

[0010] Preferably, the step 1 is specifically:

[0011] Set up the big language model as a calm assistant with judgment;

[0012] Tip: Large models require a comprehensive understanding of the internal logic of the input content;

[0013] Feel free to respond to harmless input.

[0014] Preferably, the formal expression of the system prompt is in P sys The i-th token of

[0015] The system prompts P sys With user input The concatenation forms the complete input of the large language model, which is then input into the large model for autoregressive inference, namely:

[0016]

[0017] The above formula is given by the system prompt P sys and user input P usr After the large language model samples responses from the distribution q(·), a security defense suffix based on zero-sample thought chain reasoning is constructed to enhance the security performance of the large language model based on the modification of the system prompt words.

[0018] Preferably, the step 2 is specifically:

[0019] Analysis phase: Through zero-sample thought chain prompts, the large language model is guided to analyze the logic of the input content without having to determine the specific jailbreak attack type in advance, and it can autonomously parse and identify possible jailbreak attacks;

[0020] Integration stage: Combine the model’s logical analysis results in the first stage with the zero-sample thinking chain prompt words to form a complete input, which is then processed by the large language model.

[0021] Preferably, the zero-sample thinking chain prompt words include three parts: role setting, scenario setting, and task definition. The role setting sets the large language model as an assistant that calmly analyzes the input text; the scenario setting prompts the large model that the input content may contain jailbreak attack content, and the large model is required to analyze the jailbreak content; the task definition requires the large language model not to follow the user's jailbreak input content, so as to only analyze the logic behind the user input.

[0022] Preferably, the jailbreak attack success rate is expressed by the following formula:

[0023]

[0024] There are also multiple ways to evaluate whether a jailbreak attack is successful, including manual evaluation by labelers and automatic evaluation by large language models.

[0025] Preferably, a RoBERTa model is trained to automatically evaluate the security of the output of the large language model.

[0026] A large model safety protection system based on thought chain, the system comprising:

[0027] An enhancement module, wherein the enhancement module performs large language model security defense system prompt word enhancement to ensure generation of a safe response;

[0028] A defense module, which applies a security defense suffix based on a zero-sample thinking chain to deal with various jailbreak attacks;

[0029] An evaluation module measures and improves the security of a large language model by evaluating the success rate of jailbreak attacks.

[0030] A computer-readable storage medium stores a computer program, which is executed by a processor to implement a large model security protection method based on a thought chain.

[0031] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements a large model security protection method based on a thought chain when executing the computer program.

[0032] The present invention has the following beneficial effects:

[0033] Compared with the prior art, the present invention has the following advantages:

[0034] The present invention uses the thinking chain technology to allow the large model to automatically generate security suffixes, which improves flexibility and is also more secure than the traditional security suffix method. In order to deal with complex jailbreak attacks and reduce the cost of additional training, the present invention uses the large model's own powerful language understanding ability, uses zero-sample thinking chain technology to generate security suffixes, and then guides the large model to automatically analyze the logic and intention behind the user input. This method has achieved the best security protection effect under a variety of jailbreak attack methods, and the increase in computing cost and the impact on general capabilities are relatively small.

[0035] The large language model security protection method based on thought chain adopted by the present invention includes two parts: security system prompt words and zero-sample thought chain. The present invention does not introduce additional computing costs, but makes full use of the reasoning ability of the large language model to resist jailbreak attacks, greatly enhancing the security protection ability of the large language model.

[0036] Compared with the traditional method of fine-tuning based on large-scale data or human feedback reinforcement learning, the zero-sample thinking chain security defense strategy used in this invention significantly reduces the consumption of computing resources. Through a simple and efficient prompt word modification scheme, this method has shown excellent security protection effects in a variety of jailbreak attack scenarios, proving its good generalization and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0038] Figure 1 Shown is a schematic diagram of the overall structure of the system of the present invention;

[0039] Figure 2 Displayed is a schematic diagram of the security defense system prompt words of the present invention;

[0040] Figure 3 Shown is a schematic diagram of the zero-sample thought chain prompt words of the present invention;

[0041] Figure 4 Shown is a schematic diagram of an application example of the zero-sample thinking chain of the present invention. DETAILED DESCRIPTION

[0042] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] The present invention is described in detail below in conjunction with specific embodiments. Specific embodiment one:

[0045] according to Figures 1 to 4 As shown, the specific optimization technical solution adopted by the present invention to solve the above-mentioned technical problems is: the present invention relates to a large model security protection method based on thinking chain.

[0046] A large model security protection method based on thought chain, the method comprising the following steps:

[0047] Step 1: Enhance the prompt words of the large language model security defense system to ensure that safe and expected responses are generated;

[0048] Step 2: Apply a security defense suffix mechanism based on zero-sample thinking chain to deal with various jailbreak attacks on large language models;

[0049] Step 3: Measure and improve the security of the large language model by evaluating the success rate of jailbreak attacks.

[0050] In order to deal with complex jailbreak attacks and reduce the cost of additional training, the present invention utilizes the powerful language understanding ability of the large model itself, uses zero-sample thinking chain technology to generate security suffixes, and then guides the large model to automatically analyze the logic and intention behind the user input. This method has achieved the best security protection effect under a variety of jailbreak attack methods, and the increase in computational cost and the impact on general capabilities are relatively small. Specific embodiment 2:

[0052] The difference between the second embodiment of the present application and the first embodiment is that:

[0053] The step 1 is specifically as follows:

[0054] Set up the big language model as a calm assistant with judgment;

[0055] Tip: Large models require a comprehensive understanding of the internal logic of the input content;

[0056] Feel free to respond to harmless input. Specific embodiment three:

[0058] The difference between the third embodiment of the present application and the second embodiment is that:

[0059] The system prompts are formally expressed as in P sys The i-th token of

[0060] The system prompts P sys With user input The complete input of the large language model is obtained by splicing them together, and the large model is input for autoregressive inference, that is:

[0061]

[0062] The above formula is given by the system prompt P sys and user input P usr After the large language model samples responses from the distribution q(·), a security defense suffix based on zero-sample thought chain reasoning is constructed to enhance the security performance of the large language model based on the modification of the system prompt words. Specific embodiment four:

[0064] The difference between the fourth embodiment of the present application and the third embodiment is that:

[0065] The step 2 is specifically as follows:

[0066] Analysis phase: Through zero-sample thought chain prompts, the large language model is guided to analyze the logic of the input content, so that possible jailbreak attacks can be autonomously parsed and identified without having to determine the specific jailbreak attack type in advance;

[0067] Integration stage: Combine the model’s logical analysis results in the first stage with the zero-sample thinking chain prompt words to form a complete input, which is then processed by the large language model. Specific embodiment five:

[0069] The difference between the fifth embodiment of the present invention and the fourth embodiment is that:

[0070] The zero-sample thinking chain prompts include three parts: role setting, scenario setting, and task definition. The role setting sets the large language model as an assistant that calmly analyzes the input text; the scenario setting prompts the large model that the input content may contain jailbreak attack content, and the large model needs to analyze the jailbreak content; the task definition requires the large language model not to follow the user's jailbreak input content, and only analyze the logic behind the user input. Specific embodiment six:

[0072] The difference between the sixth embodiment of the present invention and the fifth embodiment is that:

[0073] The success rate of jailbreak attack is expressed as follows:

[0074]

[0075] There are also multiple ways to evaluate whether a jailbreak attack is successful, including manual evaluation by labelers and automatic evaluation by large language models. Specific embodiment seven:

[0077] The difference between the seventh embodiment of the present invention and the sixth embodiment is that:

[0078] Train a RoBERTa model to automatically evaluate the security of the output of a large language model. Specific embodiment eight:

[0080] The difference between the eighth embodiment of the present invention and the seventh embodiment is that:

[0081] The present invention provides a large model safety protection system based on thought chain, the system comprising:

[0082] An enhancement module, wherein the enhancement module performs large language model security defense system prompt word enhancement to ensure generation of a safe response;

[0083] A defense module, wherein the defense module is based on a security defense suffix of a zero-sample thinking chain to counter jailbreak attacks;

[0084] An evaluation module is used to evaluate the security of a large language model by using a success rate of a jailbreak attack. Specific embodiment nine:

[0086] The difference between the ninth embodiment of the present invention and the eighth embodiment is that:

[0087] The present invention provides a computer-readable storage medium on which a computer program is stored. The program is executed by a processor to implement a large model security protection method based on a thinking chain. Specific embodiment ten:

[0089] The difference between the tenth embodiment of the present invention and the ninth embodiment is that:

[0090] The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements a large model security protection method based on a thought chain when executing the computer program. Specific embodiment eleven:

[0092] The difference between the eleventh embodiment of the present invention and the tenth embodiment is that:

[0093] The scheme of the present invention comprises:

[0094] 1. Enhanced prompt words for the large language model security defense system

[0095] System prompts are often used to enhance the security and reliability of large language model responses. In this invention, in order to enhance the large language model's understanding of the intention behind the jailbreak attack and ensure that it generates a safe response, we have developed a set of security defense system prompt word standards. The security defense system prompt words set by this standard must include the following:

[0096] (1) Set up the large language model as a calm assistant with judgment.

[0097] (2) It is suggested that large models need to understand the internal logic of the input content.

[0098] (3) Freedom to respond to harmless input.

[0099] This is an example of a security defense system prompt word.

[0100] In the present invention, the system prompt can be formally expressed as in P sys The i-th token of .

[0101] System prompts P sys With user input The concatenation forms the complete input of the large language model, which is then input into the large model for autoregressive inference, namely:

[0102]

[0103] The above formula is given by the system prompt P sys and user input P usr The purpose of the present invention is to construct a security defense suffix based on zero-sample thought chain reasoning technology to enhance the security performance of the large language model on the basis of modifying the system prompt word.

[0104] 2. Security defense suffix based on zero-sample thinking chain

[0105] In order to better understand and respond to jailbreak attacks, it is crucial to understand how these attacks are successful. The current large model has language comprehension capabilities similar to the level of human intelligence. Therefore, the present invention believes that under a well-designed framework, the large model can detect and understand harmful information and misleading content in the input content. In addition, the present invention improves the existing security suffix technology, which is not general enough and difficult to apply to multiple methods. The use of a large model to analyze the logic behind jailbreak attacks can deal with all jailbreak attacks that rely on text content, and does not rely on the text quality of the manually designed security suffix.

[0106] The zero-shot mind-chaining method uses a two-stage process:

[0107] (1) Analysis phase: Through the zero-sample thinking chain prompt words, the large language model is guided to analyze the logic of the input content, so that possible jailbreak attacks can be autonomously parsed and identified without having to determine the specific jailbreak attack type in advance.

[0108] (2) Integration stage: The model’s logical analysis results in the first stage are combined with the zero-sample thinking chain prompt words to form a complete input, which is further processed by the large language model to reduce the risk of generating harmful content.

[0109] The zero-sample thinking chain prompts designed by the present invention include three parts: role setting, scene setting, and task definition. The role setting sets the large language model as an assistant that calmly analyzes the input text. The scene setting prompts the large model that the input content may contain jailbreak attack content, and the large model needs to analyze the jailbreak content. The task definition requires the large language model not to follow the user's jailbreak input content, so as to only analyze the logic behind the user input. Figure 3 An example of a zero-sample thought chain prompt word

[0110] Compared with traditional methods based on large-scale data fine-tuning or human feedback reinforcement learning, the zero-shot thinking chain security defense strategy significantly reduces the consumption of computing resources. Through a simple and efficient prompt word modification scheme, this method has demonstrated excellent security protection effects in a variety of jailbreak attack scenarios, proving its good generalization and practical application value.

[0111] The large language model security protection method based on thought chain adopted by the present invention includes two parts: security system prompt words and zero-sample thought chain. The present invention does not introduce additional computing costs, but makes full use of the reasoning ability of the large language model to resist jailbreak attacks, greatly enhancing the security protection ability of the large language model.

[0112] The present invention follows the solution of most works and uses the success rate of jailbreak attacks to evaluate the security of large language models.

[0113] The formula for the success rate of jailbreak attacks is as follows:

[0114]

[0115] There are also many ways to evaluate whether a jailbreak attack is successful, such as manual evaluation by annotators and automatic evaluation by a large language model. The present invention uses a simple and efficient method to train a RoBERTa model to automatically perform security evaluation on the output of a large language model.

[0116] Table 1 and Table 2 show the effects of the security defense method used in the present invention compared with the baseline method under various jailbreak attack methods.

[0117] Table 1 Experimental results of jailbreak attack methods based on deception

[0118]

[0119] Table 2 Experimental results of the encoding-based jailbreak attack method

[0120]

[0121] The above is only a preferred implementation of a large model security protection method based on a thinking chain. The protection scope of a large model security protection method based on a thinking chain is not limited to the above embodiments. All technical solutions under this idea belong to the protection scope of the present invention. It should be pointed out that for those skilled in the art, several improvements and changes without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A large model security protection method based on thought chain, characterized by: The method comprises the following steps: Step 1: Enhance the prompt words of the large language model security defense system to ensure that safe and expected responses are generated; Step 2: Apply a security defense suffix mechanism based on zero-sample thinking chain to deal with various jailbreak attacks on large language models; Step 3: Measure and improve the security of the large language model by evaluating the success rate of jailbreak attacks.

2. The method according to claim 1, characterized in that: The step 1 is specifically as follows: Set up the big language model as a calm assistant with judgment; Tip: Large models require a comprehensive understanding of the internal logic of the input content; Feel free to respond to harmless input.

3. The method according to claim 2, characterized in that: The formal expression of the system prompt is in P sys The i-th token of The system prompts P sys With user input The concatenation forms the complete input of the large language model, which is then input into the large model for autoregressive inference, namely: The above formula is given by the system prompt P sys and user input P usr After the large language model samples responses from the distribution q(·), a security defense suffix based on zero-sample thought chain reasoning is constructed to enhance the security performance of the large language model based on the modification of the system prompt words.

4. The method according to claim 1, characterized in that: The step 2 is specifically as follows: Analysis phase: Through zero-sample thought chain prompts, the large language model is guided to analyze the logic of the input content, so that possible jailbreak attacks can be autonomously parsed and identified without having to determine the specific jailbreak attack type in advance; Integration stage: Combine the model’s logical analysis results in the first stage with the zero-sample thinking chain prompt words to form a complete input, which is then processed by the large language model.

5. The method according to claim 4, characterized in that: The zero-sample thinking chain prompts include three parts: role setting, scene setting, and task definition, among which: Role setting: the large language model is set as an assistant that calmly analyzes the input text; Scenario setting, prompting that the input content may contain jailbreak attack content, requiring a large model to analyze the jailbreak content; The task definition requires the large language model not to follow the content of the user's jailbreak input, but only to analyze the logic behind the user input.

6. The method according to claim 1, characterized in that: The success rate of jailbreak attack is expressed as follows: There are also multiple ways to evaluate whether a jailbreak attack is successful, including manual evaluation by labelers and automatic evaluation by large language models.

7. The method according to claim 6, characterized in that: Train a RoBERTa model to automatically evaluate the security of the output of a large language model.

8. A large model security protection system based on thought chain, characterized by: The system comprises: An enhancement module, wherein the enhancement module performs large language model security defense system prompt word enhancement to ensure generation of a safe response; A defense module, wherein the defense module is based on a security defense suffix of a zero-sample thinking chain to counter jailbreak attacks; An evaluation module evaluates the security of a large language model by evaluating the success rate of a jailbreak attack.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method of claims 1-7.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method of claims 1-7 is implemented.

Citation Information

Cited By

  • Large model jailbreak attack security assessment method and device based on role simulation

    CN120579192A

  • Role simulation-based large model jailbreaking attack security assessment method and device

    CN120579192B

  • Security control method and device, equipment, storage medium and computer program product

    CN122021893A