A shadow large language model-based jailbreak attack defense method, system, medium and device
Patent Information
- Application Number
- CN202610875507.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-10-09
AI Technical Summary
[0005]本发明所要解决的技术问题是,提供一种基于影子大语言模型的越狱攻击防御方法、系统、介质及设备,解决现有技术难以兼顾防御效果、延迟开销与模型兼容性的问题
[0038](1)通过训练好的防御大语言模型和预设的检测提示词对用户查询请求进行有害意图的检测,有效防御了包括基于优化、基于生成、间接攻击和多语言攻击在内的多种主流越狱手段;
Smart Images

Figure CN122885901A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a jailbreak attack defense method, system, medium, and device based on a shadow big language model, belonging to the field of artificial intelligence security technology. Background Technology
[0002] With the widespread application of large language models in natural language processing, information retrieval, and image generation, the security of these models has become an increasingly important concern. To prevent models from generating harmful, illegal, or unethical content, vendors typically perform security alignment on the models. However, jailbreak attacks bypass these security mechanisms by designing specific prompts, inducing the model to generate harmful content that should be rejected. Existing jailbreak attacks have evolved from manually designed prompts to increasingly sophisticated methods, including optimized GCG (Greedy Coordinate Gradient) attacks, generational PAIR (Prompt Automatic Iterative Refinement) attacks, indirect attacks, and multilingual attacks.
[0003] Existing defense mechanisms are mainly divided into model-based defense and plugin-based defense. Model-based defense usually requires fine-tuning of model parameters or modification of internal mechanisms. While this can improve robustness, it is often costly and unsuitable for closed-source models. Plugin-based defense, although usable as external plugins, typically cannot handle all types of jailbreak attacks, especially indirect and multilingual attacks, and often introduces significant latency, impacting user experience. Furthermore, many existing solutions fail to provide an explainable rationale for their defenses, making attribution analysis difficult in real-world deployments.
[0004] Therefore, the main problem facing existing technologies is the lack of a universal defense solution that can handle all mainstream jailbreak attacks and be compatible with both open-source and closed-source large language models while introducing negligible latency. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a jailbreak attack defense method, system, medium and device based on the shadow big language model, so as to solve the problem that the existing technology is difficult to balance defense effect, latency overhead and model compatibility.
[0006] To achieve the above objectives, the present invention employs the following technical solution:
[0007] Firstly, this invention proposes a jailbreak attack defense method based on a shadow large language model, comprising:
[0008] Get the user's query request;
[0009] The user query request is distributed to the target large language model in the normal stack and the pre-trained defensive large language model in the shadow stack;
[0010] The target large language model is used to process the user query request and generate a response output. The defense large language model is used in combination with preset detection prompt words to detect the user query request for malicious intent and generate detection results.
[0011] Based on the detection results, it is determined whether a jailbreak attack exists. If the jailbreak attack exists, the target large language model is prevented from transmitting the response output to the user; if the jailbreak attack does not exist, the target large language model is allowed to transmit the response output to the user.
[0012] Furthermore, the step of using the defense big language model in conjunction with preset detection prompts to detect malicious intent in the user query request and generate detection results includes:
[0013] Obtain preset detection prompts, including direct detection prompts and intent detection prompts;
[0014] The direct detection prompts the defense big language model to determine whether there are text fragments in the user query request that violate the preset security policy;
[0015] The intent detection prompts the defense big language model to summarize the user intent based on the user query request, and determine whether the user intent contains harmful content.
[0016] If a text fragment violates the preset security policy or contains harmful content, output a detection result indicating a jailbreak attack has occurred. If no text fragment violates the preset security policy and does not contain harmful content, output a detection result indicating no jailbreak attack has occurred.
[0017] Furthermore, the training process of the defense large language model includes:
[0018] Obtain a pre-created red team test dataset containing both harmful and harmless instructions;
[0019] The red team test dataset is processed using a general-purpose large language model with defensive capabilities, combined with the detection prompt words, to generate label data containing the detection results;
[0020] The existing open-source basic model is trained based on the labeled data to obtain a defense large language model.
[0021] Furthermore, the step of training an existing open-source basic model based on the labeled data to obtain a defensive large language model includes:
[0022] An open-source base model is obtained, and a parameter-efficient fine-tuning technique is used. The model parameters are incrementally updated by adding a low-rank matrix to the weights of the open-source base model using the label data, resulting in a trained defense large language model.
[0023] Furthermore, the step of determining whether a jailbreak attack exists based on the detection results includes:
[0024] If the detection result contains a specific marker indicating security, it is determined that there is no jailbreak attack.
[0025] If the detection results contain text fragments that indicate harmfulness, it is determined that a jailbreak attack exists.
[0026] Furthermore, the detection result is a short text shorter than a preset length, and the judgment time of the detection result is much shorter than the generation time of the response output.
[0027] Furthermore, after the target large language model generates the response output, the response output is temporarily stored in a buffer, waiting for the signal of the intercept checkpoint or the release checkpoint;
[0028] If the jailbreak attack exists, an interception checkpoint is triggered to prevent the target large language model from transmitting the response output to the user.
[0029] If the jailbreak attack does not exist, a release checkpoint is triggered, allowing the target large language model to transmit the response output to the user.
[0030] Secondly, this invention proposes a jailbreak attack defense system based on a shadow large language model, comprising:
[0031] The information acquisition module is used to acquire user query requests;
[0032] The user request sending module is used to distribute the user query request to the target large language model in the normal stack and the pre-trained defense large language model in the shadow stack.
[0033] The request response module is used to process the user query request using the target big language model to generate a response output, and to use the defense big language model in combination with preset detection prompts to detect the user query request for malicious intent and generate detection results.
[0034] The jailbreak attack defense module is used to determine whether a jailbreak attack exists based on the detection results. If the jailbreak attack exists, it prevents the target large language model from transmitting the response output to the user; if the jailbreak attack does not exist, it allows the target large language model to transmit the response output to the user.
[0035] Thirdly, the present invention also discloses a computer-readable storage medium for storing one or more programs, said one or more programs including instructions that, when executed by a computing device, cause the computing device to perform the method described in the first aspect.
[0036] Fourthly, the present invention also discloses a computer device comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method described in the first aspect.
[0037] The beneficial effects achieved by this invention are as follows:
[0038] (1) By using a well-trained defense big language model and preset detection prompts, the malicious intent of user query requests is detected, effectively defending against a variety of mainstream jailbreaking methods, including optimization-based, generation-based, indirect attack and multi-language attack.
[0039] (2) The defensive big language model used in this invention is independent of the target big language model and is applicable to most open source or closed source LLMs. It has high versatility and portability. The defensive big language model used in this invention is deployed in the shadow stack and runs synchronously with the target big language model. It does not contain additional latency and does not affect the original performance of the target big model.
[0040] (3) When an attack is detected, the system can output specific harmful fragments or intentions, providing an interpretable basis for defense, which facilitates the developer to conduct source tracing analysis and continuously optimize the existing large model. Attached Figure Description
[0041] Figure 1 A flowchart of the jailbreak attack defense method based on the shadow large language model provided by the present invention;
[0042] Figure 2 This is a structural diagram of the detection prompt words in this invention. Detailed Implementation
[0043] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0044] Example 1, as Figure 1 As shown, this invention discloses a jailbreak attack defense method based on a shadow large language model, comprising:
[0045] Get the user's query request;
[0046] The user query request is distributed to the target large language model in the normal stack and the pre-trained defensive large language model in the shadow stack;
[0047] The target large language model is used to process the user query request and generate a response output. The defense large language model is used in combination with preset detection prompt words to detect the user query request for malicious intent and generate detection results.
[0048] Based on the detection results, it is determined whether a jailbreak attack exists. If the jailbreak attack exists, the target large language model is prevented from transmitting the response output to the user; if the jailbreak attack does not exist, the target large language model is allowed to transmit the response output to the user.
[0049] The core of this invention lies in introducing the concept of shadow stack in computer system security, establishing a shadow large language model as a defense instance, which works in parallel with the target large language model in the normal stack, without containing additional latency, and has high versatility and portability.
[0050] The step of using the defense big language model in conjunction with preset detection prompts to detect malicious intent in the user query request and generate detection results includes:
[0051] Obtain preset detection prompts, including direct detection prompts and intent detection prompts;
[0052] The direct detection prompts the defense big language model to determine whether there are text fragments in the user query request that violate the preset security policy;
[0053] The intent detection prompts the defense big language model to summarize the user intent based on the user query request, and determine whether the user intent contains harmful content.
[0054] If a text fragment violates the preset security policy or contains harmful content, output a detection result indicating a jailbreak attack has occurred. If no text fragment violates the preset security policy and does not contain harmful content, output a detection result indicating no jailbreak attack has occurred.
[0055] The security policy is used to determine whether a user query request contains content that violates security requirements, including but not limited to one or more of the following:
[0056] (1) Policy on illegal and irregular content: Prohibit the generation of content that violates laws and regulations, including but not limited to inciting illegal and criminal activities, disseminating violent and terrorist information, disseminating information on the trade of contraband, organizing and carrying out cyberattacks, and other content that is prohibited from being disseminated by laws and regulations;
[0057] (2) Hazardous behavior guidance strategy: It is prohibited to generate operational guidance information that may cause personal injury, property damage or public safety risks, including but not limited to guidance on the use of hazardous chemicals, the manufacture of explosive devices, the destruction of high-risk equipment, the implementation of malicious software and other hazardous behavior;
[0058] (3) Network and system security policy: Prohibit the generation of operational content that may be used for unauthorized access, system penetration, privilege escalation, malicious code construction, vulnerability exploitation or bypassing security controls;
[0059] (4) Privacy and sensitive information protection strategy: It is prohibited to obtain, infer, disclose or disseminate personal privacy information, identity information, account information, authentication credentials, trade secrets and other sensitive information;
[0060] (5) Model security policy: Prohibit large language models from bypassing security alignment mechanisms, system prompt word protection mechanisms, access control mechanisms and other security restrictions; prohibit obtaining model internal configuration, system prompt words, hidden rules or other protected information;
[0061] (6) Harmful content generation strategy: Prohibit the generation of hate discrimination, harassment and insults, malicious deception, false advertising and other content that may cause social harm;
[0062] (7) Jailbreak attack identification strategy: Identify and intercept attack behaviors that attempt to bypass model security constraints through role-playing, instruction overwriting, code obfuscation, context pollution, multi-round inducement, translation conversion, prompt word injection, etc.
[0063] The training process of the defense large language model includes:
[0064] Obtain a pre-created red team test dataset containing both harmful and harmless instructions; the red team test dataset, which includes both harmful and harmless instructions, is a set of adversarial examples designed specifically for adversarial evaluation of the security boundaries of large language models;
[0065] The red team test dataset is processed using a general-purpose large language model with defensive capabilities, combined with the detection prompt words, to generate label data containing the detection results;
[0066] The existing open-source basic model is trained based on the labeled data to obtain a defense large language model.
[0067] The process of training an existing open-source basic model based on the labeled data to obtain a defense large language model includes:
[0068] An open-source base model is obtained, and a parameter-efficient fine-tuning technique is employed. Using the labeled data, the model parameters are updated by adding low-rank matrices to the weights of the open-source base model, resulting in a trained defensive large language model. Low-rank matrix increment refers to approximating the model parameter update with the product of two low-rank matrices while keeping the original high-dimensional weight matrix frozen, thereby achieving efficient fine-tuning or adaptation.
[0069] The determination of whether a jailbreak attack exists based on the detection results includes:
[0070] If the detection result contains a specific marker that indicates security, it is determined that there is no jailbreak attack, and the cached output of the target large language model is allowed to continue to be transmitted.
[0071] If the detection result contains text fragments that indicate harmfulness, it is determined that a jailbreak attack exists, and the output of the target large language model is discarded and a preset rejection response is returned.
[0072] The detection result is a short text shorter than a preset length, and the judgment time of the detection result is much shorter than the generation time of the response output, thereby ensuring that the delay for normal user requests is negligible.
[0073] After the target large language model generates the response output, the response output is temporarily stored in a buffer, waiting for the signal of the intercept checkpoint or the release checkpoint;
[0074] If the jailbreak attack exists, an interception checkpoint is triggered to prevent the target large language model from transmitting the response output to the user.
[0075] If the jailbreak attack does not exist, a release checkpoint is triggered, allowing the target large language model to transmit the response output to the user.
[0076] In summary, the core of this invention is:
[0077] (1) Dual-stack architecture: User query requests are distributed to the target model and the defense model at the same time. The target model is responsible for the regular dialogue generation, and the defense model is responsible for security detection.
[0078] (2) Dual-mode detection prompts: Two specific detection prompts were designed: direct prompts and intent prompts. Direct prompts are used to directly locate text fragments in user input that violate security policies; intent prompts use thought chain technology to first summarize the user's intent and then determine whether it contains harmfulness, which is specifically used to defend against indirect jailbreak attacks.
[0079] (3) Checkpoint control: The output of the defense model is usually very short, such as Yes or No, which can quickly trigger the checkpoint. If the defense model does not detect the risk, the output of the target model is released normally; if the risk is detected, the system intercepts the output of the target model and returns a rejection response.
[0080] (4) Low-cost model fine-tuning: In order to reduce the dependence on expensive closed-source models, this invention adopts knowledge distillation technology, which can use the high-quality detection data generated by GPT-5 to fine-tune the open-source model, so that it has defense capabilities close to GPT-5, while maintaining low latency. The open-source model can be Qwen-8B.
[0081] Example 2, based on the same inventive concept as Example 1, proposes a jailbreak attack defense method based on the shadow big language model, the specific steps of which are as follows.
[0082] Step S1: System initialization and request reception.
[0083] The target large language model is deployed in the normal stack, and the defensive large language model is deployed in the shadow stack. When a user enters a query request P_query, the request dispatch module immediately copies and distributes the request to the two model instances mentioned above.
[0084] Step S2: Parallel processing and dual-mode detection.
[0085] The target large language model and the defense large language model execute tasks in parallel, and are processed as follows.
[0086] (1) Target model processing: The target large language model processes the user input query request P_query as usual, and generates response content for each token. However, the response is temporarily suspended in the output buffer, waiting for the checkpoint signal. Here, token is a word element, which refers to the basic data unit of the language processed by the large language model.
[0087] (2) Defense model processing: The defense big language model receives the encapsulated detection prompt P_detect and the wrapped P_query, where the detection prompt includes the direct detection prompt Direct Prompt and the intent detection prompt IntentPrompt.
[0088] like Figure 2 As shown, in Direct Prompt mode, the system calls Direct Prompt and issues the instruction: Please identify the parts of the following sentences that may violate the security policy, and the expected output is "Harmful Fragment" or "No".
[0089] In Intent Prompt mode, the system invokes the intent prompt and issues the instruction: Please first summarize the true intent of the following request, and then determine whether the intent contains content that violates the security policy. The expected output is a description of harmful intent or No.
[0090] Step S3: Checkpoint triggering and access control.
[0091] If the defense big language model outputs a short text "No" or a similar security flag, it indicates that no jailbreak intent was detected. The system immediately triggers a release checkpoint, allowing the target big language model's cached response to be sent to the user.
[0092] If the defense big language model outputs specific harmful text fragments or intent descriptions, the system triggers an interception checkpoint. At this time, the system discards the output of the target big language model and returns a preset rejection response template to the user, such as: I cannot fulfill your request because your [harmful fragment] violates the security policy.
[0093] Step S4: Low-cost construction of the defense model.
[0094] To construct an independent defense model, this invention employs the following training process:
[0095] (1) Data distillation: Using a high-performance model as the teacher model, input a query request containing tens of thousands of red team test data, and use the Direct Prompt and Intent Prompt mentioned above to generate high-quality labels; the high-performance model can use the GPT-5 model.
[0096] (2) Model fine-tuning: Obtain the open-source base model and use LoRA (Low-Rank Adaptation) technology for efficient parameter fine-tuning. The training goal is to enable the open-source model to output detection results consistent with the teacher model when receiving user queries.
[0097] (3) Deployment: The fine-tuned model was deployed in the shadow stack as a defense big language model. Experiments showed that the model could achieve a defense effect close to that of GPT-5, and the inference latency was extremely low.
[0098] Example 3, based on the same inventive concept as other examples, proposes a jailbreak attack defense system based on a shadow large language model, comprising:
[0099] The information acquisition module is used to acquire user query requests;
[0100] The user request sending module is used to distribute the user query request to the target large language model in the normal stack and the pre-trained defense large language model in the shadow stack.
[0101] The request response module is used to process the user query request using the target big language model to generate a response output, and to use the defense big language model in combination with preset detection prompts to detect the user query request for malicious intent and generate detection results.
[0102] The jailbreak attack defense module is used to determine whether a jailbreak attack exists based on the detection results. If the jailbreak attack exists, it prevents the target large language model from transmitting the response output to the user; if the jailbreak attack does not exist, it allows the target large language model to transmit the response output to the user.
[0103] Example 4, based on the same inventive concept as other examples, discloses a computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform the method described in Example 1.
[0104] Example 5: The present invention also discloses a computer device, including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method described in Example 1.
[0105] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0106] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0107] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0108] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0109] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A jailbreak attack defense method based on a shadow large language model, characterized in that, include: Get the user's query request; The user query request is distributed to the target large language model in the normal stack and the pre-trained defensive large language model in the shadow stack; The target large language model is used to process the user query request and generate a response output. The defense large language model is used in combination with preset detection prompt words to detect the user query request for malicious intent and generate detection results. Based on the detection results, it is determined whether a jailbreak attack exists. If the jailbreak attack exists, the target large language model is prevented from transmitting the response output to the user. If the jailbreak attack does not exist, the target large language model is allowed to transmit the response output to the user.
2. The jailbreak attack defense method based on the shadow large language model according to claim 1, characterized in that, The step of using the defense big language model in conjunction with preset detection prompts to detect malicious intent in the user query request and generate detection results includes: Obtain preset detection prompts, including direct detection prompts and intent detection prompts; The direct detection prompts the defense big language model to determine whether there are text fragments in the user query request that violate the preset security policy; The intent detection prompts the defense big language model to summarize the user intent based on the user query request, and determine whether the user intent contains harmful content. If a text fragment violates the preset security policy or contains harmful content, output a detection result indicating a jailbreak attack has occurred. If no text fragment violates the preset security policy and does not contain harmful content, output a detection result indicating no jailbreak attack has occurred.
3. The jailbreak attack defense method based on the shadow large language model according to claim 2, characterized in that, The training process of the defense large language model includes: Obtain a pre-created red team test dataset containing both harmful and harmless instructions; The red team test dataset is processed using a general-purpose large language model with defensive capabilities, combined with the detection prompt words, to generate label data containing the detection results; The existing open-source basic model is trained based on the labeled data to obtain a defense large language model.
4. The jailbreak attack defense method based on the shadow large language model according to claim 3, characterized in that, The process of training an existing open-source basic model based on the labeled data to obtain a defense large language model includes: An open-source base model is obtained, and a parameter-efficient fine-tuning technique is used. The model parameters are incrementally updated by adding a low-rank matrix to the weights of the open-source base model using the label data, resulting in a trained defense large language model.
5. The jailbreak attack defense method based on the shadow large language model according to claim 1, characterized in that, The determination of whether a jailbreak attack exists based on the detection results includes: If the detection result contains a specific marker indicating security, it is determined that there is no jailbreak attack. If the detection results contain text fragments that indicate harmfulness, it is determined that a jailbreak attack exists.
6. The jailbreak attack defense method based on the shadow large language model according to claim 1, characterized in that, The detection result is a short text shorter than the preset length, and the judgment time of the detection result is much shorter than the generation time of the response output.
7. The jailbreak attack defense method based on the shadow large language model according to claim 1, characterized in that, After the target large language model generates the response output, the response output is temporarily stored in a buffer, waiting for the signal of the intercept checkpoint or the release checkpoint; If the jailbreak attack exists, an interception checkpoint is triggered to prevent the target large language model from transmitting the response output to the user. If the jailbreak attack does not exist, a release checkpoint is triggered, allowing the target large language model to transmit the response output to the user.
8. A jailbreak attack defense system based on a shadow large language model, characterized in that, include: The information acquisition module is used to acquire user query requests; The user request sending module is used to distribute the user query request to the target large language model in the normal stack and the pre-trained defense large language model in the shadow stack. The request-response module is used to process the user query request using the target big language model to generate a response output, and to use the defense big language model in combination with preset detection prompts to detect the user query request for malicious intent and generate detection results. The jailbreak attack defense module is used to determine whether a jailbreak attack exists based on the detection results. If the jailbreak attack exists, it prevents the target large language model from transmitting the response output to the user. If the jailbreak attack does not exist, the target large language model is allowed to transmit the response output to the user.
9. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods of claims 1 to 7.
10. A computer device, characterized in that, include, One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any of claims 1 to 7.