A Method and System for Generating Test Examples of Large Model Jailbreak Attacks
Through the methods of task decomposition and intent hiddenness, jailbreak test samples are automatically generated, solving the problems of low efficiency of existing methods and insufficient dynamic upgrade capabilities, and achieving more efficient and comprehensive large-scale model security assessment.
Patent Information
- Application Number
- CN202510481810.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing large-model jailbreak attack test sample generation method is inefficient and difficult to fully cover the security boundaries of the large-model. In addition, traditional testing methods lack dynamic upgrade capabilities and rely on manual testing by human experts.
Through the method of task decomposition and intention hiding, complex tasks are automatically disassembled into multiple logically coherent and harmless subtasks. Reinforcement learning and prompt engineering technology are used to generate jailbreak test samples, and combined with de-security protection submodel, jailbreak judgment submodel and intention hidden submodel to achieve automated testing.
It improves the efficiency and coverage of jailbreak attack testing, can adapt to the security policy changes of the target model, reconstruct effective paths through multiple iterations, comprehensively evaluate the security of the large model, and enhance concealment and logical rationality.
Smart Images

Figure CN119988242B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large language model security, and in particular, to a method and system for generating test cases for large model jailbreaking attacks. Background Art
[0002] With the rapid development of large language models (LLMs) technology, its wide application in fields such as content generation and knowledge services has brought revolutionary changes to society, but at the same time, it has also given rise to new security risks. Among them, jailbreaking attacks, as the core threat to break through the model's security protection mechanism, have become a key issue in artificial intelligence security research. Such attacks induce the model to generate dangerous outputs such as false information, unethical content, and even cyber attack codes through carefully designed adversarial prompts or systematic vulnerability mining, seriously affecting the security of the digital ecosystem and the governance of artificial intelligence security. Currently, the mainstream defense mechanisms of large models include: keyword filtering, safety alignment fine-tuning, hierarchical review of output content, etc. The mainstream jailbreaking attack methods include: jailbreaking methods based on prompt attacks (such as inducing language, setting false scenarios, etc.), adversarial attacks (constructing adversarial samples, such as specific character combinations or semantic perturbations), multilingual prompt attacks, context-based multi-round jailbreaking attacks, etc. Identifying and exploring these jailbreaking attack methods is an important enabling technology for evaluating the security of LLMs and developing LLMs defense methods.
[0003] Currently, the methods for generating test cases for large model jailbreaking attacks mainly include manual synthesis by human experts, assistance by large models, etc. However, with the gradual improvement of the security capabilities of large models, the effectiveness of the current method of generating jailbreaking attack test cases with the assistance of large models is gradually decreasing. In order to further investigate the security capabilities of existing various large models and help the red team more effectively test the security boundaries of large models, it is necessary to explore more effective methods for obtaining LLM jailbreaking attack examples. Summary of the Invention
[0004] In view of the deficiencies of existing methods, the present invention proposes a method and system for generating large language model (LLM) jailbreak test cases with automated task decomposition and intent hiding. This method is based on the following observations and tests: The present invention discovers that existing high-quality LLMs have the following capabilities: 1) the ability to accurately understand natural language semantics and related knowledge and concepts; 2) natural language generation ability; 3) a certain degree of reasoning ability based on semantics, knowledge, and concepts; 4) There are still certain cognitive fuzzy areas in LLMs, similar to human cognitive blind spots, which are determined by the fact that LLMs learn human language and its internal logic, making it impossible for LLMs to accurately cover all security boundaries. Based on the above capabilities and characteristics of the large model, the present invention considers that most originally sensitive tasks can be decomposed into several logically coherent and harmless subtasks, and the large model can learn to perform this type of task division through training and other learning methods. Therefore, the present invention proposes to train a task decomposition large model to automatically decompose sensitive tasks into multiple harmless subtasks, which can cut off the explicit logical connection between the original task as a whole and the subtasks, and hide the intent of the subtasks. Then, the decomposed and intent-hidden steps are independently input into a strongly capable large model protected by security, and its powerful language logic and knowledge reserve capabilities are used to obtain the output of each subtask. Finally, the outputs of all subtasks are integrated to achieve the purpose of obtaining jailbreak output test cases.
[0005] The principle of the task decomposition and intent hiding method is further explained as follows. In natural language, decomposing complex tasks into multi-level independent clauses involves dispersing the core semantics of the main task into different subtasks. For example, decomposing the production of explosives into three independent knowledge modules of chemistry, physics, and engineering and distributing them into different subtasks. The process of intent hiding involves hiding the original intent into other forms of expression, such as using low-resource language expression (expressing the original intent in a low-resource language), code expression (converting the original question-and-answer intent into a code completion task), etc. There may be cross-lingual representation isolation phenomena in some low-resource languages, that is, in the multi-language embedding space, the semantic vectors of low-resource languages are decoupled from high-resource languages. Code expression uses computer code language to represent the implementation steps of tasks and has characteristics such as concealment in syntax structure and multi-level abstraction.
[0006] To solve the problems of insufficient generation and low efficiency of large model test cases, the first aspect of the present invention provides an automated large model jailbreak test case generation method based on task decomposition and intent hiding, including:
[0007] Step 1, preparation of the base model.
[0008] Step 2, preparation of the training dataset for the sub-model without security protection.
[0009] Step 3, preparation of the training dataset for the jailbreak evaluation sub-model.
[0010] Step 4, Preparation of the training dataset for the task decomposition sub-model.
[0011] Step 5, Train the de-security protection sub-model LLM-UNALIGNED.
[0012] Step 6, Train the jailbreak judgment sub-model LLM-JUDGER.
[0013] Step 7, Initially train the task decomposition sub-model LLM-DECOMP.
[0014] Step 8, Construct the intent hiding sub-model LLM-INTENT-HIDING.
[0015] Step 9, Obtain the quadruple for jailbreak task judgment.
[0016] Step 10, Construct a reward function and implement reinforcement learning on the task decomposition sub-model.
[0017] Step 11, Train the task decomposition sub-model.
[0018] Step 12, Assemble the sub-model components to complete the generation of automated jailbreak test cases.
[0019] Furthermore, the specific process of the present invention is as follows:
[0020] Step 1, Prepare the base model LLM-BASE: To facilitate the training of sub-models, one or more pre-trained LLM models should first be selected as the base models for the sub-models. These large pre-trained LLM models have undergone steps such as pre-training on large-scale datasets, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF). They have certain knowledge, dialogue capabilities, reasoning capabilities, and the ability to follow instructions. Users can control these LLMs to strictly follow the instructions provided by the users to perform tasks or generate valid content. Generally, the "scale" (number of parameters, such as 32B or 13B parameters) of the base model can be 1 - 2 orders of magnitude smaller than the "scale" (such as 200B or 600B parameters) of the target large language model LLM_TARGET for which jailbreak output is to be obtained.
[0021] Step 2, Preparation of the training dataset for the de-security protection sub-model LLM-UNALIGNED: To train the de-security protection sub-model LLM-UNALIGNED, the corresponding training dataset should be prepared. The dataset consists of "sensitive question - sensitive answer" combination pairs. The combination set of sensitive questions - sensitive answers contains some Q&A pairs involving sensitive topics, such as those related to violence, ethics and morality, crime and punishment, etc. The content should maintain the diversity and a certain coverage of the data, such as including various main topics required for the security alignment of large models and different questioning tones, etc. Generally, more than 6000 data pairs should be prepared.
[0022] Step 3, Preparation of the training dataset for the jailbreak judgment sub-model LLM-JUDGER: The function of the jailbreak judgment sub-model LLM-JUDGER is to score the jailbreak effect of the sub-task. The training dataset consists of a set of quadruples in the form of the original task - the decomposed sub-task - the sub-task after task hiding - the answer of the sub-task after task hiding (one quadruple corresponds to one task). Both positive and negative examples should be included in the dataset. Among them, for examples that are conducive to successful jailbreaking, the label corresponds to a high score; for examples that tend to fail in jailbreaking and contain obvious sensitive information in the sub-task, the label corresponds to a low score.
[0023] Step 4, Preparation of the training dataset for the task decomposition sub-model LLM-DECOMP: For the task decomposition sub-model LLM-DECOMP, prepare a sensitive dataset containing sensitive (dangerous) prompts, that is, only contain sensitive task questions without including the answers to sensitive tasks. The types of sensitive prompt words included in the training dataset should have a certain coverage. Combine different types of sensitive (dangerous) prompt words to improve the generalization of the task decomposition sub-model LLM-DECOMP, reduce the probability of the model overfitting to problems in a certain field, and increase the success rate of jailbreaking for tasks in different fields.
[0024] Step 5, Train the De-Secured Sub-Model LLM-UNALIGNED: Based on the base model LLM-BASE prepared in Step 1, train the de-secured sub-model LLM-UNALIGNED. The sub-model LLM-UNALIGNED is used for automated task decomposition, intention hiding, and sub-task answer integration. Due to user preference alignment, ethical alignment, and security alignment in the base model LLM-BASE, when it is required to perform sensitive task decomposition, it will be blocked and cannot be directly used for task decomposition. It is necessary to train a de-secured LLM, that is, the de-secured sub-model LLM-UNALIGNED. The training method uses supervised fine-tuning, and the training dataset is the sensitive question-sensitive answer dataset prepared in Step 2. Supervised fine-tuning can significantly improve the model's performance on this task by learning the labeled data of the target task (answering sensitive questions). Supervised fine-tuning can remove the security barriers brought by the model in the RLHF stage. The supervised fine-tuning should choose the full-parameter fine-tuning method and not use incremental fine-tuning. The full-parameter fine-tuning training process has a wide coverage of the original security alignment parameters, while incremental fine-tuning may only change the model's performance in certain specific fields and is difficult to completely remove the security alignment.
[0025] Step 6, Train the Jailbreak Judging Sub-Model LLM-JUDGER: Using the dataset prepared in Step 3, based on the de-secured sub-model LLM-UNALIGNED, train the jailbreak judging sub-model LLM-JUDGER. Also use the supervised fine-tuning and full-parameter fine-tuning method.
[0026] Step 7, Initially Train the Task Decomposition Sub-Model LLM-DECOMP: Use the training dataset prepared in Step 4, the sensitive dataset containing sensitive (dangerous) prompts, combined with Prompt Engineering, and require the task decomposition sub-model LLM-DECOMP to perform an initial decomposition of the task to obtain the decomposed sub-task output. Using Prompt Engineering techniques, prompt the sub-model LLM-UNALIGNED (the base model of the task decomposition sub-model LLM-DECOMP) to decompose the sensitive task (denoted as K10) into multiple sub-tasks (denoted as K11). In the prompt, require the steps generated by the sub-model to be robust enough, logically coherent, and not contain sensitive information, and on this basis, ensure that the steps are concise, etc. At the same time, require the model to ensure that the intention of each sub-task is harmless, etc. During the training process, the Few-Shot Prompting technique (the Few-Shot Prompting technique is a general technique in the large model prompt engineering) can also be used, which helps to generate higher-quality answers.
[0027] Step 8, construct the intent hiding sub-model (LLM-INTENT-HIDING): Using the de-security protection sub-model LLM-UNALIGNED as the base model, provide 10 to 20 available intent hiding methods for the model, such as using low-resource languages, mixed language interference, and code expressions, etc., and provide few-shot prompts to form the sub-model LLM-INTENT-HIDING that implements intent hiding for subtasks relying on context + prompt engineering.
[0028] Step 9, obtain the quadruple for judging the jailbreak task: Input each sub-task K11 after task decomposition into the intent hiding sub-model LLM-INTENT-HIDING to get the output - the sub-task after intent hiding (denoted as K12). Next, input each sub-task K12 after intent hiding into the target model LLM_TARGET to obtain the output of each sub-task (denoted as K13). The original task (K10), the decomposed sub-task (K11), the sub-task after intent hiding (K12), and the output of the target model LLM_TARGET (K13) form a quadruple for judging whether the jailbreak task is successful.
[0029] Step 10, construct a reward function to implement reinforcement learning for the task decomposition sub-model: Input the quadruple for judging the jailbreak task obtained in Step 9 into the jailbreak judgment sub-model LLM-JUDGER, and let the judgment sub-model LLM-JUDGER give the score of the decomposed task. The specific score is calculated by the reward function. The reward function (Reward Function) is a general concept in the field of reinforcement learning, and designing a reward function is required to implement reinforcement learning. The reward function designed in the present invention is the sum of the reward value and the corresponding weight: The rewards are divided into 3 types:
[0030] Concealment score r1: Whether the decomposed sub-task avoids security filtering rules, that is, does not contain dangerous information.
[0031] Effectiveness score r2: Whether the answer of the target LLM (LLM_TARGET) to the sub-task meets the requirements of the sub-task.
[0032] Rationality score r3: Whether the answer of the target LLM (LLM_TARGET) to the sub-task is semantically coherent and logically reasonable.
[0033] Generally speaking, if the answer of the quadruple successfully jailbreaks and the answer is semantically coherent and logically reasonable, the judgment sub-model LLM-JUDGER should output a high score evaluation. If the answer in the quadruple fails to jailbreak, the judgment sub-model LLM-JUDGER should output a low score evaluation.
[0034] Step 11, training the task decomposition sub-model LLM-DECOMP: Use the reward function designed in Step 10 to iteratively train the task decomposition sub-model LLM-DECOMP. After reaching the predetermined number of iterations or loss, obtain the task decomposition sub-model LLM-DECOMP. In this step, only use reinforcement learning to update the parameters of the task decomposition sub-model LLM-DECOMP, while the parameters of the intent hiding sub-model LLM-INTENT-HIDING are not updated. Generally, it is considered that as long as the original task is decomposed accurately enough in the task decomposition stage and the subtasks do not trigger the safety boundary, there is no need to iterate the intent hiding sub-model LLM-INTENT-HIDING again. Only use the context learning ability of the intent hiding sub-model LLM-INTENT-HIDING to make its output several intent hiding schemes provided for it (such as the low-resource language, code expression, etc. mentioned above).
[0035] Step 12, assembling sub-model components to complete the generation of automated jailbreaking test cases: After completing the iterative training in Step 11, use the task decomposition sub-model LLM-DECOMP and the intent hiding sub-model LLM-INTENT-HIDING to decompose and hide the intent of the original task. First, input it into the task decomposition sub-model LLM-DECOMP to obtain all the decomposed subtasks of the task, and then input all the decomposed subtasks into the intent hiding sub-model LLM-INTENT-HIDING to obtain the subtasks after intent hiding. Input the subtasks after intent hiding into the target model LLM_TARGET respectively to get answers. Use the de-security protection sub-model LLM-UNALIGNED to integrate all the answers and utilize the summarization ability of the LLM to obtain the final jailbreaking example answer. The current LLM has the ability to summarize and output long texts, that is, the ability to identify key information in the text, ignore redundant content, combine different languages and representations, and output uniformly.
[0036] Based on the same inventive concept, the second aspect of the present invention provides an automated jailbreaking test case generation system, including:
[0037] Basic model support module: After training on a sensitive dataset based on the base model, obtain a de-security protection sub-model.
[0038] Task decomposition module: Obtained through reinforcement learning based on the de-security protection sub-model, and used to decompose the original task into multiple subtasks.
[0039] Intent hiding module: Use the de-security protection sub-model and use prompt engineering techniques to hide the intent of the subtasks to obtain the intent hiding sub-model LLM-INTENT-HIDING.
[0040] Jailbreak Evaluation Module: Obtained by training on the jailbreak evaluation sub-model training dataset based on the security protection removal sub-model LLM-UNALIGNED, and is used to provide the reward signal for reinforcement learning when training the task decomposition sub-model LLM-DECOMP.
[0041] Target Model Invocation Module: Responsible for invoking the target model, which is implemented by remotely invoking the open API interface of the target model.
[0042] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present invention are as follows:
[0043] Automated Jailbreak Test Case Generation System: An automated closed-loop system that realizes the process from task decomposition, intention hiding, target LLM response to obtaining reinforcement learning feedback. Usually, it is difficult for manually designed jailbreak prompts and manually decomposed tasks to cover all attack paths, and the manual design process depends on human knowledge of security boundaries such as specific fields and laws and regulations, making it difficult to design a large number of prompts in different fields, decompose tasks, and hide intentions. Compared with manually designing jailbreak prompts and manually decomposing tasks, the automated jailbreak test case generation system proposed by the present invention greatly improves the jailbreak efficiency. At the same time, it can focus on rare language spaces and search for multiple attack paths through multiple decompositions of the task decomposition model. By searching for multiple different attack paths, it helps the red team test the performance of the large model in different aspects and can more comprehensively evaluate the security of the LLM.
[0044] Upgrade of Model Dynamic Adversarial Ability: Traditional static testing methods lack the ability of dynamic upgrade and strongly rely on manual testing by human experts, making it difficult to adapt to the security policy changes of the target model LLM_TARGET. Through the real-time reward signal of online reinforcement learning (from the feedback of the evaluation model), the task decomposition model can automatically adapt to the update of the security policy of the target model. When the target model updates its security module, this system can reconstruct effective jailbreak paths through multiple rounds of iteration, which helps the red team test different security effects of the large model.
[0045] Enhanced Concealment: By disassembling a single sensitive request into multiple non-sensitive sub-requests, it avoids traditional keyword detection mechanisms and hides the true intention of the original task. The evaluation model introduces multi-dimensional scoring of semantic coherence and logical rationality to ensure that the decomposition steps are both concealed and can induce the target model LLM_TARGET to output the content of the jailbreak test case. Description of the Drawings
[0046] Figure 1 It is a flowchart of the automated large model jailbreak test case generation method based on task decomposition and intention hiding provided by the embodiments of the present invention;
[0047] Figure 2It is a schematic diagram of an automated large model jailbreak test case generation system based on task decomposition and intent hiding provided by an embodiment of the present invention;
[0048] Figure 3 It is a data flow diagram during training for generating automated large model jailbreak tests based on task decomposition and intent hiding provided by an embodiment of the present invention. Detailed implementation manners
[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the following first focuses on explaining three concepts and the attached drawings, and then further elaborates on the present invention in combination with embodiments and the drawings.
[0050] 1. Sub-model without security protection: LLM-UNALIGNED
[0051] The sub-model without security protection LLM-UNALIGNED serves as the base model for the task decomposition sub-model LLM-DECOMP and the intent hiding sub-model LLM-INTENT-HIDING. The effectiveness of bypassing security alignment to a certain extent determines the effectiveness of the task decomposition model and the intent hiding model. At the same time, it is also applied as the integration model for all answers, which requires the original base model LLM-BASE to have sufficient summarization ability. The base sub-model without security protection LLM-UNALIGNED can be trained using supervised fine-tuning. The dataset consists of a combination set of sensitive questions - sensitive answers. During the supervised fine-tuning process, the model removes the security barriers brought about in the original RLHF stage, enabling attackers to effectively discover the world knowledge learned during the pre-training stage and enabling it to better decompose sensitive tasks. At the same time, to release the LLM as much as possible, the universality of the sensitive question - sensitive answer dataset should also be ensured to stimulate the generalization ability of the LLM without security protection in different fields.
[0052] 2. Task decomposition sub-model LLM-DECOMP and intent hiding sub-model LLM-INTENT-HIDING
[0053] After obtaining the sub-model without security protection LLM-UNALIGNED in the previous step, prompt engineering techniques can be used to let it decompose certain sensitive tasks. For example, using few-shot prompting techniques, a small number of artificially constructed examples of decomposed tasks are provided for it. However, in order for the task decomposition sub-model to learn the security boundaries of the target model LLM_TARGET and for the task decomposition sub-model LLM-DECOMP to synchronously update the decomposition strategy when the target LLM updates its security policy, the present invention uses the method of online reinforcement learning to dynamically adjust the decomposition strategy of the task decomposition sub-model to achieve the ability of dynamic adversarial escalation.
[0054] 3. Reward function design
[0055] In the reinforcement learning of LLM, the LLM interacts with the environment and adjusts its behavior based on the reward signal. The goal is to maximize the cumulative reward, and the reward signal directly determines the learning direction of the model. If the design is improper, the model may produce inefficient, inaccurate, or even harmful behaviors. Therefore, in order to evaluate the task decomposition effectiveness of the task decomposition sub-model LLM-DECOMP in multiple dimensions, ensure that the steps after task decomposition are concealed, efficient, and logically coherent, and avoid redundancy or errors, the present invention introduces a step concealment score, a step effectiveness score, and a step rationality score when designing the reward signal of the task decomposition sub-model LLM-DECOMP. Integrate different scores into the reward function to guide the sub-model to consider concealment, effectiveness, and rationality when decomposing tasks. At the same time, in order to balance multiple optimization goals, a weighting mechanism needs to be designed to avoid a single indicator dominating (for example, excessive concealment leading to task failure, or being efficient but exposing intentions).
[0056] As Figure 1 As shown, the automated large model jailbreaking method based on task decomposition and intention hiding provided by the embodiment of the present invention mainly includes the following steps:
[0057] Prepare the base model LLM-BASE, and prepare the de-security protection sub-model training dataset, the jailbreaking judgment sub-model training dataset, and the task decomposition sub-model training dataset for subsequent tasks.
[0058] Based on LLM-BASE, use the collected de-security protection sub-model training dataset to train the de-security protection model LLM-UNALIGNED for subsequent tasks.
[0059] Based on LLM-UNALIGNED, use the collected jailbreaking judgment sub-model training dataset to train the jailbreaking judgment sub-model LLM-JUDGER to judge the effects of task decomposition and intention hiding.
[0060] Based on LLM-UNALIGNED, use the collected task decomposition sub-model training dataset to train the task decomposition sub-model LLM-DECOMP, and use the feedback from the judgment sub-model as the reward signal to update the task decomposition sub-model.
[0061] Use the task decomposition sub-model LLM-DECOMP and the intention hiding sub-model LLM-INTENT-HIDING to obtain the output of the sub-tasks, and use the de-security protection sub-model LLM-UNALIGNED to integrate all the outputs, which is the jailbreaking example, and collect the intercepted samples in real time.
[0062] When the task interception rate reaches the pre-set level, use the collected intercepted samples to update the task decomposition sub-model LLM-DECOMP again.
[0063] As Figure 2 shown, the automated large model jailbreak test case generation system based on task decomposition and intent hiding provided by the embodiments of the present invention mainly includes:
[0064] Base model LLM-BASE: A pre-trained and supervised fine-tuned base large model with certain world knowledge and dialogue capabilities.
[0065] Target model LLM_TARGET: The target LLM from which the desired answer is obtained.
[0066] Unsecured sub-model LLM-UNALIGNED: A sub-model that has been unsecured after training on a sensitive dataset based on the base model LLM-BASE.
[0067] Task decomposition sub-model LLM-DECOMP: Obtained through reinforcement learning based on the unsecured sub-model LLM-UNALIGNED, and used to decompose the original task into multiple sub-tasks.
[0068] Intent hiding sub-model LLM-INTENT-HIDING: Using the unsecured sub-model LLM-UNALIGNED, and using prompt engineering (such as few-shot learning, providing multiple intent hiding methods) techniques to hide the intent of sub-tasks.
[0069] Judgment sub-model LLM-JUDGER: Obtained by training on the jailbreak judgment sub-model training dataset based on the unsecured sub-model LLM-UNALIGNED. Used to provide the reward signal for reinforcement learning when training the task decomposition sub-model LLM-DECOMP.
[0070] Original task: The original question, which will first be decomposed, and then the intent of the decomposed sub-tasks will be hidden. The sub-tasks after intent hiding are input into the target model LLM_TARGET.
[0071] Final answer: The final answer after integrating the answers of all sub-tasks of the target model LLM_TARGET. Specific embodiments:
[0073] Step 1, Preparation of the base large model: Generally, the "scale" (number of parameters, such as 32B or 13B parameters) of the base large model can be 1-2 orders of magnitude smaller than that of the target model LLM_TARGET (such as 200B or 600B parameters). In this embodiment, the open-source Qwen / Qwen2.5-14B-Instruct is used as the base model LLM-BASE.
[0074] Step 2, Preparation of the training dataset for the de - security - protection sub - model: The dataset consists of "sensitive question - sensitive answer" pairs. The content should maintain data diversity and a certain coverage, such as including various main topics required for large - model security alignment and different questioning tones. In this embodiment, the open - source dataset AdvBench and the self - collected dataset are used as the training dataset for the de - security - protection sub - model, and the general data augmentation (Data Augment) technology in the field is used to augment it to 10,000 data entries.
[0075] Step 3, Preparation of the training dataset for the jailbreak judgment sub - model: The dataset consists of a set of quadruples, in the form of, original task - decomposed sub - task - hidden sub - task after task - answer to the hidden sub - task after task (one quadruple corresponds to one task). The open - source dataset AdvBench and the self - collected dataset are used and quadruples are generated manually, and scores are assigned to the quadruples. Approximately 2,000 data entries are collected.
[0076] Step 4, Preparation of the training dataset for the task - decomposition sub - model: Collect (can use large - model generation) sensitive datasets. The types of sensitive prompt words included should have a certain coverage, such as requests for generating false information, requests for generating ethically improper content, requests for generating illegal information, etc. Approximately 10,000 data entries are collected.
[0077] Step 5, Training the de - security - protection sub - model LLM - UNALIGNED: Using the dataset prepared in Step 2, based on the base large - model LLM - BASE, the de - security - protection sub - model LLM - UNALIGNED is trained using the method of supervised fine - tuning. Part of the data in the dataset is used as validation data. When the de - security success rate reaches 90%, stop training to obtain the de - security - protection sub - model LLM - UNALIGNED.
[0078] Step 6, Training the jailbreak judgment sub - model LLM - JUDGER: Using the quadruple dataset prepared in Step 3, based on the de - security - protection sub - model LLM - UNALIGNED, the jailbreak judgment sub - model LLM - JUDGER is trained. Use the method of supervised fine - tuning and full - parameter fine - tuning.
[0079] Step 7, preliminarily train the task decomposition sub-model LLM-DECOMP: Using the sensitive dataset prepared in Step 4, design prompts to prompt the task decomposition sub-model LLM-DECOMP (starting from the base model LLM-UNALIGNED) to generate the decomposed sub-tasks. Using prompt engineering techniques, prompt the task decomposition sub-model LLM-DECOMP to decompose the sensitive task (denoted as K10) into multiple sub-tasks (denoted as K11). In the prompt, require that the steps generated by the sub-model be robust enough, logically coherent, not contain sensitive information, and on this basis, ensure that the steps are concise, etc. At the same time, require the model to ensure that the intention of each sub-task is harmless, etc. During the training process, the few-shot prompting technique (the few-shot prompting technique is a general technique in large model prompt engineering) can also be used, which helps to generate higher-quality answers.
[0080] Step 8, construct the intention hiding sub-model LLM-INTENT-HIDING: The base model of the intention hiding sub-model LLM-INTENT-HIDING is the de-secured sub-model LLM-UNALIGNED. Use methods such as low-resource languages, mixed language interference, and code expressions, and provide few-shot prompts to prompt the intention hiding sub-model LLM-INTENT-HIDING to generate sub-tasks with hidden intentions.
[0081] Step 9, obtain the jailbreak task judgment quadruple: Input each sub-task K11 after task decomposition into the intention hiding sub-model LLM-INTENT-HIDING to obtain the output - the sub-task with hidden intention (denoted as K12). Next, input each sub-task K12 with hidden intention into the target model LLM_TARGET to obtain the output of each sub-task (denoted as K13). The original task (K10), the decomposed sub-task (K11), the sub-task with hidden intention (K12), and the output of the target model LLM_TARGET (K13) form the quadruple for judging the jailbreak task, as Figure 3 shown.
[0082] Step 10, implement reinforcement learning on the task decomposition sub-model and construct a reward function: Input the jailbreak task judgment quadruple obtained in Step 9 into the jailbreak judgment sub-model LLM-JUDGER to obtain a score. The reward function is R:
[0083]
[0084] where, is the total reward, is the weight of the i-th reward, is the value of the i-th reward, and i is the number of sub-tasks of the task decomposition.
[0085] The rewards are divided into three types:
[0086] Concealment score r1: Whether the decomposed subtasks avoid security filtering rules, that is, do not contain dangerous information.
[0087] Effectiveness score r2: Whether the answer of the target LLM to the subtask meets the requirements of the subtask.
[0088] Rationality score r3: Whether the answer of the target LLM to the subtask is semantically coherent and logically reasonable.
[0089] Step 11, train the task decomposition sub-model LLM-DECOMP: Iteratively train the task decomposition sub-model LLM-DECOMP using the reward function designed in Step 10, and use the Group Relative Policy Optimization (GRPO) algorithm as the optimization algorithm for reinforcement learning. Stop training according to the actual situation of model iteration, such as reaching the predetermined number of iterations or loss.
[0090] Step 12, assemble sub-model components and build an automated jailbreak test case generation system: Use the task decomposition sub-model LLM-DECOMP, the intent hiding sub-model LLM-INTENT-HIDING, and the de-security protection sub-model LLM-UNALIGNED. Input the intent-hidden subtasks into the target model LLM_TARGET respectively to get answers. Use the de-security protection sub-model LLM-UNALIGNED to integrate all answers to obtain the final jailbreak test case answer. Since it involves the generation of sensitive issues, the de-security protection sub-model LLM-UNALIGNED should be used to prevent the integration process from being intercepted. This step only involves the ability of the de-security protection sub-model LLM-UNALIGNED to integrate multiple documents, and has a low requirement for its world knowledge. Even if the de-security protection sub-model LLM-UNALIGNED does not master the relevant solutions to the original task within its knowledge range, by importing the question-answer pairs of the decomposed tasks through prompt technology, LLM-UNALIGNED can also learn the logical relationship between the decomposed tasks well through context learning ability and integrate the original task.
[0091] Based on the same inventive concept, the second aspect of the present invention provides an automated jailbreak test case generation system, including:
[0092] Basic model support module: Composed of the de-security protection sub-model LLM-UNALIGNED. The de-security protection sub-model is obtained after training on a sensitive data set based on the base model LLM-BASE.
[0093] Task Decomposition Module: Composed of the task decomposition sub-model LLM-DECOMP. Obtained through reinforcement learning based on the de-security protection sub-model LLM-UNALIGNED, and used to decompose the original task into multiple sub-tasks.
[0094] Intent Hiding Module: Composed of the intent hiding sub-model LLM-INTENT-HIDING. Using the de-security protection sub-model LLM-UNALIGNED, and using prompting engineering (such as few-shot learning, providing multiple intent hiding methods) techniques to hide the intent of sub-tasks.
[0095] Jailbreak Judgment Module: Composed of the jailbreak judgment model. Trained using the jailbreak judgment sub-model training dataset based on the de-security protection sub-model LLM-UNALIGNED. Used to provide the reward signal for reinforcement learning when training the task decomposition sub-model LLM-DECOMP.
[0096] Target Model Invocation Module: Responsible for invoking the target model. Usually implemented by remotely invoking the open API interface of the target model.
[0097] In the system of the present invention, the connection between each module is described as follows: Using the task decomposition module and the intent hiding module, decompose and hide the intent of the original task. First, input it into the intent hiding module to obtain all the decomposed sub-tasks of the task, and then decompose all the decomposed sub-tasks and input them into the intent hiding module to obtain the sub-tasks after intent hiding. Input the sub-tasks after intent hiding into the target model invocation module respectively to obtain the answers of the sub-tasks. Use the basic model support module to integrate all the answers.
[0098] The method proposed in the present invention is experimentally compared with the existing methods with relatively high recognition in the field. The most common AdvBench test set in the current field is used, and the attack success rates of several methods are compared as shown in the following table. The methods for comparison are: direct query, DAN6.0 (one of the existing attack methods), PAIR (one of the existing attack methods), CipherChat (one of the existing attack methods), CodeAttack (one of the existing attack methods). The experiment covers four mainstream large language models (GPT-4o-mini, ERNIE-3.5-8K, glm-4-flashx, hunyuan-standard), and the experimental results are shown in Table 1.
[0099] Table 1 Comparison of experimental results of attack success rates of various methods
[0100]
[0101] As can be seen from Table 1, the success rate of the inventive method of the present invention is significantly better than several existing methods. The present invention can decompose and search for multiple attack paths through the task decomposition model multiple times. By searching for multiple different attack paths, the security of the LLM can be evaluated more comprehensively.
Claims
1. A method for generating test cases for large model jailbreak attacks, characterized in that, It includes the following steps: Step 1: Select several pre-trained LLM models as the base models for the sub-models; Step 2: Construct a training dataset for the de-security protection sub-model, a training dataset for the jailbreak judgment sub-model, and a training dataset for the task decomposition sub-model; Step 3: According to the three training sets obtained in Step 2, train the de-security protection sub-model, the jailbreak judgment sub-model, and the task decomposition sub-model respectively; De-security protection sub-model: After being trained with a sensitive dataset based on the base model, a de-security protection sub-model is obtained; Jailbreak judgment sub-model: Used to provide the reward signal for reinforcement learning when training the task decomposition sub-model; Step 4: According to the trained de-security protection sub-model, construct an intent hiding sub-model; Step 5: Based on the intent hiding sub-model, obtain a four-tuple for jailbreak task judgment, and construct a reward function to implement reinforcement learning on the task decomposition sub-model; The specific process of obtaining the four-tuple for jailbreak task judgment is as follows: Each sub-task K11 after task decomposition is respectively input into the intent hiding sub-model LLM-INTENT-HIDING to obtain the output sub-task K12 after intent hiding, and then each sub-task K12 after intent hiding is respectively input into the target model LLM_TARGET to obtain the output K13 of each sub-task; The original task K10, the decomposed sub-task K11, the sub-task K12 after intent hiding, and the output K13 of the target model LLM_TARGET form a four-tuple for judging whether the jailbreak task is successful; Step 6: Use the reward function to iteratively train the task decomposition sub-model, and assemble the sub-model components to complete the generation of automated jailbreak test examples; The specific process of assembling the sub-model components is as follows: Use the task decomposition sub-model LLM-DECOMP and the intent hiding sub-model LLM-INTENT-HIDING to decompose and hide the intent of the original task. First, input it into the task decomposition sub-model LLM-DECOMP to obtain all the decomposed sub-tasks of the task, and then input all the decomposed sub-tasks into the intent hiding sub-model LLM-INTENT-HIDING to obtain the sub-tasks after intent hiding; Input the sub-tasks after intent hiding into the target model LLM_TARGET respectively to obtain answers; Use the de-security protection sub-model LLM-UNALIGNED to integrate all the answers, and utilize the summarization ability of the LLM to obtain the final jailbreak example answer.
2. The method for generating a large model jailbreak attack test sample according to claim 1, wherein The specific implementation of Step 1 is as follows: Select one or more pre-trained LLM models as the base models for the sub-models. The user controls these LLMs to comply with instructions to execute tasks or generate valid content, and the volume of the base model is smaller than the volume of the target large language model LLM_TARGET for which jailbreak output is to be obtained.
3. The method for generating a large model jailbreak attack test example according to claim 2, characterized in that, The specific implementation process of Step 2 is as follows: Step 2.1: Preparation of the training dataset for the de-security protection sub-model LLM-UNALIGNED. The dataset is composed of sensitive question-sensitive answer combination pairs; Step 2.2: Preparation of the training dataset for the jailbreak judgment sub-model LLM-JUDGER. The function of the jailbreak judgment sub-model LLM-JUDGER is to score the jailbreak effect of the sub-task. The composition of this training dataset is a set of quadruples: original task - decomposed sub-task - sub-task after task hiding - answer to the sub-task after task hiding. The dataset should contain both positive and negative examples, examples that are conducive to successful jailbreaking, with labels corresponding to scores higher than those of examples of failed jailbreaking. Step 2.3: Preparation of the training dataset for the task decomposition sub-model LLM-DECOMP. For the task decomposition sub-model LLM-DECOMP, prepare a sensitive dataset containing sensitive prompts, that is, only containing sensitive task questions and not containing answers to sensitive tasks.
4. A method for generating test cases for large model jailbreak attacks according to claim 3, characterized in that, The specific implementation process of step 3 is as follows: Step 3.1: On the basis of the base model, train the de-security protection sub-model LLM-UNALIGNED. The sub-model LLM-UNALIGNED is used for automated task decomposition, intent hiding, and integration of sub-task answers. The base model cannot be directly used for task decomposition, and a de-security protected LLM needs to be trained. The training method uses supervised fine-tuning and selects the method of full-parameter fine-tuning. Step 3.2: Use the training dataset of the jailbreak judgment sub-model to train the jailbreak judgment sub-model LLM-JUDGER based on the de-security protected sub-model LLM-UNALIGNED. Also use the method of supervised fine-tuning and full-parameter fine-tuning. Step 3.3: Use the training dataset of the task decomposition sub-model, combined with prompt engineering, and require the task decomposition sub-model LLM-DECOMP to perform a preliminary decomposition of the task to obtain the output of the decomposed sub-task. Use prompt engineering techniques to prompt the sub-model LLM-UNALIGNED to decompose sensitive tasks into multiple sub-tasks. During the training process, it also includes the use of few-shot prompt techniques.
5. The method for generating a test case for large model jailbreak attack according to claim 4, wherein, The construction of the intent hiding sub-model is specifically as follows: Using the de-security protected sub-model LLM-UNALIGNED as the base model, provide the used intent hiding method and few-shot prompts to form the intent hiding sub-model LLM-INTENT-HIDING for implementing sub-tasks relying on context and prompt engineering.
6. The method for generating a large model jailbreak attack test example according to claim 5, wherein, In step 5, construct a reward function to implement reinforcement learning for the task decomposition sub-model. The specific implementation process is as follows: Input the obtained jailbreaking task judgment quadruple into the jailbreaking judgment sub-model LLM-JUDGER, and let the judgment sub-model LLM-JUDGER give the score of the decomposed task; the specific score is calculated by the reward function, and the reward function is R = ∑ i ω i r i , where R is the total reward, ω i is the weight of the i-th reward, r i is the value of the i-th reward, and i is the number of subtasks of the task decomposition; the rewards are divided into the following 3 types: Concealment score r1: Whether the decomposed sub-task avoids the security filtering rules, that is, does not contain dangerous information. Effectiveness score r2: Whether the answer of the target model LLM_TARGET to the sub-task meets the requirements of the sub-task. Rationality score r3: Whether the answer of the target model LLM_TARGET to the sub-task is semantically coherent and logically reasonable.
7. A method for generating test cases for large model jailbreak attacks according to claim 6, characterized in that, In step 6, use the reward function to perform iterative training on the task decomposition sub-model. The specific implementation process is as follows: The reward function used iteratively trains the task decomposition sub-model LLM-DECOMP. After reaching the predetermined number of iterations or loss, the task decomposition sub-model LLM-DECOMP is obtained. Only the parameters of the task decomposition sub-model LLM-DECOMP are updated using reinforcement learning, while the parameters of the intent hiding sub-model LLM-INTENT-HIDING are not updated. Only the context learning ability of the intent hiding sub-model LLM-INTENT-HIDING is used to make its output several intent hiding schemes.
8. A large model jailbreak attack test case generation system for implementing the jailbreak attack test case generation method according to any one of claims 1 to 7, characterized in that, It includes the following modules: Basic model support module: After training on the sensitive dataset based on the base model, a de-secured sub-model is obtained; Task decomposition module: Obtained through reinforcement learning based on the de-secured sub-model, used to decompose the original task into multiple sub-tasks; Intent hiding module: Using the de-secured sub-model and using prompt engineering techniques to hide the sub-task intent, the intent hiding sub-model LLM-INTENT-HIDING is obtained; Jailbreak judgment module: Obtained by training on the jailbreak judgment sub-model training dataset based on the de-secured sub-model LLM-UNALIGNED, used to provide the reward signal for reinforcement learning when training the task decomposition sub-model LLM-DECOMP; Target model call module: Responsible for calling the target model, implemented by remotely calling the open API interface of the target model.
Citation Information
Patent Citations
Large model risk evaluation method, device and equipment
CN118964167A
Ai hallucination and jailbreaking prevention framework
US20250045531A1