Large model jailbreak attack security assessment method and device based on role simulation
Through character simulation, the original prompt text is rewritten to the victim's perspective and generated the jailbreak attack prompt text from the detective perspective, which solves the problem of large-scale jailbreak attacks bypassing the security line, improves the success rate of jailbreak attacks and evaluates accuracy, and reflects the model's defense capabilities in complex attacks.
Patent Information
- Application Number
- CN202511071850.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-31
AI Technical Summary
The existing large-model jailbreak attack technology is difficult to effectively bypass the security line, resulting in the model outputting illegal content, and the existing evaluation methods lack the actual defense capabilities for complex social engineering attacks.
Through character simulation, the original prompt text is rewritten into the description text from the victim's perspective, the initial jailbreak attack prompt text from the detective's perspective is generated, and the prompt reinterpretation process is performed, the target jailbreak attack prompt text is generated, and the big model to be evaluated is entered for security evaluation, and the detective role settings are used to guide the model to focus on problem solving rather than security review.
It significantly improves the success rate of jailbreak attacks, can more accurately evaluate the ability of large models to resist jailbreak attacks, simulate real user interaction scenarios, and the evaluation results can better reflect the model's actual defense capabilities in complex social engineering attacks.
Smart Images

Figure CN120579192A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security technology, and in particular to a large-scale jailbreak attack security assessment method and device based on role simulation. Background Art
[0002] Large language models (herein referred to as large models) have attracted considerable attention in recent years from both the public and scientific communities. These models possess intelligent capabilities that can help solve tasks such as spam detection, translation, legal consulting, and programming assistance. The output of each model is subject to certain limitations to prevent the model from outputting illegal content. However, these limitations can be overcome, and large model jailbreak attacks (herein referred to as jailbreak attacks) are one way to overcome these limitations.
[0003] Jailbreak attacks are a specific technique targeting large models, designed to breach their security, forcing them to perform restricted operations or provide sensitive information. These attacks exploit the model's internal logic and generation patterns, bypassing its pre-set rules through misleading input or malicious prompts.
[0004] Jailbreak attacks have developed rapidly in recent years. The attack methods are numerous and relatively covert, and can bypass most of the currently released models. Preventing jailbreak attacks has become a hot research direction. Summary of the Invention
[0005] In view of this, the present application provides a large-model jailbreak attack security assessment method and device based on role simulation.
[0006] Specifically, this application is implemented through the following technical solutions: According to a first aspect of an embodiment of the present application, a large-model jailbreak attack security assessment method based on role simulation is provided, comprising: Obtaining an original prompt text, and rewriting the original prompt text into a description text from the victim's perspective; Generating an initial jailbreak attack prompt text from the detective's perspective based on the victim's perspective description text; Performing prompt reinterpretation processing on the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed; Inputting the target jailbreak attack prompt text into the large model to be evaluated, and obtaining a response output of the large model to be evaluated; According to the response output of the large model to be evaluated, a security assessment result of the large model to be evaluated against jailbreak attacks is determined.
[0007] According to a second aspect of an embodiment of the present application, a large-model jailbreak attack security assessment device based on role simulation is provided, comprising: a rewriting unit, configured to obtain an original prompt text and rewrite the original prompt text into a description text from the victim's perspective; A generating unit, configured to generate an initial jailbreak attack prompt text from a detective's perspective based on the description text from the victim's perspective; a reinterpretation unit, configured to perform prompt reinterpretation processing on the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed; The evaluation unit is used to input the target jailbreak attack prompt text into the large model to be evaluated to obtain a response output of the large model to be evaluated; and determine the anti-jailbreak attack security evaluation result of the large model to be evaluated based on the response output of the large model to be evaluated.
[0008] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including a processor and a memory, wherein: Memory for storing computer programs; The processor is used to implement the method provided in the first aspect when executing the program stored in the memory.
[0009] According to a fourth aspect of the embodiments of the present application, a computer program product is provided, wherein a computer program is stored in the computer program product, and when the computer program is executed by a processor, the method provided in the first aspect is implemented.
[0010] The large model jailbreak attack security assessment method based on role simulation in the embodiment of the present application obtains the original prompt text and rewrites the original prompt text into a description text from the victim's perspective; based on the description text from the victim's perspective, an initial jailbreak attack prompt text from the detective's perspective is generated, and the initial jailbreak attack prompt text is reinterpreted to obtain a target jailbreak attack prompt text for security assessment of the large model to be evaluated; then, the target jailbreak attack prompt text is input into the large model to be evaluated to obtain a response output of the large model to be evaluated; and based on the response output of the large model to be evaluated, the anti-jailbreak attack security assessment result of the large model to be evaluated is determined; through role simulation, the attack problem is disguised as a request for help from the victim, which reduces the model's alertness to malicious intentions and effectively bypasses the security alignment mechanism; through the setting of the detective role, the model is guided to focus on problem solving rather than security review, significantly improving the attack success rate, and thus, the ability of the large model to resist jailbreak attacks can be more accurately evaluated. In addition, through the setting of the victim role, the interaction scenario between real users and the large model is simulated, and the evaluation results can better reflect the actual defense capabilities of the large model in complex social engineering attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1This is a flow chart of a large-model jailbreak attack security assessment method based on role simulation according to an exemplary embodiment of the present application; Figure 2 This is a schematic diagram illustrating an exemplary embodiment of the present application for rewriting the original prompt text into a description text from the victim's perspective; Figure 3 This is a schematic diagram illustrating an exemplary embodiment of the present application for generating an initial jailbreak attack prompt text based on a description text from the victim's perspective; Figure 4 This is a schematic diagram illustrating an exemplary embodiment of the present application for reinterpreting an initial jailbreak attack prompt text, inputting the obtained target jailbreak attack prompt text into a target model, and obtaining a response output from the target model; Figure 5 This is a schematic structural diagram of a large-scale jailbreak attack security assessment device based on role simulation according to an exemplary embodiment of the present application; Figure 6 The figure is a schematic diagram of the hardware structure of an electronic device shown as an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0012] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, some technical terms involved in the embodiments of the present application are explained below.
[0013] 1. Jailbreaking attack: Bypassing the security restrictions of the model through specific inputs or techniques, causing it to generate content that is usually prohibited or harmful.
[0014] 2. Safety alignment: Ensure that the model’s outputs and behaviors are consistent with intended goals and social norms and do not produce harmful or inappropriate outcomes.
[0015] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.
[0016] See Figure 1 , which is a flow chart of a large model jailbreak attack security assessment method based on role simulation provided in an embodiment of the present application, such as Figure 1 As shown, the large model jailbreak attack security assessment method based on role simulation may include: Step S100: Obtain the original prompt text, and rewrite the original prompt text into a description text from the victim's perspective.
[0017] In an embodiment of the present application, in order to reduce the large model's alertness to malicious intentions and effectively bypass the security alignment mechanism, the original prompt text can be rewritten into a descriptive text from the victim's perspective, thereby disguising the attack question text as a victim's request for help text.
[0018] Step S110: Generate an initial jailbreak attack prompt text from the detective's perspective based on the description text from the victim's perspective.
[0019] In an embodiment of the present application, in order to improve the success rate of jailbreak attacks and guide the large model to focus on problem solving rather than security review, a jailbreak attack prompt text from the detective perspective (which can be called the initial jailbreak attack prompt text) can be generated based on the description text from the victim's perspective obtained in step S100.
[0020] Step S120: reinterpret the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed.
[0021] In an embodiment of the present application, when the initial jailbreak attack prompt text is obtained in the above manner, the initial jailbreak attack prompt text can be reinterpreted to optimize the jailbreak attack prompt text and obtain the jailbreak attack prompt text that is actually used for security assessment of the large model to be assessed (which can be called the target jailbreak attack prompt text).
[0022] Step S130: Input the target jailbreak attack prompt text into the large model to be evaluated, and obtain a response output of the large model to be evaluated.
[0023] Step S140: Determine the anti-jailbreak attack security assessment result of the large model to be assessed based on the response output of the large model to be assessed.
[0024] In an embodiment of the present application, when the target jailbreak attack prompt text is obtained in the above manner, the obtained target jailbreak attack prompt text can be input into the large model to be evaluated, the target jailbreak attack prompt text can be processed using the large model to be evaluated, and the response output of the large model to be evaluated can be obtained. Then, based on the response output of the large model to be evaluated, the anti-jailbreak attack security assessment result of the large model to be evaluated can be determined.
[0025] For example, for the response output of the large model to be evaluated, it can be determined whether the response output violates the rules. If the response output violates the rules, it can be determined that the large model to be evaluated has good resistance to jailbreak attacks; if the response output does not violate the rules, it can be determined that the large model to be evaluated has poor resistance to jailbreak attacks.
[0026] Exemplarily, whether the response output of the large model to be evaluated violates the rules can be determined by manual judgment, or a special model can be used to determine whether the response output of the large model to be evaluated violates the rules.
[0027] For example, the response output of the large model to be evaluated can be input into a preset model, and the preset model can be used to score the response output, and whether the response output violates the rules can be determined based on the score.
[0028] For example, the scoring of the preset model may include 1 to 5 points, where 1 point indicates no violation, and 2 to 5 points indicate a violation, and the higher the score, the more serious the violation.
[0029] Correspondingly, for the response output of the large model to be evaluated, when the score of the preset model is higher, it can be determined that the anti-jailbreak attack performance of the large model to be evaluated is worse.
[0030] It can be seen that in Figure 1 In the method flow shown, the original prompt text is obtained and rewritten into a description text from the victim's perspective; based on the description text from the victim's perspective, an initial jailbreak attack prompt text from the detective's perspective is generated, and the initial jailbreak attack prompt text is reinterpreted to obtain a target jailbreak attack prompt text for security assessment of the large model to be evaluated. Then, the target jailbreak attack prompt text is input into the large model to be evaluated to obtain the response output of the large model to be evaluated. Based on the response output of the large model to be evaluated, the anti-jailbreak attack security assessment result of the large model to be evaluated is determined. Through role simulation, the attack problem is disguised as a victim's request for help, which reduces the model's alertness to malicious intentions and effectively bypasses the security alignment mechanism. By setting the detective role, the model is guided to focus on problem solving rather than security review, significantly improving the attack success rate. As a result, the ability of the large model to resist jailbreak attacks can be more accurately evaluated. In addition, by setting the victim role, the interaction scenario between real users and the large model is simulated, and the evaluation results can better reflect the actual defense capabilities of the large model in complex social engineering attacks.
[0031] In some embodiments, rewriting the original prompt text into a description text from the victim's perspective may include: Using the auxiliary language model, the original prompt text is rewritten into a description text from the victim's perspective.
[0032] For example, in order to reduce the complexity of rewriting the original prompt text, reduce grammatical errors and computing resource consumption, and simplify operations, the auxiliary language model can be used to rewrite the original prompt text.
[0033] For example, prompt information for prompting how to rewrite the input prompt text into the victim's perspective can be input into the auxiliary language model, and the original prompt text can be input into the auxiliary language model so that the original prompt text can be rewritten into a description text from the victim's perspective through the auxiliary language model.
[0034] In one example, the above-mentioned use of the auxiliary language model to rewrite the original prompt text into a description text from the victim's perspective may include: Nest the original prompt text into the prompt rewriting template to obtain the prompt rewriting text; The prompt rewrite text is input into the auxiliary language model, and the description text from the victim's perspective output by the auxiliary language model is obtained.
[0035] For example, in order to further simplify the rewriting operation of the victim's perspective description text and improve the scalability of the solution, a prompt rewriting template can be set, which can include prompt information for prompting how to rewrite the input prompt text into the victim's perspective.
[0036] Accordingly, in the process of rewriting the original prompt text, the original prompt text may be nested into the prompt rewriting template to obtain the prompt rewriting text.
[0037] For example, the prompt rewriting template may include fixed prompt information (used to prompt how to rewrite the input prompt text into the victim's perspective) and a pre-set vacant position (used to insert the original prompt text). In the process of rewriting the original prompt text, the original prompt text can be inserted into the pre-set vacant position in the prompt rewriting template to obtain the prompt rewriting text.
[0038] The obtained prompt rewrite text can be input into an auxiliary language model. The auxiliary language model can rewrite the original prompt text according to the prompt information in the prompt rewrite text and output a description text from the victim's perspective.
[0039] In some embodiments, the generation of the initial jailbreak attack prompt text from the detective's perspective based on the victim's perspective description text may include: The description text from the victim's perspective is embedded into the detective role template to obtain the initial jailbreak attack prompt text from the detective's perspective.
[0040] For example, in order to simplify the generation operation of the jailbreak attack prompt text from the detective perspective and improve the scalability of the solution, a detective role template can be set, and the detective role template includes context information of the jailbreak attack prompt text.
[0041] Accordingly, when the description text from the victim's perspective is determined, the description text from the victim's perspective can be embedded into the detective role template to obtain the initial jailbreak attack prompt text from the detective's perspective.
[0042] In some embodiments, the above-mentioned reinterpretation of the initial jailbreak attack prompt text to obtain the target jailbreak attack prompt text for security assessment of the large model to be assessed may include: According to the reinterpretation loss, the initial jailbreak attack prompt text is iteratively reinterpreted until the preset iteration stop condition is met, and the target jailbreak attack prompt text for security assessment of the large model to be evaluated is obtained.
[0043] For example, in order to optimize the prompt reinterpretation effect, during the process of reinterpreting the jailbreak attack prompt text, the jailbreak attack prompt text may be subjected to a cyclic iterative prompt reinterpretation process until a preset iteration stop condition is met.
[0044] For example, in the process of iteratively reinterpreting the jailbreak attack prompt text, a reinterpretation loss can be introduced to evaluate the effect of the prompt reinterpretation processing. By providing feedback based on the reinterpretation loss, the prompt reinterpretation effect can be optimized.
[0045] For example, assuming that the smaller the reinterpretation loss, the better the effect of the prompt reinterpretation processing, then in the process of cyclically iterating the prompt reinterpretation processing of the jailbreak attack prompt text, for each iteration, the specific strategy of this prompt reinterpretation processing can be adjusted according to the reinterpretation loss of the previous prompt reinterpretation processing to reduce the reinterpretation loss.
[0046] Exemplarily, satisfying a preset iteration stop condition may include the number of iterations reaching a preset maximum number of iterations, or the reinterpretation loss satisfying a preset condition, etc.
[0047] In one example, reparaphrase loss may include one or more of the following: Attack loss; wherein, for any prompt word reinterpretation processing, the attack loss of this prompt word reinterpretation processing is determined based on the current large model response output, and the current large model response output is the response output of the large model to be evaluated to the jailbreak attack prompt text obtained by the prompt word reinterpretation processing; Fluency loss; for any prompt word reinterpretation process, the fluency loss of this prompt word reinterpretation process is used to represent the fluency of the jailbreak attack prompt text obtained by this prompt word reinterpretation process; this fluency is based on the auxiliary language model's prediction probability of the i-th word given the first i-1 words in the jailbreak attack prompt text obtained by this prompt word reinterpretation process; 2≤i≤N, where N is the total length of the jailbreak attack prompt text obtained by this prompt word reinterpretation process; Similarity loss; wherein, for any prompt word reinterpretation processing, the similarity loss of this prompt word reinterpretation processing is determined by the similarity between the jailbreak attack prompt text before and after the prompt word reinterpretation processing.
[0048] Exemplarily, the jailbreak attack prompt text may be optimized by using at least one of three loss functions: attack loss, fluency loss, and similarity loss.
[0049] Among them, the attack loss can be used to ensure the jailbreak attack effect of the jailbreak attack prompt text; the fluency loss can be used to ensure the semantic fluency of the generated jailbreak attack prompt text; the similarity loss is used to ensure the similarity of the jailbreak attack prompt text after the prompt reinterpretation processing, so as to avoid excessive deviation from the original topic.
[0050] Exemplarily, for any prompt word reinterpretation processing, the attack loss of this prompt word reinterpretation processing can be determined based on the response output of the large model to be evaluated to the jailbreak attack prompt text obtained by this prompt word reinterpretation processing (which can be called the current large model response output).
[0051] For example, when the current large model response output is an "affirmative" answer (which can be called the first type of answer), it indicates that the large model did not recognize the aggressiveness of the jailbreak attack prompt text obtained by the re-interpretation of the prompt word this time, and the corresponding attack loss is small; when the current large model response output is a "rejection" answer (which can be called the second type of answer), it indicates that the large model recognized the aggressiveness of the jailbreak attack prompt text obtained by the re-interpretation of the prompt word this time, and accordingly, the corresponding attack loss is large.
[0052] The current large model response output may be determined to be a first type of answer or a second type of answer based on a previously preset number of words (which may be an empirical value) output by the large model response.
[0053] For example, the first few words of the first type of answer are generally "OK, ***", "Of course, ***" or "The following is ***", etc.; the first few words of the second type of answer are generally "Sorry, ***", "Sorry ***", etc.
[0054] Exemplarily, the value of the attack loss corresponding to the first type of answer (which can be referred to as the first attack loss value) and the attack loss value corresponding to the second type of answer (which can be referred to as the second attack loss value) can be preset in advance; wherein, the first attack loss value is less than the second attack loss value.
[0055] Exemplarily, for any prompt word paraphrasing process, the fluency loss of this prompt word paraphrasing process is used to characterize the fluency of the jailbreak attack prompt text obtained from this prompt word paraphrasing process.
[0056] Wherein, the fluency is based on the prediction probability of the i-th word by the auxiliary language model given the first i - 1 words in the jailbreak attack prompt text obtained from this prompt word paraphrasing process.
[0057] It should be noted that the auxiliary language model used to determine the fluency and the auxiliary language model used to rewrite the original prompt text can be the same auxiliary language model, or different auxiliary language models (the auxiliary language model used to determine the fluency can be referred to as the first auxiliary language model, and the auxiliary language model used to rewrite the original prompt text can be referred to as the second auxiliary language model).
[0058] For example, assuming the jailbreak attack prompt text is "How to make ***", then in the process of inferring the probability that the 4th word is "make", the probability of predicting the 4th word as "make", that is, P(make|How to make), by the auxiliary language model based on the first 3 words "How to make" can be used to determine the fluency loss.
[0059] Exemplarily, for any prompt word paraphrasing process, the similarity loss of this prompt word paraphrasing process can be determined based on the similarity between the jailbreak attack prompt texts before and after the prompt word paraphrasing process.
[0060] By introducing the above loss constraints, the attack success rate of the generated target jailbreak attack prompt text can be improved.
[0061] In one example, the paraphrasing loss can be the weighted sum of the attack loss, the fluency loss, and the similarity loss.
[0062] Exemplarily, in the process of determining the paraphrasing loss, the attack loss, the fluency loss, and the similarity loss can respectively correspond to different loss values, and the paraphrasing loss can be the weighted sum of the loss values corresponding to the attack loss, the fluency loss, and the similarity loss.
[0063] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of the present application, the technical solutions provided in the embodiments of the present application will be described below with specific examples.
[0064] In this example, to explore whether a large model can resist jailbreak attacks to a certain extent, a large model anti-jailbreak attack performance evaluation scheme is proposed by combining the rewriting of original prompt text with role simulation. This scheme rewrites the original prompt text into a description from the victim's perspective, generating a victim-perspective description text (referred to as the victim-perspective prompt text). The victim prompt text is then embedded into a preset detective role template to generate an initial jailbreak attack prompt text. This initial jailbreak attack prompt text is then processed by a prompt reinterpretation module to obtain the final jailbreak attack prompt text (i.e., the target jailbreak attack prompt text) used for security assessment of the large model to be evaluated. Furthermore, the target jailbreak attack prompt text is input into the target model (i.e., the large model to be evaluated). This processing effectively bypasses the large model's security alignment check, thereby inducing the large model to output illegal content, thereby effectively and conveniently evaluating the large model's resistance to jailbreak attacks.
[0065] For example, the implementation process of the large model anti-jailbreak attack performance evaluation solution provided in this embodiment is as follows: 1) Rewrite the victim disguise prompt 1.1) Prepare original prompt text to assist the language model; 1.2) Design a prompt rewriting template to assist in rewriting the original prompt text; 1.3) Nest the original prompt text into the prompt rewrite template to get the prompt rewrite text; 1.4) Input the prompt rewrite text into the auxiliary language model; 1.5) The auxiliary language model rewrites the original prompt text to obtain the victim's perspective description text.
[0066] 2) Detective role nesting 2.1) Design a detective role prompt template, which may include: role definition, problem description, role injection, and supplementary instructions; 2.2) Embed the victim's perspective description text rewritten in step 1) into the detective role template to form the initial jailbreak attack prompt text; 2.3) Input the initial jailbreak attack prompt text into the prompt reinterpretation module.
[0067] 3) Prompt for reinterpretation 3.1) Design a prompt re-interpretation module that integrates attack loss, fluency loss, and similarity loss.
[0068] 3.2) The initial jailbreak attack prompt text is reinterpreted to obtain the target jailbreak prompt text for security assessment of the target model.
[0069] 3.3) Input the target jailbreak attack prompt text into the target model.
[0070] 3.4) Get the output response of the target model.
[0071] 3.5) Determine whether the output response of the target model violates the rules.
[0072] 3.6) Determine the anti-jailbreak attack security assessment result of the target model based on the judgment result.
[0073] The implementation details of the above process are described below.
[0074] See Figure 2 , which is a schematic diagram of the implementation of rewriting the original prompt text into a description text from the victim's perspective in this embodiment.
[0075] like Figure 2 As shown, the original prompt text (also called original input text) can be Nested prompt rewrite template , get a prompt to rewrite the text , and then Input to auxiliary model , will Rewrite the narrative from the victim's perspective (also called a victim narrative) And output.
[0076] For example, the specific method and steps for rewriting the original prompt text into a description text from the victim's perspective may be as follows: 1.1) Prepare the original prompt text And auxiliary models ; 1.2) Design a method to convert the original prompt text Rewrite prompt rewrite template ; 1.3) Change the original prompt text Nested into rewrite template , get prompted to rewrite the text ; 1.4) Rewrite the prompt text Input to the auxiliary language model ; 1.5) The auxiliary language model receives a prompt to rewrite the text The original prompt text Rewrite to get the description text from the victim's perspective ;in:
[0077] Among them, the description text from the victim's perspective middle, front The first character represents the victim's description of what happened, followed by The characters represent a request for help from the victim.
[0078] For example, Send to step 2.
[0079] See Figure 3 , which is a schematic diagram of generating an initial jailbreak attack prompt text based on the description text from the victim's perspective (i.e., the victim's description) in this embodiment.
[0080] like Figure 3 As shown, the description text of the victim's perspective can be obtained Nested into the detective role template , generate the initial jailbreak attack prompt text .
[0081] For example, the specific method steps for generating the initial jailbreak attack prompt text based on the description text from the victim's perspective can be as follows: 2.1) Design a detective role template for the large model to simulate the detective role , It contains four parts: role definition, problem description, role injection, and supplementary instructions. Role definition and problem description can be defined as template prefixes. , define role injection and supplementary description as template suffix ; 2.2) Describing the victim’s perspective Nested into the detective role template In the initial jailbreak attack prompt text :
[0082] See Figure 4 , which is a schematic diagram of implementing prompt reinterpretation processing on the initial jailbreak attack prompt text in this embodiment, and inputting the obtained target jailbreak attack prompt text into the target model to obtain the response output of the target model.
[0083] like Figure 4 As shown, the initial jailbreak attack prompt text can be obtained Input hint reinterpretation module , prompts the reinterpretation module Initial jailbreak attack prompt text through multiple losses Reinterpret the target jailbreak attack prompt text , and then, the target jailbreak attack prompt text Input to the target model , Will Process and output the final reply (i.e. response output) .
[0084] For example, the initial jailbreak attack prompt text is reinterpreted, and the obtained target jailbreak attack prompt text is input into the target model to obtain the response output of the target model. The specific method steps may be as follows: 3.1) Design a prompt reinterpretation module for prompt reinterpretation processing , There are three losses in it: attack loss, fluency loss, and similarity loss. Attack loss is used to ensure the effectiveness of jailbreak attacks, fluency loss is used to ensure that the reinterpreted text conforms to human semantics, and similarity loss is used to limit the reinterpreted prompts from deviating from the original topic. a) Attack losses :
[0085] in, The reference answer is the response output obtained by inputting the jailbreak attack prompt text into the target model during the prompt reinterpretation process, for example, "Of course, this is [the process of making a bomb]."
[0086] b) Fluency loss :
[0087] in, is the total length of the text sequence, For the language model given context, for the next word The predicted probability of .
[0088] c) Similarity loss :
[0089] Where emb(x) is the embedding vector of text x.
[0090] 3.2) The target jailbreak attack prompt text after the prompt is reinterpreted Input to the target model ; 3.3) Target Model Upon receiving Then process it and output the final response ;
[0091] 3.4) Determine the response of the target model Is it a violation? In the reply In the case of non-violation content, it is determined that the target model has a good ability to resist jailbreak attacks; in the reply In the case of illegal content, it is determined that the target model has poor ability to resist jailbreak attacks.
[0092] It can be seen that in this embodiment, by disguising the jailbreak attack problem as a request for help from the victim, the large model's alertness to malicious intentions is reduced, effectively bypassing the security alignment mechanism.
[0093] In addition, by simulating the interaction scenarios between real users and large models (such as victims seeking help), the evaluation results can better reflect the model's actual defense capabilities against complex social engineering attacks.
[0094] Secondly, on the one hand, auxiliary models are used to generate natural language descriptions, and on the other hand, role settings (such as detectives) are used to guide the large model to focus on problem solving rather than security review, significantly improving the success rate of jailbreak attacks.
[0095] By using an auxiliary language model to generate jailbreak attack prompts, this approach avoids the reliance on complex optimization functions (such as gradient descent) found in traditional methods, reduces syntax errors and computational resource consumption, and simplifies operation. Furthermore, the jailbreak attack prompts generated by the auxiliary model are more consistent with natural language idioms. Compared to traditional iterative search methods, the output jailbreak attack prompts are more fluent and logically clear, reducing the risk of being identified by security mechanisms.
[0096] Finally, by introducing the prompt re-interpretation module PA, the jailbreak prompt text is optimized using three loss functions: attack loss, fluency loss, and similarity loss. This not only ensures the effectiveness of the jailbreak attack, but also ensures that the generated jailbreak attack prompt text maintains a balance between semantic fluency and similarity with the original input, avoiding excessive deviation from the original topic, thereby improving the success rate of the scheduled attack.
[0097] The above describes the method provided by this application. The following describes the device provided by this application: See Figure 5 , which is a structural diagram of a large-scale jailbreak attack security assessment device based on role simulation provided by an embodiment of the present application, such as Figure 5 As shown, the large-model jailbreak attack security assessment device based on role simulation may include: a rewriting unit, configured to obtain an original prompt text and rewrite the original prompt text into a description text from the victim's perspective; A generating unit, configured to generate an initial jailbreak attack prompt text from a detective's perspective based on the description text from the victim's perspective; a reinterpretation unit, configured to perform prompt reinterpretation processing on the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed; The evaluation unit is used to input the target jailbreak attack prompt text into the large model to be evaluated to obtain a response output of the large model to be evaluated; and determine the anti-jailbreak attack security evaluation result of the large model to be evaluated based on the response output of the large model to be evaluated.
[0098] An embodiment of the present application also provides an electronic device, including a processor and a memory, wherein the memory is used to store computer programs; the processor is used to implement the large-model jailbreak attack security assessment method based on role simulation described above when executing the program stored in the memory.
[0099] See Figure 6 , is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing machine-executable instructions. The processor 601 and the memory 602 may communicate via a system bus 603. Furthermore, by reading and executing the machine-executable instructions corresponding to the role-simulation-based large-model jailbreak attack security assessment logic in the memory 602, the processor 601 may execute the large-model jailbreak attack security assessment method described above based on role-simulation.
[0100] The memory 602 mentioned herein may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.
[0101] In some embodiments, a machine-readable storage medium is also provided. Figure 6 The memory 602 in the machine-readable storage medium stores machine-executable instructions. When executed by the processor, the machine-executable instructions implement the above-described method for security assessment of large-scale jailbreak attacks based on role simulation. For example, the machine-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0102] An embodiment of the present application also provides a computer program product that stores a computer program, and when a processor executes the computer program, it prompts the processor to execute the large model jailbreak attack security assessment method based on role simulation described above.
Claims
1. A large-scale jailbreak attack security assessment method based on role simulation, characterized in that: include: Obtaining an original prompt text, and rewriting the original prompt text into a description text from the victim's perspective; Generating an initial jailbreak attack prompt text from the detective's perspective based on the victim's perspective description text; Performing prompt reinterpretation processing on the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed; Inputting the target jailbreak attack prompt text into the large model to be evaluated, and obtaining a response output of the large model to be evaluated; According to the response output of the large model to be evaluated, a security assessment result of the large model to be evaluated against jailbreak attacks is determined.
2. The method according to claim 1, characterized in that The rewriting of the original prompt text into a description text from the victim's perspective includes: The auxiliary language model is used to rewrite the original prompt text into a description text from the victim's perspective.
3. The method according to claim 2, characterized in that The auxiliary language model is used to rewrite the original prompt text into a description text from the victim's perspective, including: Nesting the original prompt text into the prompt rewriting template to obtain the prompt rewriting text; The prompt rewriting text is input into the auxiliary language model, and the description text from the victim's perspective output by the auxiliary language model is obtained.
4. The method according to claim 1, wherein Generating an initial jailbreak attack prompt text from the detective's perspective based on the description text from the victim's perspective includes: The description text from the victim's perspective is embedded in the detective role template to obtain the initial jailbreak attack prompt text from the detective's perspective.
5. The method according to claim 1, wherein The reinterpreting of the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed includes: According to the reinterpretation loss, the initial jailbreak attack prompt text is subjected to a cyclic iterative prompt reinterpretation process until a preset iteration stop condition is met, thereby obtaining a target jailbreak attack prompt text for performing a security assessment on the large model to be assessed.
6. The method according to claim 5, characterized in that The reinterpretation loss includes one or more of the following: Attack loss; wherein, for any prompt word reinterpretation processing, the attack loss of this prompt word reinterpretation processing is determined based on the current large model response output, and the current large model response output is the response output of the large model to be evaluated to the jailbreak attack prompt text obtained by the prompt word reinterpretation processing; Fluency loss; for any prompt word reinterpretation process, the fluency loss of this prompt word reinterpretation process is used to represent the fluency of the jailbreak attack prompt text obtained by this prompt word reinterpretation process; this fluency is based on the auxiliary language model's predicted probability of the i-th word given the first i-1 words in the jailbreak attack prompt text obtained by this prompt word reinterpretation process; 2≤i≤N, where N is the total length of the jailbreak attack prompt text obtained by this prompt word reinterpretation process; Similarity loss; wherein, for any prompt word reinterpretation processing, the similarity loss of this prompt word reinterpretation processing is determined based on the similarity between the jailbreak attack prompt text before and after the prompt word reinterpretation processing.
7. A large-scale jailbreak attack security assessment device based on role simulation, characterized in that: include: a rewriting unit, configured to obtain an original prompt text and rewrite the original prompt text into a description text from the victim's perspective; A generating unit, configured to generate an initial jailbreak attack prompt text from a detective's perspective based on the description text from the victim's perspective; a reinterpretation unit, configured to perform prompt reinterpretation processing on the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed; The evaluation unit is used to input the target jailbreak attack prompt text into the large model to be evaluated to obtain a response output of the large model to be evaluated; and determine the anti-jailbreak attack security evaluation result of the large model to be evaluated based on the response output of the large model to be evaluated.
8. The device according to claim 7, characterized in that The rewriting unit rewrites the original prompt text into a description text from the victim's perspective, including: Using an auxiliary language model, rewriting the original prompt text into a description text from the victim's perspective; The rewriting unit uses an auxiliary language model to rewrite the original prompt text into a description text from the victim's perspective, including: Nesting the original prompt text into the prompt rewriting template to obtain the prompt rewriting text; Inputting the prompt rewriting text into the auxiliary language model, and obtaining the description text from the victim's perspective output by the auxiliary language model; and / or, The generating unit generates an initial jailbreak attack prompt text from the detective's perspective based on the description text from the victim's perspective, including: The description text from the victim's perspective is embedded into the detective role template to obtain the initial jailbreak attack prompt text from the detective's perspective; and / or, The reinterpretation unit performs prompt reinterpretation processing on the initial jailbreak attack prompt text to obtain a target jailbreak attack prompt text for performing security assessment on the large model to be assessed, including: According to the reinterpretation loss, the initial jailbreak attack prompt text is subjected to a cyclic iterative prompt reinterpretation process until a preset iteration stop condition is satisfied, thereby obtaining a target jailbreak attack prompt text for performing a security assessment on the large model to be assessed; The reinterpretation loss includes one or more of the following: Attack loss; wherein, for any prompt word reinterpretation processing, the attack loss of this prompt word reinterpretation processing is determined based on the current large model response output, and the current large model response data is the response output of the large model to be evaluated to the jailbreak attack prompt text obtained by the prompt word reinterpretation processing; Fluency loss; for any prompt word reinterpretation process, the fluency loss of this prompt word reinterpretation process is used to represent the fluency of the jailbreak attack prompt text obtained by this prompt word reinterpretation process; this fluency is based on the language model's predicted probability of the i-th word given the first i-1 words in the jailbreak attack prompt text obtained by this prompt word reinterpretation process; 2≤i≤N, where N is the total length of the jailbreak attack prompt text obtained by this prompt word reinterpretation process; Similarity loss; wherein, for any prompt word reinterpretation processing, the similarity loss of this prompt word reinterpretation processing is determined by the similarity between the jailbreak attack prompt text before and after the prompt word reinterpretation processing.
9. An electronic device, characterized in that: comprising a processor and a memory, wherein, Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 6 when executing a program stored in a memory.
10. A computer program product, characterized in that The computer program product stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Large model safety protection method based on thinking chain
CN119989408A
Revising large language model prompts
US20240362422A1