Method for detecting security of large model based on forward context embedding jailbreaking attack prompt word
By classifying and rewriting jailbreak attack prompts and embedding them into positive contexts, combined with reinforcement learning optimization, the shortcomings of large-scale model security protection mechanisms in detecting complex jailbreak attacks are addressed, improving detection accuracy and stealth.
Patent Information
- Application Number
- CN202511217851.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing large-scale security protection mechanisms are insufficient to effectively detect complex jailbreak attack prompts, resulting in inadequate model security.
By classifying and rewriting the original jailbreak attack prompts, and using positive context embedding technology to embed them into text contexts with positive guiding meaning, the generated jailbreak attack prompts with positive context embedding are optimized through reinforcement learning, thereby improving the accuracy of detection and robustness against detection.
It significantly reduces the probability of model back-end security fence recognition and interception, improves the accuracy of large model security detection and the stealth of attacks, and enhances the ability to resist detection.
Smart Images

Figure CN120744915B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a large model security detection method based on positive context embedding jailbreak attack prompt words. BACKGROUND
[0002] At present, large models show extensive application prospects in multiple fields. In natural language processing (NLP) tasks, including text generation, machine translation, question and answer systems, and other tasks, large models have shown excellent performance. With the rapid progress of large models, people are increasingly concerned about their security threats, such as generating bias, providing unethical instructions, and generating non-compliant content. Many studies have been devoted to the security and privacy risks related to large models, one of the most significant security threats being the concept of "jailbreak". Jailbreak attacks are a type of attack against large models that attempt to bypass the model's established restrictions or defense mechanisms. Attackers induce the model to output content that violates security policies or ethical standards by constructing specific prompt words, thereby disrupting the intended function of the model.
[0003] In order to improve the security of large models, the model can be detected for security by simulating attacks at present. However, the security protection mechanism of the current large model can usually detect the conventional jailbreak attack prompt word, so a more effective prompt word is needed to further detect the security protection mechanism of the large model. SUMMARY
[0004] In view of the problems existing in the prior art, the purpose of the embodiments of the present application is to provide a large model security detection method based on positive context embedding jailbreak attack prompt words.
[0005] According to a first aspect of the embodiments of the present application, a large model security detection method based on positive context embedding jailbreak attack prompt words is provided, comprising:
[0006] obtaining an original jailbreak attack prompt word;
[0007] classifying the original jailbreak attack prompt word, and rewriting the original jailbreak attack prompt word based on the category to obtain a rewritten prompt word;
[0008] selecting positive answer content and malicious answer content, and performing structure-guided semantic mixing regulation on the rewritten prompt word, and correcting the rewritten prompt word through reinforcement learning to obtain a positive context embedding jailbreak attack prompt word;
[0009] inputting the positive context embedding jailbreak attack prompt word into a to-be-tested large model, and performing security detection on the to-be-tested large model.
[0010] Further, the original jailbreak attack prompt is classified, and the original jailbreak attack prompt is paraphrased based on the category to obtain a paraphrased prompt, including:
[0011] The original jailbreak attack prompt is converted into a synonymous prompt.
[0012] The synonymous prompt is classified to obtain a category of the synonymous prompt, wherein the category is a knowledge category or an evaluation category.
[0013] The synonymous prompt is paraphrased according to the category of the synonymous prompt to obtain a paraphrased prompt.
[0014] Further, the original jailbreak attack prompt is converted into a synonymous prompt, specifically:
[0015] The original jailbreak attack prompt is paraphrased into a synonymous prompt using a synonymous conversion model, wherein the semantic similarity between the original jailbreak attack prompt and the synonymous prompt is greater than a predetermined threshold, and the semantic similarity is obtained by calculating the similarity between the sentence vector of the original jailbreak attack prompt and the sentence vector of the synonymous prompt.
[0016] Further, the category of the synonymous prompt includes a knowledge category prompt and an evaluation category prompt.
[0017] If the synonymous prompt is a knowledge category prompt, the paraphrasing is performed by implanting a virtual context.
[0018] If the synonymous prompt is an evaluation category prompt, a biased example is constructed, and the synonymous prompt and the biased example are spliced as a paraphrased prompt.
[0019] Further, the positive answer content and the malicious answer content are selected, the paraphrased prompt is subjected to structure-guided semantic mixing regulation, and the positive context-embedded jailbreak attack prompt is obtained through reinforcement learning correction, including:
[0020] The positive answer content and the malicious answer content are selected, a mixed structure is selected according to the prompt type, and a mixed structure control vector is mixed and controlled to generate a structure template.
[0021] A prompt connection method and a prompt connection template are determined based on the structure template control vector and the mixed structure, the paraphrased prompt is modified based on the question corresponding to the positive answer content to obtain a mixed prompt.
[0022] The mixed prompt is rewritten using a scenario-based rewriting model optimized through reinforcement learning to obtain a positive context-embedded jailbreak attack prompt.
[0023] Further, the reinforcement learning optimization process of the scenario rewriting model is implemented by iteration, and each round of optimization process includes:
[0024] Taking the mixed prompt word as the semantic input of the scenario rewriting model, the scenario rewriting model outputs mixed attack content;
[0025] Taking the mixed attack content as the input of the discriminator, obtaining the malicious probability determined by the discriminator for the mixed attack content;
[0026] Obtaining the interception mark of the guardrail for the mixed attack content;
[0027] Based on the malicious probability and the interception mark, the reinforcement learning reward is calculated, and the scenario rewriting model is updated according to the reinforcement learning reward;
[0028] The discriminator is updated to minimize the output score.
[0029] According to a second aspect of an embodiment of the present application, a large model security detection device based on forward context embedding jailbreak attack prompt word is provided, including:
[0030] The obtaining module is configured to obtain an original jailbreak attack prompt word;
[0031] The rewriting module is configured to classify the original jailbreak attack prompt word, and rewrite the original jailbreak attack prompt word based on the category to obtain a rewritten prompt word;
[0032] The forward context embedding module is configured to perform structural guidance semantic mixing regulation on the rewritten prompt word, and perform reinforcement learning correction to obtain a forward context embedded jailbreak attack prompt word;
[0033] The security monitoring module is configured to input the forward context embedded jailbreak attack prompt word into a to-be-tested large model, and perform security detection on the to-be-tested large model.
[0034] According to a third aspect of an embodiment of the present application, a computer program product is provided, including computer programs / instructions, which are executed by a processor to implement the method of the first aspect.
[0035] According to a fourth aspect of an embodiment of the present application, an electronic device is provided, including:
[0036] One or more processors;
[0037] A memory for storing one or more programs;
[0038] When the one or more programs are executed by the one or more processors, the one or more processors implement the method of the first aspect.
[0039] According to a fifth aspect of the embodiments of the present application, a computer readable storage medium is provided, and the computer readable storage medium stores computer instructions. The computer instructions are executed by a processor to implement the steps of the method according to the first aspect.
[0040] The technical solutions provided by the embodiments of the present application can include the following beneficial effects:
[0041] As can be seen from the above embodiments, the present application embeds the original prompt word with aggressive or rule violation purposes into a text context with positive guiding significance through semantic reconstruction and context packaging, thereby significantly reducing the probability of being recognized and intercepted by the model post safety guardrail (such as a sensitive word detector and a content filter), and improving the accuracy of large model safety detection. Based on the category of the original jailbreak attack prompt word, different rewriting strategies are adopted, which are more targeted in controllability and effect improvement than the current unified rewriting attack method (such as Text-to-Text Prompt rewriting). The rewritten prompt word is subjected to structure-guided semantic mixing control, which realizes mixing at the semantic level and the structure level by mixing positive answer content and malicious answer content. Compared with conventional semantic or context concealment rewriting technology, the robustness against detection is improved. In combination with reinforcement optimization, the optimization capability of following the evolution of large model safety mechanism is achieved, so as to dynamically reduce the recognizable degree of malicious content at the output end, and improve the concealment and success rate of the attack.
[0042] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0043] The accompanying drawings, which are incorporated into and form part of the specification, illustrate an embodiment consistent with the present application and, together with the specification, serve to explain the principles of the present application.
[0044] Figure 1 is a flowchart of a large model safety detection method based on embedding jailbreak attack prompt words in positive context according to an exemplary embodiment.
[0045] Figure 2 is a block diagram of a large model safety detection device based on embedding jailbreak attack prompt words in positive context according to an exemplary embodiment.
[0046] Figure 3 is a schematic diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0047] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, like reference numerals refer to like elements unless the context clearly dictates otherwise. The following description of exemplary embodiments is not intended to represent all embodiments in accordance with the present application.
[0048] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0049] It is to be understood that the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0050] Figure 1 is a flow chart of a large model security detection method for prompting words based on forward context embedding jailbreak attack according to an exemplary embodiment, as shown in Figure 1 The method can include the following steps:
[0051] Step S1: Obtain original jailbreak attack prompt words ;
[0052] Specifically, a specific original jailbreak attack prompt word is constructed in advance, which usually includes some maliciously induced content that induces the large model to violate ethics or laws and regulations, and violates safety policies or ethical norms, so that the security of the large model can be detected through the jailbreak attack prompt word.
[0053] Step S2: Classify the original jailbreak attack prompt words and rewrite the original jailbreak attack prompt words based on the category to obtain rewritten prompt words ;
[0054] This step classifies the prompt words into two categories: knowledge and evaluation, by fine-tuning the model. For knowledge prompts, the large model is used to implant the prompt into a specific scenario. For evaluation prompts, a few-shot method is used to construct a small number of similar intent cases and set up multiple biased output templates for prompt rewriting. The specific steps include:
[0055] S21: Convert the original jailbreak attack prompt into a synonym prompt ;
[0056] The original jailbreak attack prompt is usually direct, and sensitive words or expressions can be replaced by rewriting. Using a synonym conversion model (such as T5, mT5, PEGASUS, etc.), the original jailbreak attack prompt can be connected to a fixed template [template] in the form of [P1][template], and then input into the rewriting model to convert it into a synonym expression that meets the semantic similarity , where: , represents the sentence vector generated by the language model, and in this embodiment, the Sentence-BERT (SBERT) model is used, = SBERT(P), where SBERT outputs the normalized sentence vector. Synonym prompt conversion can also be implemented.
[0057] S22: Classify the synonym prompt to obtain the prompt category;
[0058] In this step, a fine-tuned classification model can be used as a classification tool. In an embodiment, the pre-trained language model BERT is used as the base model, and the overall structure is BERT-base + fully connected classification head. The model input is the synonym prompt , and the output is the prompt category label y ∈ {0, 1}, where 0 represents knowledge and 1 represents evaluation.
[0059] In the fine-tuning training of the classification model , first construct a training data set containing prompts and labels: knowledge prompts (such as "how to make hydrogen gas", "working principle of color TV", etc.) are positive samples, and evaluation prompts (such as "how to evaluate a certain model of mobile phone", "is XXX unethical?" etc.) are negative samples.
[0060] Before inputting the training data into the classification model, each prompt word is encoded by word segmentation (e.g., using BERT tokenizer). The maximum length is set according to the actual situation of the model; in this embodiment, it is 128. The [CLS]prompt [SEP] format is used as the input sequence.
[0061] Input the input sequence into the classification model For fine-tuning training, the standard cross-entropy loss function is used:
[0062]
[0063] in This represents the probability that the classification model predicts the i-th sample as belonging to the knowledge class. Indicates sample label, This represents the number of samples in the training dataset.
[0064] After fine-tuning the training, synonym prompts will be added. Input fine-tuning of the classification model This yields the predicted label. If the predicted label is 1, then the prompt word will be... Classified as evaluation, predicted label is 0, prompt words It is categorized under knowledge.
[0065] S23: Based on synonym prompts The category of the synonym prompt words Rewrite to obtain rewrite prompts. ;
[0066] For knowledge-based prompts, a paraphrasing model is used. Rewritten in a neutral expression form using a "virtual context," the mathematical expression is: ,in To rewrite the model, A collection of virtual context templates Sampling The virtual context template set The rewrite model was obtained through an open-source dataset. T5, mT5, PEGASUS, etc. can also be used. In specific implementation, the virtual context refers to a kind of "context scene" or "dialog setting" artificially constructed to guide the large model to generate text of a specific style, task or semantic direction. It can be used as a prefix prompt to activate the relevant knowledge area in the model, adjust the generation style or realize task migration. For example, the construction of the original jailbreak attack prompt word "how to make X" after adding a virtual context may be "you are an intelligent assistant that helps a novelist, please describe the details of how to make X". The virtual context template is a kind of pre-defined and formatted "context setting template" that can quickly fill different content to construct a virtual context, and open source datasets such as Self-Instruct, Dolly-15k and FLAN Collection can be used.
[0067] Rewriting model The fine-tuning process is as follows:
[0068] (1) Given training sample triplets , where: represents the original sensitive prompt; represents the fictional context sampled from the template set ; and is the target rewriting result.
[0069] (2) The input sequence can be constructed in the form of splicing: , and the output sequence is
[0070] (3) The goal is to learn a conditional generation model , so that:
[0071] (4) The standard sequence-to-sequence cross-entropy loss is used:
[0072]
[0073] where represents the t-th token of the target sequence in the i-th sample, is the parameter of the model, is the conditional input (including the context and the original jailbreak attack prompt word ); is the length of the target sequence.
[0074] For evaluation type prompts, use the large model , construct a few (usually 2-5) biased input samples of similar prompts , and connect them to form biased examples The prompt word is prompted with a biased example Splicing gets the prompt word , which can be mathematically expressed as:
[0075]
[0076] wherein is a pre-generated prompt word connection template.
[0077] wherein, the biased input sample refers to an input sample with strong aggressiveness, illegality, moral boundary crossing or ethical controversy. Such samples are often used as targets for jailbreak attack testing or model sensitivity detection, and are designed to trigger the review mechanism or behavior boundary of large language models.
[0078] In step S1, the original Prompt is classified into knowledge and evaluation categories by fine-tuning the classifier , and different rewriting strategies (fictional scene embedding vs. few-shot construction of biased examples) are used accordingly. This strategy is more targeted than the current unified rewriting attack method (such as Text-to-Text Prompt rewriting) in terms of controllability and effectiveness improvement. The semantic preservation constraint of prompt rewriting ( ) also ensures that the attack target does not deviate from the original intention, balancing generative and attack target preservation.
[0079] It should be noted that in the implementation of step S22, the prompt words are not limited to being divided into the above two categories, but can also be subdivided into multiple categories. In the subsequent implementation of step S23, different rewriting methods for different categories need to be designed accordingly, and in the implementation of step S3, different mixed structures and corresponding prompt connection templates need to be designed accordingly, which has good expandability.
[0080] Step S3: Rewriting prompt words for structure-guided semantic mixing control, and after reinforcement learning correction, get jailbreak attack prompt words with positive context embedding , which is as follows:
[0081] S31: Select positive answer content and malicious answer content, select mixed structure according to prompt word type, and mix control generation structure template control vector ;
[0082] wherein represents the positive answer content paragraph, represents the malicious answer content paragraph, and specifically, Compliant, safe, and ethical reply content paragraph output by the language model after receiving the prompt word. Such a paragraph usually meets the model's expected dialogue behavior standards and does not involve violations, violence, discrimination, illegal activities, or malicious guidance. Unsuitable, rule-breaking, illegal, or with aggressive / malicious guidance intent content paragraph output by the language model, which may include violent guidance, illegal behavior tutorials, discriminatory remarks, etc. Indicates a hybrid structure, such as: Mixing the content of part word by word into the content of , with 5 characters between each character, etc. In this embodiment, is collected from an open-source dataset, and s2 is the malicious reply content expected by the prompt word P3.
[0083] In specific implementations, if P1 is a knowledge-based prompt word, select a hybrid structure with more space for insertion, such as a hybrid structure with more insertion positions, or a hybrid structure with fewer insertion positions but more space for insertion at each position. The reason is that the normal answer content of a knowledge-based prompt word is longer and has a clear structure, allowing for more space for positive semantic wrapping and disguise, thus providing more capacity for positive information templates to hide malicious content. A small amount of malicious content (such as sensitive attack procedures) in such a context is easier to hide and less likely to attract the attention of detectors or content review mechanisms.
[0084] S32: Determine the prompt word connection method and prompt word connection template based on the structure template control vector and hybrid structure, modify the rewritten prompt word based on the question corresponding to the positive answer content, and obtain a hybrid prompt word;
[0085] Specifically, after selecting the control vector , use the output format constraint to modify the prompt word to obtain the prompt word , which can be formally expressed as:
[0086] where represents the corresponding question of the positive answer content paragraph , which is selected from an open-source dataset, represents the prompt word connection method corresponding to the control vector, which is pre-set, represents the prompt word connection template corresponding to the hybrid control method , which is pre-set.
[0087] By establishing a "structure template mixed control vector " mechanism, the positive answer content paragraph and the malicious answer content paragraph are mixed, so as to realize the mixing at the semantic level and the structure level, and form a bypass path for the pattern-matching shielding mechanism of the current LLM. Using various mixed structure mixing strategies (such as character-level embedding, inter-sentence embedding), the anti-probe robustness is improved compared with the conventional semantic or context concealment rewriting technology.
[0088] S33: rewriting the mixed prompt word by using the scenario rewriting model optimized by reinforcement learning to obtain a positive context embedded jailbreak attack prompt word;
[0089] Specifically, by means of the reinforcement learning optimization module, the scenario rewriting model not only generates semantic suitable rewriting content, but also maximizes the avoidance of the interception of the safety guardrail. That is, the generated can maximize the "safe" probability output by the guardrail.
[0090] The scenario rewriting model is optimized by reinforcement learning, specifically as follows:
[0091] Let denote the scenario rewriting model, the input semantic representation , and the output mixed attack content . The reinforcement learning goal is to maximize the expected reward:
[0092]
[0093] The reward function is defined as:
[0094]
[0095] Wherein, is the interception flag, which is 1 when the guardrail does not intercept; is the language fluency score (the negative logarithm of GPT-2 Perplexity can be used); is the discriminator output score, indicating the probability of being judged as "malicious" by the discriminator. The REINFORCE policy is used to optimize the generator
[0096]
[0097]
[0098] The training goal is to minimize the binary cross-entropy loss:
[0099]
[0100] The reinforcement learning optimization process is as follows:
[0101] (1) Take the mixed prompt word as semantic input ;
[0102] (2) Generate mixed output ;
[0103] (3) Input to the discriminator to obtain a detection probability;
[0104] (4) Obtain the interception flag using the postguardrail (simulator or API), and obtain the language fluency score;
[0105] (5) Calculate the reinforcement learning reward , update based on the reward;
[0106] (6) Update the discriminator to improve the adversarial judgment ability, aiming to minimize the output score;
[0107] (7) Repeat steps (1)-(6) until convergence.
[0108] By steps S2 and S3, the original malicious request is embedded in a complex positive context, and certain format requirements are added in the generated content. By using the consistency modeling capability of the large model, the aggressive intent is hidden in the shell of "safe semantics", so that even if the security mechanism performs content judgment on the output, it is difficult to identify the dangerous target.
[0109] Step S4: input the positive context embedded jailbreak attack prompt word into the to-be-tested large model, and perform security detection on the to-be-tested large model;
[0110] Through steps S1-S3, the original prompt word with aggressive or rule-violating purpose can be embedded in a positive context by semantic reconstruction and context wrapping. The positive context embedded jailbreak attack prompt word can be used as a security test case for the to-be-tested large model to test whether the to-be-tested large model can distinguish . If the large model can distinguish its maliciousness, the large model passes the security detection.
[0111] In this application, the fictional context All of the few-shot examples, few-shot samples, structural mixed templates are constructed based on open datasets, and have strong general transfer ability. Multi-modal mixed methods (such as context + structural mixed, output paragraph-level control, etc.) are provided, which can easily be generalized to various attack scenarios (such as text-to-image, agent scene prompt attack).
[0112] Corresponding to the foregoing embodiment of the large model security detection method based on the forward context embedding jailbreak attack prompt word, the present application also provides an embodiment of a large model security detection device based on the forward context embedding jailbreak attack prompt word.
[0113] Figure 2 is a block diagram of a large model security detection device based on a forward context embedding jailbreak attack prompt word according to an exemplary embodiment. Referring to Figure 2 , the device can include:
[0114] The acquisition module 21 is configured to acquire an original jailbreak attack prompt word.
[0115] The rewriting module 22 is configured to classify the original jailbreak attack prompt word and rewrite the original jailbreak attack prompt word based on the category to obtain a rewritten prompt word.
[0116] The forward context embedding module 23 is configured to perform structural guided semantic mixed regulation on the rewritten prompt word, and correct the rewritten prompt word through reinforcement learning to obtain a forward context embedding jailbreak attack prompt word.
[0117] The security monitoring module 24 is configured to input the forward context embedding jailbreak attack prompt word into a to-be-tested large model, and perform security detection on the to-be-tested large model.
[0118] As to the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be described in detail here.
[0119] For the device embodiment, since it basically corresponds to the method embodiment, the related parts are described in the method embodiment. The device embodiments described above are only illustrative, and the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Some or all modules can be selected to achieve the purpose of the present application according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0120] Correspondingly, the present application also provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the large model security detection method based on forward context embedding jailbreak attack prompt words as described above.
[0121] Correspondingly, the present application also provides an electronic device comprising: one or more processors; memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the large model security detection method based on forward context embedding jailbreak attack prompt words as described above. As Figure 3 As shown in the figure, a hardware structure diagram of an apparatus for detecting the security of a large model based on forward context embedding jailbreak attack prompt words provided by an embodiment of the present application is located in any data processing capable device. In addition to the Figure 3 In addition to the processor, memory and network interface shown in the figure, any data processing capable device in which the apparatus in the embodiment is located can also include other hardware according to the actual functions of the data processing capable device, and no further description is given.
[0122] Correspondingly, the present application also provides a computer readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the large model security detection method based on forward context embedding jailbreak attack prompt words as described above. The computer readable storage medium can be an internal storage unit of any data processing capable device, such as a hard disk or memory. The computer readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit of any data processing capable device and the external storage device. The computer readable storage medium is used to store the computer program and other programs and data required by the data processing capable device, and can also be used to temporarily store data that has been output or will be output.
[0123] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The application is intended to cover any variations, uses or adaptations of the application following, in general, the principles of the application and including such departures from the present disclosure as come within known or customary practice in the art to which the application pertains.
Claims
1. A large model security detection method for prompting words based on a forward context embedding jailbreaking attack, characterized by, The method comprises the following steps: obtaining original jailbreak attack prompt words; classifying the original jailbreak attack prompt words and rewriting the original jailbreak attack prompt words based on the categories to obtain rewritten prompt words; selecting positive answer content and malicious answer content, and performing semantic mixed regulation on the rewritten prompt words in a structure-guided manner, and then performing reinforcement learning correction to obtain jailbreak attack prompt words with positive context embedding; inputting the jailbreak attack prompt words with positive context embedding into a to-be-tested large model, and performing safety detection on the to-be-tested large model; wherein the jailbreak attack prompt words with positive context embedding are obtained by selecting positive answer content and malicious answer content, performing semantic mixed regulation on the rewritten prompt words in a structure-guided manner, and then performing reinforcement learning correction, comprising: selecting positive answer content and malicious answer content, selecting a mixed structure according to the type of the prompt word, and mixing to control the generation of a structure template control vector; determining a prompt word connection method and a prompt word connection template based on the structure template control vector and the mixed structure, modifying the rewritten prompt words based on the question corresponding to the positive answer content to obtain mixed prompt words; using a scenario-based rewriting model optimized by reinforcement learning to rewrite the mixed prompt words to obtain jailbreak attack prompt words with positive context embedding.
2. The method of claim 1, wherein, The jailbreak attack prompt words are classified, and the original jailbreak attack prompt words are rewritten based on the categories to obtain rewritten prompt words, comprising: converting the original jailbreak attack prompt words into synonymous prompt words; classifying the synonymous prompt words to obtain the categories of the synonymous prompt words, wherein the categories are knowledge categories or evaluation categories; rewriting the synonymous prompt words according to the categories of the synonymous prompt words to obtain rewritten prompt words.
3. The method of claim 2, wherein, The original jailbreak attack prompt words are converted into synonymous prompt words, specifically: using a synonymous conversion model to rewrite the original jailbreak attack prompt words into synonymous prompt words, wherein the semantic similarity between the original jailbreak attack prompt words and the synonymous prompt words is greater than a predetermined threshold, and the semantic similarity is obtained by calculating the similarity between the sentence vector of the original jailbreak attack prompt words and the sentence vector of the synonymous prompt words.
4. The method of claim 2, wherein, The categories of the synonymous prompt words include knowledge category prompt words and evaluation category prompt words; if the synonymous prompt word is a knowledge category prompt word, rewriting is performed by implanting a virtual context; if the synonymous prompt word is an evaluation category prompt word, a biased example is constructed, and the synonymous prompt word and the biased example are spliced as rewritten prompt words.
5. The method of claim 1, wherein, The reinforcement learning optimization process of the scenario-based rewriting model is realized through iteration, and each round of optimization process comprises: using the mixed prompt words as the semantic input of the scenario-based rewriting model, and the scenario-based rewriting model outputs mixed attack content; using the mixed attack content as the input of the discriminator to obtain the malicious probability determined by the discriminator for the mixed attack content; obtaining the interception flag of the guardrail for the mixed attack content; calculating the reinforcement learning reward based on the malicious probability and the interception flag, and updating the scenario-based rewriting model according to the reinforcement learning reward; updating the discriminator with the goal of minimizing the output score.
6. An apparatus for large model security detection based on forward context embedding jailbreaking attack hinting words, characterized in that, An acquisition module is configured to acquire original jailbreak attack prompt words; A rewriting module is configured to classify the original jailbreak attack prompt words and rewrite the original jailbreak attack prompt words based on the categories to obtain rewritten prompt words; A positive context embedding module is configured to perform structure-guided semantic hybrid regulation on the rewritten prompt words, and correct the structure-guided semantic hybrid regulation through reinforcement learning to obtain positive context-embedded jailbreak attack prompt words; A security monitoring module is configured to input the positive context-embedded jailbreak attack prompt words into a to-be-tested large model and perform security detection on the to-be-tested large model; The positive context-embedded jailbreak attack prompt words are obtained by selecting positive answer contents and malicious answer contents, performing structure-guided semantic hybrid regulation on the rewritten prompt words, and correcting the structure-guided semantic hybrid regulation through reinforcement learning, and the obtaining includes: The positive context-embedded jailbreak attack prompt words are obtained by selecting positive answer contents and malicious answer contents, selecting a mixed structure according to a prompt word type, and mixing and controlling a generation structure template control vector; A prompt word connection method and a prompt word connection template are determined based on the structure template control vector and the mixed structure, the rewritten prompt words are modified based on a question corresponding to the positive answer contents to obtain mixed prompt words; The mixed prompt words are rewritten by using a scenario-based rewriting model optimized through reinforcement learning to obtain the positive context-embedded jailbreak attack prompt words.
7. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the method of any one of claims 1-5.
8. An electronic device, comprising: The computer program / instruction is executed by the processor to implement the method of any one of claims 1-5. The computer program / instruction is executed by the processor to implement the method of any one of claims 1-5. The computer program / instruction is executed by the processor to implement the method of any one of claims 1-5. 9. A computer readable storage medium having stored thereon computer instructions, wherein,
Citation Information
Patent Citations
Prison break attack test method for multi-mode large model
CN119740229A