Security confrontation test method and device for large language model and electronic equipment
By obtaining the prompt text collection and adversarial strategy library to generate adversarial prompt text, and using the adversarial prompt generation model to perform security adversarial testing, the problem of high efficiency and low cost of artificial construction of adversarial prompt text in the prior art is solved, and efficient and accurate security adversarial testing is achieved.
Patent Information
- Application Number
- CN202510337146.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
AI Technical Summary
The security adversarial testing of large language models in the prior art requires manual construction of adversarial prompt text, which is costly and inefficient.
By obtaining the cue text collection, adversarial strategy library and adversarial cue generation model of the target large language model, the adversarial cue text is generated, and the adversarial cue generation model is used for security adversarial testing, combining the risk assessment model to improve testing efficiency and accuracy.
It reduces the generation cost of adversarial prompt text, improves generation efficiency and test accuracy, and enhances the processing accuracy of large language models in specific tasks.
Smart Images

Figure CN120297360A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, and in particular, to technologies such as deep learning, cloud computing, natural language processing, large models, etc. In particular, it relates to a security adversarial testing method, device, and electronic device for large language models. Background Art
[0002] Currently, when performing security adversarial testing on large language models, it is necessary for humans to construct adversarial prompt texts based on the original prompt texts of the large language models; and then perform security adversarial testing on the large language models in combination with the adversarial prompt texts. Among them, the generation cost of the adversarial prompt texts is high and the efficiency is poor. Summary of the Invention
[0003] The present disclosure provides a security adversarial testing method, device, and electronic device for large language models.
[0004] According to one aspect of the present disclosure, a security adversarial testing method for a large language model is provided. The method includes: obtaining a set of prompt texts corresponding to a target large language model, an adversarial strategy library, and an adversarial prompt generation model; generating adversarial prompt texts according to the original prompt texts in the set of prompt texts, the adversarial strategies in the adversarial strategy library, and the adversarial prompt generation model; and performing security adversarial testing on the target large language model according to the adversarial prompt texts.
[0005] According to another aspect of the present disclosure, a training method for an adversarial prompt generation model is provided. The method includes: obtaining first training data; the first training data includes: sample original prompt texts, sample strategy description texts corresponding to sample adversarial strategies, and sample adversarial prompt texts corresponding to the sample original prompt texts under the sample adversarial strategies; obtaining an initial adversarial prompt generation model; and training the adversarial prompt generation model according to the sample original prompt texts, the sample strategy description texts corresponding to the sample adversarial strategies, and the sample adversarial prompt texts corresponding to the sample original prompt texts under the sample adversarial strategies to obtain a trained adversarial prompt generation model for generating adversarial prompt texts when performing security adversarial testing on a target large language model.
[0006] According to another aspect of the present disclosure, a security adversarial testing device for a large language model is provided. The device includes: an obtaining module for obtaining a set of prompt texts corresponding to a target large language model, an adversarial strategy library, and an adversarial prompt generation model; a generating module for generating adversarial prompt texts according to the original prompt texts in the set of prompt texts, the adversarial strategies in the adversarial strategy library, and the adversarial prompt generation model; and a testing module for performing security adversarial testing on the target large language model according to the adversarial prompt texts.
[0007] According to another aspect of the present disclosure, there is provided a training device for an adversarial prompt generation model, the device comprising: a first acquisition module for acquiring first training data; the first training data includes: sample original prompt text, sample policy description text corresponding to a sample adversarial policy, and sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy; a second acquisition module for acquiring an initial adversarial prompt generation model; a training processing module for training the adversarial prompt generation model according to the sample original prompt text, the sample policy description text corresponding to the sample adversarial policy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy, to obtain a trained adversarial prompt generation model for generating adversarial prompt text when a target large language model performs a security adversarial test.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the security adversarial test method for a large language model proposed above in the present disclosure; or, execute the training method for an adversarial prompt generation model proposed above in the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, the computer instructions for causing a computer to execute the security adversarial test method for a large language model proposed above in the present disclosure; or, execute the training method for an adversarial prompt generation model proposed above in the present disclosure.
[0010] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, the computer program when executed by a processor implements the steps of the security adversarial test method for a large language model proposed above in the present disclosure; or, implements the steps of the training method for an adversarial prompt generation model proposed above in the present disclosure.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0013] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0014] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;
[0015] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;
[0016] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0017] Figure 5 is a schematic framework diagram for the training of an adversarial prompt generation model;
[0018] Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0019] Figure 7 is a schematic diagram according to the sixth embodiment of the present disclosure;
[0020] Figure 8 is a block diagram of an electronic device for implementing the security adversarial testing method of the large language model or the training method of the adversarial prompt generation model according to the embodiments of the present disclosure. Detailed implementation manners
[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0022] Currently, when performing security adversarial testing on a large language model, it is necessary for humans to construct adversarial prompt texts based on the original prompt texts of the large language model; and then perform security adversarial testing on the large language model in combination with the adversarial prompt texts. Among them, the generation cost of the adversarial prompt texts is high and the efficiency is poor.
[0023] In view of the above problems, the present disclosure proposes a security adversarial testing method, device and electronic device for a large language model.
[0024] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the security adversarial testing method of the large language model in the embodiments of the present disclosure can be applied to a security adversarial testing device for a large language model, and the device can be configured in an electronic device so that the electronic device can perform the security adversarial testing function of the large language model.
[0025] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC for short), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens, and / or display screens.
[0026] Among them, the security adversarial testing device of the large language model can also be software in the electronic device, such as the security adversarial testing software of the large language model, etc. In the following embodiments, the electronic device is taken as an example of the execution subject for illustration.
[0027] As Figure 1 shown, the security adversarial testing method of the large language model may include the following steps:
[0028] Step 101, obtain the prompt text set, adversarial strategy library, and adversarial prompt generation model corresponding to the target large language model.
[0029] In the embodiments of the present disclosure, the original prompt text in the prompt text set corresponding to the target large language model is the prompt text involved in the processing task of the target large language model. Among them, the scenarios to which the processing tasks of the target large language model belong include at least one of the following: intelligent customer service scenario, code generation scenario, medical diagnosis scenario.
[0030] Among them, the original prompt text, for example, the question text in the intelligent customer service scenario, the code requirement text in the code generation scenario, the symptom text in the medical diagnosis scenario, etc.
[0031] Among them, the setting of multiple scenarios enables the electronic device to generate adversarial prompt texts for security adversarial testing in multiple scenarios, thereby improving the security adversarial testing efficiency of the target large language model in multiple scenarios.
[0032] In the embodiments of the present disclosure, the adversarial strategy library may include at least one of the following strategies: perturbation strategy, attack strategy, and mutation strategy; the perturbation strategy is used to perform perturbation processing on at least one of the following levels: character level, word level, sentence level, and semantic level; the attack strategy includes at least one of the following: prompt injection strategy, prompt leakage strategy, and jailbreak attack strategy; the mutation strategy is used to perform mutation processing on at least one of the following: language structure, style, sentence pattern, language type.
[0033] In the embodiments of the present disclosure, the perturbation strategy at the character level is used to perform at least one of the following processes on the original prompt text: replacing at least one character in the original prompt text; adding redundant characters to the original prompt text; shuffling the characters in the original prompt text. Among them, the perturbation strategy at the word level is used to perform at least one of the following processes on the original prompt text: replacing at least one word in the original prompt text; adding redundant words to the original prompt text; shuffling the words in the original prompt text. Among them, the perturbation strategy at the sentence level is used to perform at least one of the following processes on the original prompt text: replacing at least one sentence in the original prompt text; adding redundant sentences to the original prompt text; shuffling the sentences in the original prompt text. Among them, the perturbation strategy at the semantic level, for example, restating the original prompt text without modifying its semantics, etc.
[0034] In the embodiments of the present disclosure, the prompt injection strategy may refer to injecting malicious instruction text, etc. into the original prompt text. The jailbreak attack strategy may refer to replacing the original prompt text with jailbreak instruction text, etc. The prompt leakage strategy may refer to injecting inducement instruction text, etc. into the original prompt text to induce the target large language model to leak sensitive information, etc.
[0035] In addition, the attack strategy may further include at least one of the following: role-playing strategy, negative induction strategy, reverse inhibition strategy, thought chain injection strategy, etc. Among them, the role-playing strategy may refer to guiding the target large language model to respond in a specific manner according to the set role information. The negative induction strategy may refer to giving negative or incorrect preconditions to induce the target large language model to generate inaccurate or harmful responses. The reverse inhibition strategy may refer to suppressing the normal functional performance of the target large language model by introducing concepts or instructions contrary to the expected output. The thought chain injection strategy may refer to guiding the target large language model to think along a designed path and draw a specific conclusion by constructing a series of misleading reasoning steps.
[0036] Among them, the setting of multiple confrontation strategies in the confrontation strategy library enables the electronic device to select confrontation strategies from the confrontation strategy library according to needs for generating confrontation prompt text, thereby improving the flexibility of generating confrontation prompt text and further improving the generation efficiency of confrontation prompt text.
[0037] Step 102, generate confrontation prompt text according to the original prompt text in the prompt text set, the confrontation strategy in the confrontation strategy library, and the confrontation prompt generation model.
[0038] In an embodiment of the present disclosure, the adversarial prompt generation model can be obtained by fine-tuning and training based on the original prompt text of the sample, the sample strategy description text corresponding to the sample adversarial strategy, and the sample adversarial prompt text corresponding to the original prompt text of the sample under the sample adversarial strategy.
[0039] Among them, the fine-tuning and training process of the adversarial prompt generation model can be, for example, inputting the original prompt text of the sample and the sample strategy description text into the adversarial prompt generation model to obtain the predicted adversarial prompt text output by the adversarial prompt generation model; determining the value of the first loss function according to the predicted adversarial prompt text, the sample adversarial prompt text, and the first loss function of the adversarial prompt generation model; and performing parameter adjustment processing on the adversarial prompt generation model according to the value of the first loss function to obtain the trained adversarial prompt generation model.
[0040] Among them, the adversarial prompt generation model is obtained by fine-tuning and training based on the original prompt text of the sample, the sample strategy description text corresponding to the sample adversarial strategy, and the sample adversarial prompt text corresponding to the original prompt text of the sample under the sample adversarial strategy, so that the adversarial prompt generation model can generate adversarial prompt text in combination with the adversarial strategy, thereby further improving the generation accuracy and generation efficiency of the adversarial prompt text.
[0041] Step 103, perform a security adversarial test on the target large language model according to the adversarial prompt text.
[0042] The security adversarial test method for the large language model in the embodiment of the present disclosure includes obtaining a prompt text set, an adversarial strategy library, and an adversarial prompt generation model corresponding to the target large language model; generating adversarial prompt text according to the original prompt text in the prompt text set, the adversarial strategy in the adversarial strategy library, and the adversarial prompt generation model; performing a security adversarial test on the target large language model according to the adversarial prompt text; among them, by combining the adversarial prompt generation model to generate adversarial prompt text and then performing a security adversarial test on the target large language model, the generation cost of the adversarial prompt text can be reduced, the generation efficiency of the adversarial prompt text can be improved, and further the security adversarial test efficiency of the target large language model can be improved; the adversarial prompt text generated by combining the adversarial prompt generation model is relatively comprehensive, so that the security adversarial test accuracy of the target large language model can be improved, and further the task processing accuracy of the target large language model when applied to a specific task can be improved.
[0043] Among them, in order to improve the richness of the generated adversarial prompt text and further improve the security adversarial test efficiency of the target large language model, the electronic device can obtain the first original prompt text from the prompt text set and the first adversarial strategy from the adversarial strategy library, and then generate the adversarial prompt text. As Figure 2 shown, Figure 2 is a schematic diagram according to the second embodiment of the present disclosure.Figure 2 The illustrated embodiment may include the following steps:
[0044] Step 201: Obtain a set of prompt texts corresponding to the target large language model, an adversarial strategy library, and an adversarial prompt generation model.
[0045] Step 202: Obtain a first original prompt text from the set of prompt texts.
[0046] In the embodiments of the present disclosure, the electronic device may randomly extract the first original prompt text from the set of prompt texts; alternatively, the electronic device may sequentially extract the first original prompt text from the set of prompt texts.
[0047] It should be noted that, based on the same original prompt text and the same adversarial strategy, multiple text generation processes may result in the same adversarial prompt text or different adversarial prompt texts. Therefore, in order to expand the number of adversarial prompt texts, the original prompt texts in the set of prompt texts may be repeatedly extracted. That is to say, the same original prompt text may be extracted multiple times.
[0048] Step 203: Obtain a first adversarial strategy from the adversarial strategy library.
[0049] In the embodiments of the present disclosure, the process of the electronic device executing Step 203 may be, for example, to obtain an adversarial strategy selection condition; determine a second adversarial strategy from the adversarial strategy library according to the adversarial strategy selection condition; and determine the first adversarial strategy according to the second adversarial strategy.
[0050] In an example of the embodiments of the present disclosure, the adversarial strategy selection condition may include information such as the identifier of the adversarial strategy to be selected. The electronic device may perform an adversarial strategy selection process according to the identifier in the adversarial strategy selection condition.
[0051] In another example, the adversarial strategy selection condition may include the identifier of the adversarial strategy to be selected, etc., and the selection probability of the adversarial strategy to be selected. Correspondingly, for each obtained first original prompt text, the electronic device may select an adversarial strategy from each adversarial strategy according to the selection probability of each adversarial strategy for use in the generation process of the adversarial prompt text in combination with the first original prompt text.
[0052] The setting of the adversarial strategy selection condition enables the electronic device to determine or change the adversarial strategy selection condition through interaction with the object, and then perform the generation process of the adversarial prompt text, thereby improving the flexibility of the adversarial strategy selection and further improving the richness of the generated adversarial prompt text.
[0053] In an embodiment of the present disclosure, in one example, the process by which the electronic device determines the first adversarial strategy according to the second adversarial strategy may be, for example, extracting a strategy from the second adversarial strategy as the first adversarial strategy. In another example, the process by which the electronic device determines the first adversarial strategy according to the second adversarial strategy may be, for example, directly determining the second adversarial strategy as the first adversarial strategy.
[0054] Step 204: Input the first original prompt text and the strategy description text corresponding to the first adversarial strategy into the adversarial prompt generation model, and obtain the adversarial prompt text output by the adversarial prompt generation model.
[0055] In an embodiment of the present disclosure, in one example, the number of the first adversarial strategies may be one. For example, the first adversarial strategy may be one of a perturbation strategy, an attack strategy, and a mutation strategy. Correspondingly, the electronic device may directly input the first original prompt text and the strategy description text corresponding to the first adversarial strategy into the adversarial prompt generation model, and obtain the adversarial prompt text output by the adversarial prompt generation model.
[0056] In another example, the number of the first adversarial strategies may be multiple, including a perturbation strategy and / or an attack strategy. Correspondingly, for each first adversarial strategy, the electronic device may input the first original prompt text and the strategy description text corresponding to the first adversarial strategy into the adversarial prompt generation model, and obtain the adversarial prompt text.
[0057] In another example, the number of the first adversarial strategies may be multiple, including any one of a perturbation strategy and an attack strategy, and a mutation strategy. For example, the first adversarial strategy may include a perturbation strategy and a mutation strategy; or, the first adversarial strategy may include an attack strategy and a mutation strategy. Correspondingly, the process for the electronic device to execute step 204 may be, for example, obtaining the first strategy description text corresponding to any one of the strategies, and the second strategy description text corresponding to the mutation strategy in the first adversarial strategy; inputting the first original prompt text and the first strategy description text into the adversarial prompt generation model, and obtaining the intermediate prompt text output by the adversarial prompt generation model; inputting the intermediate prompt text and the second strategy description text into the adversarial prompt generation model, and obtaining the adversarial prompt text output by the adversarial prompt generation model.
[0058] Among them, combining the mutation strategy, and adding one of the perturbation strategy and the attack strategy to process the original prompt text can further expand the richness of the adversarial prompt text, thereby further improving the security adversarial test efficiency of the target large language model.
[0059] Step 205: Perform a security adversarial test on the target large language model according to the adversarial prompt text.
[0060] Among them, it should be noted that for the detailed content of steps 202 to 204, reference can be made to Figure 1 step 102 in the illustrated embodiment, and no further detailed description will be given here.
[0061] The security confrontation test method for the large language model of the present disclosure embodiment includes obtaining a set of prompt texts, an adversarial strategy library, and an adversarial prompt generation model corresponding to the target large language model; obtaining a first original prompt text from the set of prompt texts; obtaining a first adversarial strategy from the adversarial strategy library; inputting the first original prompt text and the strategy description text corresponding to the first adversarial strategy into the adversarial prompt generation model to obtain the adversarial prompt text output by the adversarial prompt generation model; performing a security confrontation test on the target large language model according to the adversarial prompt text; among them, obtaining the first original prompt text from the set of prompt texts and obtaining the first adversarial strategy from the adversarial strategy library, and then generating the adversarial prompt text can improve the richness of the generated adversarial prompt text, and further improve the efficiency of the security confrontation test of the target large language model; combining the adversarial prompt text generated by the adversarial prompt generation model is relatively comprehensive, so as to improve the accuracy of the security confrontation test of the target large language model, and further improve the task processing accuracy when the target large language model is applied to specific tasks.
[0062] Among them, in order to obtain the security confrontation test result of the target large language model and improve the accuracy of the security confrontation test result, the electronic device can input the adversarial prompt text into the target large language model, and determine the security confrontation test result in combination with the predicted answer text output by the target large language model and the risk assessment large model. As Figure 3 shown, Figure 3 is a schematic diagram according to the third embodiment of the present disclosure, Figure 3 The illustrated embodiment may include the following steps:
[0063] Step 301, obtain a set of prompt texts, an adversarial strategy library, and an adversarial prompt generation model corresponding to the target large language model.
[0064] Step 302, generate an adversarial prompt text according to the original prompt text in the set of prompt texts, the adversarial strategy in the adversarial strategy library, and the adversarial prompt generation model.
[0065] Step 303, input the adversarial prompt text into the target large language model to obtain the predicted answer text output by the target large language model.
[0066] Step 304, perform risk assessment processing on the predicted answer text according to the risk assessment large model to obtain a risk assessment result.
[0067] In an embodiment of the present disclosure, the process of the electronic device executing step 304 may be, for example, inputting the predicted answer text into the risk assessment large model to obtain the risk score value and risk reason output by the risk assessment large model; and determining the risk assessment result corresponding to the predicted answer text based on the risk score value and the risk reason.
[0068] Among them, the risk reason may be, for example, a countermeasure; or, it may be the expected effect corresponding to the countermeasure. Among them, the expected effect is, for example, outputting harmful content, leaking sensitive data, etc. The risk assessment large model may be trained by combining the sample text, the sample risk score value of the sample text, and the risk reason.
[0069] Among them, the risk score value can indicate the risk degree of the predicted answer text. The lower the risk score value, the safer the predicted answer text, that is, the lower the possibility of outputting harmful content or leaking sensitive data, etc. The higher the risk score value, the more dangerous the predicted answer text, that is, the higher the possibility of outputting harmful content or leaking sensitive data, etc.
[0070] Among them, determining the risk score value and risk reason of the predicted answer text by combining the risk assessment large model can improve the accuracy of the determined risk assessment result.
[0071] Step 305, determining the security adversarial test result of the target large language model according to the risk assessment result.
[0072] In an embodiment of the present disclosure, the process of the electronic device executing step 305 may be, for example, obtaining the first risk assessment result among each risk assessment result; the risk score value in the first risk assessment result is greater than or equal to the risk score threshold; and determining the security adversarial test result of the target large language model according to the risk reason in the first risk assessment result.
[0073] In an embodiment of the present disclosure, the process of the electronic device determining the security adversarial test result by combining the risk reason may be, for example, obtaining the third countermeasure indicated by the risk reason in the first risk assessment result; for each third countermeasure, determining the risk degree of the third countermeasure according to the risk score value in the first risk assessment result to which the risk reason indicating the third countermeasure belongs; and determining the security adversarial test result of the target large language model according to the risk degrees of each third countermeasure.
[0074] Among them, in the security confrontation test results determined according to the risk reasons in the first risk assessment result, it can be seen which confrontation prompt texts the target large language model has weak confrontation capabilities against under which confrontation strategies. Furthermore, the target large language model can be specifically optimized for confrontation to improve the confrontation capabilities of the target large language model against the confrontation prompt texts under each confrontation strategy. Thus, when the target large language model is applied to a specific task, the probability of generating risky answer texts under this specific task can be reduced. Among them, the determination of the risk level enables the target large language model to perform different degrees of confrontation optimization processing for different confrontation strategies, thereby improving the optimization processing efficiency of the target large language model.
[0075] Among them, it should be noted that for the detailed content of steps 303 to 305, reference can be made to Figure 1 step 103 in the illustrated embodiment, and no further detailed description will be provided here.
[0076] The security confrontation test method for the large language model of the present disclosure embodiment includes obtaining a set of prompt texts, a confrontation strategy library, and a confrontation prompt generation model corresponding to the target large language model; generating confrontation prompt texts according to the original prompt texts in the set of prompt texts, the confrontation strategies in the confrontation strategy library, and the confrontation prompt generation model; inputting the confrontation prompt texts into the target large language model to obtain the predicted answer texts output by the target large language model; performing risk assessment processing on the predicted answer texts according to the risk assessment large model to obtain a risk assessment result; and determining the security confrontation test result of the target large language model according to the risk assessment result. Among them, combining the predicted answer texts output by the target large language model and the risk assessment large model to determine the security confrontation test result can further improve the accuracy of the security confrontation test result.
[0077] Figure 4 It is a schematic diagram according to the fourth embodiment of the present disclosure. It should be noted that the training method of the confrontation prompt generation model in the present disclosure embodiment can be applied to a training device of the confrontation prompt generation model, and this device can be configured in an electronic device so that the electronic device can perform the training function of the confrontation prompt generation model.
[0078] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC for short), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens, and / or display screens.
[0079] Among them, the training device for the adversarial prompt generation model can also be software in an electronic device, such as the training software for the adversarial prompt generation model, etc. In the following embodiments, the electronic device is taken as the execution subject for illustration.
[0080] As Figure 4 shown, the training method for the adversarial prompt generation model may include the following steps:
[0081] Step 401, obtain the first training data; the first training data includes: the sample original prompt text, the sample strategy description text corresponding to the sample adversarial strategy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial strategy.
[0082] In the embodiments of the present disclosure, the sample adversarial strategy may include at least one of the following strategies: perturbation strategy, attack strategy, and mutation strategy; the perturbation strategy is used to perform perturbation processing on at least one of the following levels: character level, word level, sentence level, and semantic level; the attack strategy includes at least one of the following: prompt injection strategy, prompt leakage strategy, and jailbreak attack strategy; the mutation strategy is used to perform mutation processing on at least one of the following: language structure, style, sentence pattern, and language type.
[0083] Step 402, obtain an initial adversarial prompt generation model.
[0084] In the embodiments of the present disclosure, in order to further improve the training accuracy of the adversarial prompt generation model, after step 402, the electronic device may perform pre-training processing on the adversarial prompt generation model according to the knowledge text, so that the adversarial prompt generation model can well understand the knowledge text. Correspondingly, the electronic device may perform the following process: obtain the knowledge text; the knowledge text may include security knowledge text and / or adversarial knowledge text; perform pre-training processing on the adversarial prompt generation model according to the security knowledge text and / or adversarial knowledge text in the knowledge text.
[0085] Among them, in one example, the knowledge text may include security knowledge text. The security knowledge text is determined based on relevant security guidelines and is the knowledge text related to security stipulated in the relevant security guidelines. Correspondingly, the electronic device may perform pre-training processing on the adversarial prompt generation model according to the security knowledge text.
[0086] In another example, the knowledge text may include adversarial knowledge text. The adversarial knowledge text is determined based on relevant security guidelines and is the knowledge text related to risks stipulated in the relevant security guidelines. Correspondingly, the electronic device may perform pre-training processing on the adversarial prompt generation model according to the adversarial knowledge text.
[0087] In another example, the knowledge text may include security knowledge text and adversarial knowledge text. The electronic device may perform training processing on the adversarial prompt generation model according to the security knowledge text and the adversarial knowledge text.
[0088] In an embodiment of the present disclosure, the process of the electronic device performing pre-training processing on the adversarial prompt generation model according to the security knowledge text and / or the adversarial knowledge text in the knowledge text may be, for example, obtaining a first text segment and a second text segment in the knowledge text; inputting the first text segment into the adversarial prompt generation model to obtain a predicted text segment output by the adversarial prompt generation model; and performing parameter adjustment processing on the adversarial prompt generation model according to the second text segment and the predicted text segment to obtain the pre-trained adversarial prompt generation model.
[0089] The first text segment may be at least one discontinuous character in the knowledge text, or may be a segment composed of at least two consecutive characters in the knowledge text. Among them, parameter adjustment processing is performed on the adversarial prompt generation model by combining the first text segment and the second text segment in the knowledge text, so that the adversarial prompt generation model can learn the semantic knowledge in the knowledge text.
[0090] Step 403: Perform training processing on the adversarial prompt generation model according to the sample original prompt text, the sample policy description text corresponding to the sample adversarial policy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy, to obtain the trained adversarial prompt generation model, which is used to generate the adversarial prompt text when the target large language model performs security adversarial testing.
[0091] In an embodiment of the present disclosure, the process of the electronic device executing step 403 may be, for example, inputting the sample original prompt text and the sample policy description text into the adversarial prompt generation model to obtain a predicted adversarial prompt text output by the adversarial prompt generation model; determining the value of the first loss function according to the predicted adversarial prompt text, the sample adversarial prompt text, and the first loss function of the adversarial prompt generation model; and performing parameter adjustment processing on the adversarial prompt generation model according to the value of the first loss function to obtain the trained adversarial prompt generation model.
[0092] The first loss function may specifically be the reciprocal of the vector similarity between the text vector of the predicted adversarial prompt text and the text vector of the sample adversarial prompt text. The more similar the text vector of the predicted adversarial prompt text is to the text vector of the sample adversarial prompt text, the smaller the value of the first loss function; the greater the difference between the text vector of the predicted adversarial prompt text and the text vector of the sample adversarial prompt text, the greater the value of the first loss function.
[0093] Among them, according to the predicted adversarial prompt text, the sample adversarial prompt text, and the first loss function of the adversarial prompt generation model, the value of the first loss function is determined, and then parameter adjustment processing is performed on the adversarial prompt generation model, which can make the difference between the predicted adversarial prompt text output by the adversarial prompt generation model and the sample adversarial prompt text smaller and smaller, thereby improving the accuracy of the trained adversarial prompt generation model.
[0094] In the embodiments of the present disclosure, in order to further improve the accuracy of the trained adversarial prompt generation model, when the number of predicted adversarial prompt texts corresponding to the sample original prompt text is at least two, the reward scores of at least two predicted adversarial prompt texts are determined according to the reward model, and then the adversarial prompt generation model is retrained according to the reward scores of the two predicted adversarial prompt texts. Correspondingly, the electronic device can also perform the following process: determining the reward scores for at least two predicted adversarial prompt texts according to at least two predicted adversarial prompt texts and the reward model; determining the value of the second loss function according to at least two predicted adversarial prompt texts, the reward scores, and the second loss function; and performing parameter adjustment processing on the adversarial prompt generation model according to the value of the second loss function to obtain the trained adversarial prompt generation model. Among them, the second loss function encourages the adversarial prompt generation model to generate adversarial prompt texts that can obtain higher reward scores.
[0095] The training method of the adversarial prompt generation model in the embodiments of the present disclosure includes obtaining first training data; the first training data includes: a sample original prompt text, a sample strategy description text corresponding to a sample adversarial strategy, and a sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial strategy; obtaining an initial adversarial prompt generation model; and training the adversarial prompt generation model according to the sample original prompt text, the sample strategy description text corresponding to the sample adversarial strategy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial strategy to obtain a trained adversarial prompt generation model for generating adversarial prompt texts during the security adversarial test of the target large language model; among them, the adversarial prompt generation model is fine-tuned and trained by combining the sample original prompt text, the sample strategy description text corresponding to the sample adversarial strategy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial strategy to ensure the accuracy of the trained adversarial prompt generation model.
[0096] The following is an example for illustration. As Figure 5 shown, it is a schematic framework diagram of the training of the adversarial prompt generation model. In Figure 5It includes the following steps. Step 501, obtain a base large model (i.e., the initial adversarial prompt generation model). Step 502, perform incremental pre-training processing on the base large model in combination with the security knowledge corpus (i.e., perform pre-training processing on the adversarial prompt generation model according to the knowledge text). Step 503, determine Q&A pairs Q&APair (the first training data) in combination with the adversarial generation instruction set, mutation rewriting instruction set, etc., and then perform fine-tuning training processing on the adversarial prompt generation model. Step 504, perform training processing on the adversarial prompt generation model in combination with the preference annotation data obtained through instruction output sampling and preference annotation (the replacement solution can be the reward score of the predicted adversarial prompt text output by the adversarial prompt generation model), so as to obtain the trained adversarial prompt generation model (i.e., Figure 5 the attack large model in).
[0097] To implement the above embodiments, the present disclosure also provides a security adversarial test device for a large language model. As Figure 6 shown, Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure. The security adversarial test device 60 for the large language model may include: an acquisition module 601, a generation module 602, and a test module 603.
[0098] Among them, the acquisition module 601 is used to acquire a set of prompt texts, an adversarial strategy library, and an adversarial prompt generation model corresponding to the target large language model; the generation module 602 is used to generate adversarial prompt texts according to the original prompt texts in the set of prompt texts, the adversarial strategies in the adversarial strategy library, and the adversarial prompt generation model; the test module 603 is used to perform a security adversarial test on the target large language model according to the adversarial prompt texts.
[0099] As a possible implementation manner of the embodiment of the present disclosure, the generation module 602 includes: a first acquisition unit, a second acquisition unit, and a third acquisition unit; the first acquisition unit is used to acquire a first original prompt text from the set of prompt texts; the second acquisition unit is used to acquire a first adversarial strategy from the adversarial strategy library; the third acquisition unit is used to input the first original prompt text and the strategy description text corresponding to the first adversarial strategy into the adversarial prompt generation model, and acquire the adversarial prompt text output by the adversarial prompt generation model.
[0100] As a possible implementation manner of the embodiment of the present disclosure, the second acquisition unit is specifically used to acquire an adversarial strategy selection condition; determine a second adversarial strategy from the adversarial strategy library according to the adversarial strategy selection condition; and determine the first adversarial strategy according to the second adversarial strategy.
[0101] As a possible implementation manner of the embodiments of the present disclosure, the adversarial strategies in the adversarial strategy library include at least one of the following strategies: perturbation strategy, attack strategy, and mutation strategy; the perturbation strategy is used to perform perturbation processing on at least one of the following levels: character level, word level, sentence level, and semantic level; the attack strategy includes at least one of the following: prompt injection strategy, prompt leakage strategy, and jailbreak attack strategy; the mutation strategy is used to perform mutation processing on at least one of the following: language structure, style, sentence pattern, and language type.
[0102] As a possible implementation manner of the embodiments of the present disclosure, the first adversarial strategy includes any one of the perturbation strategy and the attack strategy, and the mutation strategy; the third acquisition unit is specifically configured to acquire the first strategy description text corresponding to the any one of the strategies, and the second strategy description text corresponding to the mutation strategy in the first adversarial strategy; input the first original prompt text and the first strategy description text into the adversarial prompt generation model to obtain the intermediate prompt text output by the adversarial prompt generation model; input the intermediate prompt text and the second strategy description text into the adversarial prompt generation model to obtain the adversarial prompt text output by the adversarial prompt generation model.
[0103] As a possible implementation manner of the embodiments of the present disclosure, the test module 603 includes: a fourth acquisition unit, a risk assessment unit, and a determination unit; the fourth acquisition unit is configured to input the adversarial prompt text into the target large language model to obtain the predicted answer text output by the target large language model; the risk assessment unit is configured to perform risk assessment processing on the predicted answer text according to the risk assessment large model to obtain a risk assessment result; the determination unit is configured to determine the security adversarial test result of the target large language model according to the risk assessment result.
[0104] As a possible implementation manner of the embodiments of the present disclosure, the risk assessment unit is specifically configured to input the predicted answer text into the risk assessment large model to obtain the risk score value and the risk reason output by the risk assessment large model; determine the risk assessment result corresponding to the predicted answer text with the risk score value and the risk reason.
[0105] As a possible implementation manner of the embodiments of the present disclosure, the determination unit is specifically configured to acquire the first risk assessment result among the various risk assessment results; the risk score value in the first risk assessment result is greater than or equal to the risk score threshold; determine the security adversarial test result of the target large language model according to the risk reason in the first risk assessment result.
[0106] As a possible implementation manner of the embodiments of the present disclosure, the determining unit is further specifically configured to obtain a third countermeasure indicated by a risk reason in the first risk assessment result; for each third countermeasure, determine the risk level of the third countermeasure according to the risk score value in the first risk assessment result to which the risk reason indicating the third countermeasure belongs; and determine the security countermeasure test result of the target large language model according to the risk levels of the respective third countermeasures.
[0107] As a possible implementation manner of the embodiments of the present disclosure, the scenarios to which the processing tasks of the target large language model belong include at least one of the following: intelligent customer service scenario, code generation scenario, and medical diagnosis scenario.
[0108] The security countermeasure test device for the large language model according to the embodiments of the present disclosure obtains a set of prompt texts, a countermeasure library, and a countermeasure prompt generation model corresponding to the target large language model; generates countermeasure prompt texts according to the original prompt texts in the set of prompt texts, the countermeasures in the countermeasure library, and the countermeasure prompt generation model; and performs a security countermeasure test on the target large language model according to the countermeasure prompt texts. Among them, by combining the countermeasure prompt generation model to generate countermeasure prompt texts and then performing a security countermeasure test on the target large language model, the generation cost of the countermeasure prompt texts can be reduced, the generation efficiency of the countermeasure prompt texts can be improved, and further the security countermeasure test efficiency of the target large language model can be improved; the countermeasure prompt texts generated by combining the countermeasure prompt generation model are relatively comprehensive, so that the security countermeasure test accuracy of the target large language model can be improved, and further the task processing accuracy of the target large language model when applied to specific tasks can be improved.
[0109] To implement the above embodiments, the present disclosure also provides a training device for a countermeasure prompt generation model. As Figure 7 shown, Figure 7 is a schematic diagram according to the sixth embodiment of the present disclosure. The training device 70 for the countermeasure prompt generation model may include: a first acquisition module 701, a second acquisition module 702, and a training processing module 703.
[0110] Among them, the first acquisition module 701 is used to acquire first training data; the first training data includes: a sample original prompt text, a sample strategy description text corresponding to a sample adversarial strategy, and a sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial strategy; the second acquisition module 702 is used to acquire an initial adversarial prompt generation model; the training processing module 703 is used to perform training processing on the adversarial prompt generation model according to the sample original prompt text, the sample strategy description text corresponding to the sample adversarial strategy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial strategy, so as to obtain a trained adversarial prompt generation model for generating adversarial prompt texts during the security adversarial test of the target large language model.
[0111] As a possible implementation manner of the embodiments of the present disclosure, the device further includes: a third acquisition module and a pre-training module; the third acquisition module is used to acquire knowledge texts; the knowledge texts include security knowledge texts and / or adversarial knowledge texts; the pre-training module is used to perform pre-training processing on the adversarial prompt generation model according to the security knowledge texts and / or the adversarial knowledge texts in the knowledge texts.
[0112] As a possible implementation manner of the embodiments of the present disclosure, the pre-training module is specifically configured to acquire a first text segment and a second text segment in the knowledge text; input the first text segment into the adversarial prompt generation model to obtain a predicted text segment output by the adversarial prompt generation model; perform parameter adjustment processing on the adversarial prompt generation model according to the second text segment and the predicted text segment to obtain a pre-trained adversarial prompt generation model.
[0113] As a possible implementation manner of the embodiments of the present disclosure, the training processing module 703 is specifically configured to input the sample original prompt text and the sample strategy description text into the adversarial prompt generation model to obtain a predicted adversarial prompt text output by the adversarial prompt generation model; determine the value of the first loss function according to the predicted adversarial prompt text, the sample adversarial prompt text, and the first loss function of the adversarial prompt generation model; perform parameter adjustment processing on the adversarial prompt generation model according to the value of the first loss function to obtain a trained adversarial prompt generation model.
[0114] As a possible implementation manner of the embodiments of the present disclosure, the number of predicted adversarial prompt texts corresponding to the sample original prompt text is at least two; the apparatus further includes: a first determination module, a second determination module, and a parameter adjustment module; the first determination module is configured to determine reward scores for at least two of the predicted adversarial prompt texts according to at least two of the predicted adversarial prompt texts and a reward model; the second determination module is configured to determine a value of the second loss function according to at least two of the predicted adversarial prompt texts, the reward scores, and the second loss function; the parameter adjustment module is configured to perform parameter adjustment processing on the adversarial prompt generation model according to the value of the second loss function to obtain a trained adversarial prompt generation model.
[0115] The training apparatus for the adversarial prompt generation model of the embodiments of the present disclosure obtains first training data; the first training data includes: a sample original prompt text, a sample policy description text corresponding to a sample adversarial policy, and a sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy; obtains an initial adversarial prompt generation model; and performs training processing on the adversarial prompt generation model according to the sample original prompt text, the sample policy description text corresponding to the sample adversarial policy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy to obtain a trained adversarial prompt generation model for generating adversarial prompt texts when a target large language model performs a security adversarial test; wherein, the adversarial prompt generation model is fine-tuned and trained in combination with the sample original prompt text, the sample policy description text corresponding to the sample adversarial policy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy to ensure the accuracy of the trained adversarial prompt generation model.
[0116] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing are all carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0117] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0118] Figure 8FIG. 0 shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0119] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0120] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as, for example, a keyboard, a mouse, etc.; an output unit 807, such as, for example, various types of displays, speakers, etc.; a storage unit 808, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 809, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0121] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the security adversarial testing method of a large language model or the training method of an adversarial prompt generation model. For example, in some embodiments, the security adversarial testing method of a large language model or the training method of an adversarial prompt generation model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the security adversarial testing method of the large language model or the training method of the adversarial prompt generation model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the security adversarial testing method of the large language model or the training method of the adversarial prompt generation model in any other suitable way (e.g., by means of firmware).
[0122] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0124] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronics, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0125] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic, speech, or tactile input).
[0126] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0127] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0128] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0129] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A security adversarial testing method for a large language model, the method comprising: Obtaining a set of prompt texts, an adversarial strategy library, and an adversarial prompt generation model corresponding to the target large language model; Generating adversarial prompt texts based on the original prompt texts in the set of prompt texts, the adversarial strategies in the adversarial strategy library, and the adversarial prompt generation model; Performing a security adversarial test on the target large language model according to the adversarial prompt texts.
2. The method according to claim 1, wherein, The generating of the adversarial prompt texts based on the original prompt texts in the set of prompt texts, the adversarial strategies in the adversarial strategy library, and the adversarial prompt generation model comprises: Obtaining a first original prompt text from the set of prompt texts; Obtaining a first adversarial strategy from the adversarial strategy library; Inputting the first original prompt text and the strategy description text corresponding to the first adversarial strategy into the adversarial prompt generation model to obtain the adversarial prompt texts output by the adversarial prompt generation model.
3. The method according to claim 2, wherein The obtaining of the first adversarial strategy from the adversarial strategy library comprises: Obtaining adversarial strategy selection conditions; Determining a second adversarial strategy from the adversarial strategy library according to the adversarial strategy selection conditions; Determining the first adversarial strategy according to the second adversarial strategy.
4. The method according to claim 1 or 2, wherein The adversarial strategy library includes at least one of the following strategies: perturbation strategy, attack strategy, and mutation strategy; The perturbation strategy is used to perform perturbation processing on at least one of the following levels: character level, word level, sentence level, and semantic level; The attack strategy includes at least one of the following: prompt injection strategy, prompt leakage strategy, and jailbreak attack strategy; The mutation strategy is used to perform mutation processing on at least one of the following: language structure, style, sentence pattern, language type.
5. The method according to claim 2, wherein The first adversarial strategy includes: any one of the perturbation strategy and the attack strategy, and the mutation strategy; The inputting of the first original prompt text and the strategy description text corresponding to the first adversarial strategy into the adversarial prompt generation model to obtain the adversarial prompt texts output by the adversarial prompt generation model comprises: Obtaining a first strategy description text corresponding to the any one of the strategies, and a second strategy description text corresponding to the mutation strategy; Inputting the first original prompt text and the first strategy description text into the adversarial prompt generation model to obtain intermediate prompt texts output by the adversarial prompt generation model; Inputting the intermediate prompt texts and the second strategy description text into the adversarial prompt generation model to obtain the adversarial prompt texts output by the adversarial prompt generation model.
6. The method according to claim 1, wherein The performing of the security adversarial test on the target large language model according to the adversarial prompt texts comprises: Inputting the adversarial prompt texts into the target large language model to obtain answer texts output by the target large language model; Performing a risk assessment process on the answer texts according to a risk assessment large model to obtain a risk assessment result; Determining a security adversarial test result of the target large language model according to the risk assessment result.
7. The method according to claim 6, wherein, The performing of the risk assessment process on the answer texts according to a risk assessment large model to obtain a risk assessment result comprises: Input the answer text into the risk assessment large model to obtain the risk score value and risk reason output by the risk assessment large model; Determine the risk assessment result corresponding to the answer text based on the risk score value and the risk reason.
8. The method according to claim 6, wherein Determine the security confrontation test result of the target large language model according to the risk assessment result, including: Obtain the first risk assessment result among all the risk assessment results; the risk score value in the first risk assessment result is greater than or equal to the risk score threshold; Determine the security confrontation test result of the target large language model according to the risk reason in the first risk assessment result.
9. The method according to claim 8, wherein Determine the security confrontation test result of the target large language model according to the risk reason in the first risk assessment result, including: Obtain the third confrontation strategy indicated by the risk reason in the first risk assessment result; For each third confrontation strategy, determine the risk level of the third confrontation strategy according to the risk score value in the first risk assessment result to which the risk reason indicating the third confrontation strategy belongs; Determine the security confrontation test result of the target large language model according to the risk levels of all the third confrontation strategies.
10. The method according to claim 1, wherein, The scenarios to which the processing tasks of the target large language model belong include at least one of the following: intelligent customer service scenario, code generation scenario, medical diagnosis scenario.
11. A training method for an adversarial prompt generation model, the method comprising: Obtain first training data; The first training data includes: sample original prompt text, sample strategy description text corresponding to the sample confrontation strategy, and sample adversarial prompt text corresponding to the sample original prompt text under the sample confrontation strategy; Obtain an initial adversarial prompt generation model; Train the adversarial prompt generation model according to the sample original prompt text, the sample strategy description text corresponding to the sample confrontation strategy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample confrontation strategy to obtain a trained adversarial prompt generation model for generating adversarial prompt text when the target large language model performs a security confrontation test.
12. The method according to claim 11, wherein, The method further includes: Obtain knowledge text; the knowledge text includes security knowledge text and / or adversarial knowledge text; Perform pre-training processing on the adversarial prompt generation model according to the security knowledge text and / or the adversarial knowledge text in the knowledge text.
13. The method according to claim 12, wherein, Performing pre-training processing on the adversarial prompt generation model according to the security knowledge text and / or the adversarial knowledge text in the knowledge text includes: Obtain a first text segment and a second text segment in the knowledge text; Input the first text segment into the adversarial prompt generation model to obtain a predicted text segment output by the adversarial prompt generation model; Perform parameter adjustment processing on the adversarial prompt generation model according to the second text segment and the predicted text segment to obtain a pre-trained adversarial prompt generation model.
14. The method according to claim 11, wherein, Training the adversarial prompt generation model based on the sample original prompt text, the sample policy description text corresponding to the sample adversarial policy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy to obtain a trained adversarial prompt generation model, including: Inputting the sample original prompt text and the sample policy description text into the adversarial prompt generation model to obtain a predicted adversarial prompt text output by the adversarial prompt generation model; Determining the value of the first loss function according to the predicted adversarial prompt text, the sample adversarial prompt text, and the first loss function of the adversarial prompt generation model; Performing parameter adjustment processing on the adversarial prompt generation model according to the value of the first loss function to obtain a trained adversarial prompt generation model.
15. The method according to claim 14, wherein The number of predicted adversarial prompt texts corresponding to the sample original prompt text is at least two; the method further includes: Determining reward scores for at least two predicted adversarial prompt texts according to the at least two predicted adversarial prompt texts and a reward model; Determining the value of the second loss function according to the at least two predicted adversarial prompt texts, the reward scores, and the second loss function; Performing parameter adjustment processing on the adversarial prompt generation model according to the value of the second loss function to obtain a trained adversarial prompt generation model.
16. A security adversarial test device for a large language model, the device includes: An acquisition module, configured to acquire a prompt text set, an adversarial policy library, and an adversarial prompt generation model corresponding to a target large language model; A generation module, configured to generate adversarial prompt texts according to the original prompt texts in the prompt text set, the adversarial policies in the adversarial policy library, and the adversarial prompt generation model; A test module, configured to perform a security adversarial test on the target large language model according to the adversarial prompt texts.
17. A training device for an adversarial prompt generation model, the device includes: A first acquisition module, configured to acquire first training data; The first training data includes: a sample original prompt text, a sample policy description text corresponding to a sample adversarial policy, and a sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy; A second acquisition module, configured to acquire an initial adversarial prompt generation model; A training processing module, configured to train the adversarial prompt generation model according to the sample original prompt text, the sample policy description text corresponding to the sample adversarial policy, and the sample adversarial prompt text corresponding to the sample original prompt text under the sample adversarial policy to obtain a trained adversarial prompt generation model for generating adversarial prompt texts when performing a security adversarial test on a target large language model.
18. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 10; or, execute the method according to any one of claims 11 to 15.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10; or, execute the method according to any one of claims 11 to 15.
20. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10; or, implements the method according to any one of claims 11 to 15.
Citation Information
Cited By
Large model output data security detection method and system based on adversarial attack
CN121309093A