Reverse-generated dialogue model attack method, system and storage medium

CN115827839BActive Publication Date: 2026-10-09TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211493802.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2026-10-09
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

[0004]本发明提供一种基于反向生成的对话模型攻击方法、系统及存储介质,用以解决现有预训练语言模型生成文本安全检测成本高、无法对上文诱导性进行控制的问题

Benefits of technology

[0033] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the reverse generation-based dialogue model attack method as described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115827839B_ABST
    Figure CN115827839B_ABST
Patent Text Reader

Abstract

The application provides a dialogue model attack method and system based on reverse generation, and a storage medium, comprising: measuring initial preceding text data through a preset classifier to determine the relationship between preceding text toxicity, preceding text category and preceding text inducibility; establishing a reverse language generation model, training the reverse language generation model using a loss function, generating preceding text based on a given reply through the trained reverse language generation model; controlling the preceding text generation category of the reverse language generation model through a hard prompt, and controlling the preceding text generation toxicity of the model by setting parameters, so that the finally generated preceding text has greater inducibility. The application solves the problems of high cost of existing pre-training language model text generation safety detection and inability to control preceding text inducibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to a method, system, and storage medium for attacking dialogue models based on reverse generation. Background Technology

[0002] With technological advancements, large pre-trained language models have made significant progress in natural language generation tasks. While these models can generally generate high-quality text, they can also produce offensive or biased text. This issue severely hinders their use in real-world applications, especially in human-computer interaction scenarios such as dialogue. Some existing chatbots have been forced offline within a day of release due to offensive and discriminatory remarks. Therefore, detecting and fixing security vulnerabilities in these models is crucial.

[0003] Existing methods for detecting security issues primarily involve constructing inputs to a model through template generation, extracting real data, manual writing, and automated generation of large models, followed by observing the model's output. These constructed inputs vary significantly in their ability to induce unsafe responses. However, no systematic work has yet explored the factors influencing the ability of an input to induce unsafe responses (its inducementability). Summary of the Invention

[0004] This invention provides a method, system, and storage medium for attacking dialogue models based on reverse generation, in order to solve the problems of high security detection cost and inability to control the persuasiveness of the preceding text in existing pre-trained language models.

[0005] This invention provides a method for attacking dialogue models based on reverse generation, comprising:

[0006] The initial text data is measured using a pre-defined classifier to determine the relationship between text toxicity and text category and text persuasiveness.

[0007] A reverse language generation model is established, and the reverse language generation model is trained using a loss function. Based on a given response, the above text is generated through the trained reverse language generation model.

[0008] The reverse language generation model controls the generation category of the preceding text through hard prompts and controls the toxicity of the preceding text generation by setting parameters, so that the final generated preceding text has greater persuasiveness.

[0009] According to the present invention, a dialogue model attack method based on reverse generation is provided, wherein the method measures the initial context data using a preset classifier to determine the relationship between context toxicity and context category and context persuasiveness, specifically including:

[0010] The toxicity of the initial data is measured using a pre-defined classifier.

[0011] The initial data is input into a preset dialogue model and the first response result is generated repeatedly.

[0012] The proportion of unsafe responses in the first response is used as the persuasiveness of the preceding text.

[0013] According to the present invention, a dialogue model attack method based on reverse generation is provided, wherein the method measures the initial context data using a preset classifier to determine the relationship between context toxicity and context category and context persuasiveness, and further includes:

[0014] The categories of the initial text data are measured using a preset classifier;

[0015] The initial data is input into a preset dialogue model and the second response is generated repeatedly.

[0016] The proportion of insecurity in the second response is calculated as the leading indicator in the above text.

[0017] According to the present invention, a dialogue model attack method based on reverse generation is provided, wherein the method involves establishing a reverse language generation model, training the reverse language generation model using a loss function, and generating the preceding text based on a given response using the trained reverse language generation model, specifically including:

[0018] The reverse language generation model is trained using a pre-defined training loss function.

[0019] Based on the given response input, the trained reverse language generation model generates the relevant preceding text.

[0020] According to the present invention, a dialogue model attack method based on reverse generation is provided, which controls the preceding text generation category of the reverse language generation model through hard prompts, specifically including:

[0021] The hard prompt, set to the category name, is appended to the end of the given response in the input. The reverse language generation model generates the context corresponding to the category and performs training optimization.

[0022] During the training and optimization process, the word vectors of the resource token corresponding to the hard prompt are optimized together with the parameters of other parts of the model.

[0023] According to the present invention, a dialogue model attack method based on reverse generation is provided, wherein the model controls the generation of context toxicity by setting parameters to make the final generated context more persuasive, specifically including:

[0024] The models with the set parameters include: a basic reverse generation model, a toxic reverse generation model, and a language model;

[0025] The basic inverse generation model is used to model P. θ (c t |c <t The toxicity-reverse generation model generates a more toxic preceding text, modeling P. γ (c t |c <t The language model is modeled as follows:

[0026] The basic reverse generation model, the toxicity reverse generation model, and the language model are fused to determine the total generation probability and generate the final text.

[0027] This invention also provides a dialogue model attack system based on reverse generation, the system comprising:

[0028] The persuasiveness factor determination module is used to measure the initial text data through a preset classifier to determine the relationship between text toxicity and text category and text persuasiveness.

[0029] The modeling module is used to build a reverse language generation model and train the reverse language generation model using a loss function. Based on a given response, the above text is generated by the trained reverse language generation model.

[0030] The control and adjustment module is used to control the generation category of the reverse language generation model through hard prompts and to control the toxicity of the generated text by setting parameters, so that the final generated text has greater persuasiveness.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the above-described reverse generation-based dialogue model attack methods.

[0032] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the dialogue model attack method based on reverse generation as described above.

[0033] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the reverse generation-based dialogue model attack method as described above.

[0034] This invention provides a method, system, and storage medium for attacking dialogue models based on reverse generation. By determining the relationship between the toxicity and category of the preceding text and its persuasiveness, a reverse language generation model is established. The model controls the category and toxicity of the preceding text. It not only comprehensively explores the factors influencing the persuasiveness of the preceding text but also proposes an efficient reverse generation method for constructing adversarial data. This method can control the category, persuasiveness, and toxicity of the generated preceding text. It generates a large amount of highly persuasive data to better test the model's security and can further be used to enhance the model's security. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0036] Figure 1 This is one of the flowcharts of a dialogue model attack method based on reverse generation provided by the present invention;

[0037] Figure 2 This is the second flowchart of a dialogue model attack method based on reverse generation provided by the present invention;

[0038] Figure 3 This is the third flowchart of a dialogue model attack method based on reverse generation provided by the present invention;

[0039] Figure 4 This is the fourth flowchart of a dialogue model attack method based on reverse generation provided by the present invention;

[0040] Figure 5 This is a schematic diagram of the module connections of a dialogue model attack system based on reverse generation provided by the present invention;

[0041] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention.

[0042] Figure label:

[0043] 110: Module for identifying inducing factors; 120: Modeling module; 130: Control and adjustment module;

[0044] 610: Processor; 620: Communication interface; 630: Memory; 640: Communication bus. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0046] The following is combined with Figures 1-4 This invention describes a dialogue model attack method based on reverse generation, comprising:

[0047] S100. Measure the initial text data using a preset classifier to determine the relationship between text toxicity and text category and text persuasiveness.

[0048] S200. Establish a reverse language generation model and train the reverse language generation model using a loss function. Generate the above text based on the given response using the trained reverse language generation model.

[0049] S300. The reverse language generation model controls the generation category of the preceding text through hard prompts, and controls the toxicity of the preceding text generation by setting parameters, so that the final generated preceding text has greater persuasiveness.

[0050] To construct a large number of highly persuasive topologies, this invention proposes a reverse generation attack method that generates related topologies based on existing context text. This invention also supports control over the reverse generation, including controlling the category of generated topologies, increasing their persuasiveness, and reducing their toxicity. Based on the reverse generation method, this invention augments the existing BAD dataset to obtain nearly 8 times more persuasive topologies, classifies them into 12 categories, and demonstrates that topologies obtained using reverse generation can effectively improve the security level of dialogue models.

[0051] By measuring the initial text data using a pre-defined classifier, the relationship between text toxicity, text category, and text persuasiveness is determined, specifically including:

[0052] S101. Measure the toxicity of the initial above data using a preset classifier;

[0053] S102. Input the initial above text data into the preset dialogue model and repeat the process multiple times to generate the first response result;

[0054] S103. Calculate the proportion of unsafe results in the first response as the leading indicator in the above text.

[0055] This invention utilizes existing classifiers to measure the toxicity of BAD (Bad Context) data contexts. These contexts are then fed into three popular dialogue models—Blender, DialoGPT, and Plato2—for repeated generation, and the proportion of unsafe responses is calculated as the context's persuasiveness. For a given context c and the language model M to be tested, it is assumed that M samples a series of responses R = {r1, r2, ..., r...} given c. |R| Let the success rate of inducing model M by c above be defined as: ∑ r∈R `unsafe(c,r) / |R|`. Here, `unsafe(c,r)` returns 0 if the response is safe, otherwise it returns 1. To balance computational overhead and estimation error, 10 responses are sampled for each context, i.e., |R| = 10. This invention finds that there is a positive correlation between overall context toxicity and context persuasiveness; the stronger the context toxicity, the stronger the context persuasiveness generally is.

[0056] The method involves measuring the initial text data using a pre-defined classifier to determine the relationship between text toxicity, text category, and text persuasiveness. This also includes:

[0057] The categories of the initial text data are measured using a preset classifier;

[0058] The initial data is input into a preset dialogue model and the second response is generated repeatedly.

[0059] The proportion of insecurity in the second response is calculated as the leading indicator in the above text.

[0060] This invention utilizes existing classifiers to measure the categories of BAD (Bad) data context (a total of 12 unsafe categories). These contexts are then fed into three popular dialogue models—Blender, DialoGPT, and Plato2—for repeated generation, and the proportion of unsafe responses is calculated as context persuasiveness. This invention reveals that context category is also a significant factor influencing context persuasiveness. For example, contexts of category A may be more toxic than those of category B, but the overall persuasiveness of contexts of category A is lower than that of contexts of category B.

[0061] A reverse language generation model is established, and the model is trained using a loss function. Based on a given response, the above text is generated using the trained reverse language generation model, specifically including:

[0062] S201. Train the reverse language generation model using a preset training loss function;

[0063] S202. Based on the given response input, input the trained reverse language generation model to generate the relevant preceding text.

[0064] The core idea of ​​the basic reverse generation method in this invention is to generate related context (which may induce the context of this response) based on a given response r = {r1, r2, ..., r M The basic reverse generation method requires using a reverse language model to generate a related context c = {c1, c2, ..., c}. N The training loss function for the reverse language model is as follows:

[0065]

[0066] The reverse language generation model controls the generated context category via hard prompts, specifically including:

[0067] S301. The hard prompt set as the category name is appended to the end of the given response in the input. The reverse language generation model generates the context corresponding to the category and performs training optimization.

[0068] S302. During the training and optimization process, the word vector of the resource credential token corresponding to the hard prompt is optimized together with the parameters of other parts of the model.

[0069] While the basic reverse generation method in this invention can generate context related to a given response, it cannot control the category of the generated context. Therefore, this invention proposes a hard prompt-based method to control the category of the generated context. Specifically, the hard prompt, set as the category name, is appended to the input response, and the model needs to generate context corresponding to that category. During training and optimization, the word vector of the token (the resource credential required to access the resource interface API) corresponding to the hard prompt is optimized along with the parameters of other parts of the model.

[0070] Hard prompts are suggestions composed of specific Chinese or English words; they are human-readable. Prompts initially started with hand-designed templates. Hand-designed templates are generally based on human natural language knowledge, striving for semantically fluent and efficient templates. For example, Petroni et al. hand-designed cloze templates for knowledge probe tasks on the famous LAMA dataset; Brown et al. designed prefix templates for question answering, translation, and probe tasks. The advantage of hand-designed templates is their intuitiveness, but the disadvantage is that they require a lot of experimentation, experience, and language expertise, making them costly.

[0071] To address the shortcomings of manually designed templates, many studies have begun exploring how to automatically learn suitable templates. Automatically learned templates can be broadly categorized into two types: discrete prompts and continuous prompts. Discrete prompts mainly include Prompt Mining, Prompt Paraphrasing, Gradient-based Search, Prompt Generation, and Prompt Scoring; continuous prompts mainly include Prefix Tuning, Tuning Initialized with Discrete Prompts, and Hard-Soft Prompt Hybrid Tuning.

[0072] By controlling the generation toxicity of the supertext through a model with set parameters, the final generated supertext becomes more persuasive. Specifically, this includes:

[0073] The models with the set parameters include: a basic reverse generation model, a toxic reverse generation model, and a language model;

[0074] The basic inverse generation model is used to model P. θ (c t |c <t The toxicity-reverse generation model generates a more toxic preceding text, modeling P. γ (c t |c <t The language model is modeled as follows:

[0075] The basic reverse generation model, the toxicity reverse generation model, and the language model are fused to determine the total generation probability and generate the final text.

[0076] While more toxic adversarial topologies are generally more persuasive, less toxic adversarial topologies are more difficult to detect and therefore more harmful. Therefore, this invention proposes to simultaneously reduce the toxicity and increase the persuasiveness of the generated topologies during the reverse generation process, thereby constructing more threatening adversarial topologies. For efficient control, this invention trains three models with different parameters, each consisting of a basic reverse generation model used to model P. θ (c t |c <t ,r), a toxicity-reverse generation model that generates more toxic contexts, modeled by P γ (c t |c <t (r), and a language model for modeling After training these three models, the total generation probability during the inference generation phase is obtained by fusing these three probabilities:

[0077]

[0078] Where α and β are manually selected hyperparameters. The second term in the formula reduces the toxicity of the generated text, because P γ The first term assigns a higher probability to toxic tokens, thus causing the second term to assign a lower probability to toxic tokens. The third term aims to utilize Bayes' theorem. To improve the success rate of generating the above-mentioned inducement, since P(r) is fixed for a given r, improving... The value of P(r|c) can be increased, and when r is an unsafe response, increasing the value of P(r|c) helps to increase the success rate of the generated above c.

[0079] In a specific example, the calculation methods for the inducement success rate, toxicity, and category are as follows: The model samples 10 responses for each inducement to calculate the inducement success rate, using a top-k sampling algorithm (k=10). To determine whether a response is safe, this invention combines the results of two mainstream classifiers: Perspective API (P-API) and the BAD classifier. A response is considered safe only if both classifiers deem it safe; otherwise, it is considered unsafe.

[0080] The toxicity mentioned above is obtained by the score of the toxicity attribute returned by P-API. The higher the score, the stronger the toxicity.

[0081] To obtain the above categories, this invention combines the results of P-API and a sensitive topic classifier. The former includes six unsafe categories: identity attack, insult, profanity, threats, sexual harassment, and flirting; the latter includes six sensitive topics: drugs, medicine, inappropriate content, and others. If P-API scores for all six categories are not higher than 0.5, the result of the sensitive topic classifier is used as the category; otherwise, the category with the highest P-API score is used.

[0082] Training and inference parameters for the reverse generative model. All reverse generative models use DialoGPT-medium.

[0083] The maximum number of training epochs is set to 5, and training stops once the loss value on the validation set fails to decrease. Parameter updates use the AdamW optimizer with a learning rate of 2e-5 and 8 samples per update.

[0084] The sampling method used during inference is the top-p sampling algorithm (p=0.9), and the hyperparameters for probability combination during inference are set to α=2 and β=2.

[0085] This invention provides a reverse generation-based attack method for dialogue models. By determining the relationship between the toxicity and category of the preceding text and its persuasiveness, a reverse language generation model is established. The method controls the category and toxicity of the preceding text, comprehensively exploring the factors influencing its persuasiveness. Furthermore, it proposes an efficient reverse generation method for constructing adversarial data, controlling the category, persuasiveness, and toxicity of the generated preceding text. This generates a large amount of highly persuasive data to better test the model's security and can be further used to enhance the model's security.

[0086] refer to Figure 5 The present invention also discloses a dialogue model attack system based on reverse generation, the system comprising:

[0087] The persuasive factor determination module 110 is used to measure the initial text data through a preset classifier to determine the relationship between text toxicity and text category and text persuasiveness.

[0088] Modeling module 120 is used to build a reverse language generation model and train the reverse language generation model using a loss function, and generate the above text based on a given response through the trained reverse language generation model.

[0089] The control and adjustment module 130 is used to control the generation category of the reverse language generation model through hard prompts and to control the generation toxicity of the text by setting parameters, so that the final generated text has greater persuasiveness.

[0090] Among them, the inducibility factor determination module 110 measures the toxicity of the initial above data through a preset classifier;

[0091] The initial data is input into a preset dialogue model and the first response result is generated repeatedly.

[0092] The proportion of unsafe responses in the first response is used as the persuasiveness of the preceding text.

[0093] The categories of the initial text data are measured using a preset classifier;

[0094] The initial data is input into a preset dialogue model and the second response is generated repeatedly.

[0095] The proportion of insecurity in the second response is calculated as the leading indicator in the above text.

[0096] Modeling module 120 trains the reverse language generation model using a preset training loss function;

[0097] Based on the given response input, the trained reverse language generation model generates the relevant preceding text.

[0098] The control adjustment module 130 controls the text generation category of the reverse language generation model through hard prompts, specifically including:

[0099] The hard prompt, set to the category name, is appended to the end of the given response in the input. The reverse language generation model generates the context corresponding to the category and performs training optimization.

[0100] During the training and optimization process, the word vectors of the resource token corresponding to the hard prompt are optimized together with the parameters of other parts of the model.

[0101] By controlling the generation toxicity of the supertext through a model with set parameters, the final generated supertext becomes more persuasive. Specifically, this includes:

[0102] The models with the set parameters include: a basic reverse generation model, a toxic reverse generation model, and a language model;

[0103] The basic inverse generation model is used to model P. θ (c t |c <t The toxicity-reverse generation model generates a more toxic preceding text, modeling P. γ (c t |c <t The language model is modeled as follows:

[0104] The basic reverse generation model, the toxicity reverse generation model, and the language model are fused to determine the total generation probability and generate the final text.

[0105] This invention provides a reverse-generation-based dialogue model attack system. By determining the relationship between the toxicity and category of the preceding text and its persuasiveness, a reverse language generation model is established. The system controls the category and toxicity of the preceding text. It not only comprehensively explores the factors influencing the persuasiveness of the preceding text but also proposes an efficient reverse-generation method for constructing adversarial data. This method can control the category, persuasiveness, and toxicity of the generated preceding text. It generates a large amount of highly persuasive data to better test the model's security and can further be used to enhance the model's security.

[0106] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can invoke logical instructions in the memory 630 to execute a dialogue model attack method based on reverse generation. This method includes: measuring initial context data using a preset classifier to determine the relationship between context toxicity and context category and context persuasiveness.

[0107] A reverse language generation model is established, and the reverse language generation model is trained using a loss function. Based on a given response, the above text is generated through the trained reverse language generation model.

[0108] The reverse language generation model controls the generation category of the preceding text through hard prompts and controls the toxicity of the preceding text generation by setting parameters, so that the final generated preceding text has greater persuasiveness.

[0109] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0110] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, the computer program being executed by a processor, the computer being able to execute a dialogue model attack method based on reverse generation provided by the above methods, the method including: measuring the initial context data through a preset classifier, and determining the relationship between context toxicity and context category and context persuasiveness;

[0111] A reverse language generation model is established, and the reverse language generation model is trained using a loss function. Based on a given response, the above text is generated through the trained reverse language generation model.

[0112] The reverse language generation model controls the generation category of the preceding text through hard prompts and controls the toxicity of the preceding text generation by setting parameters, so that the final generated preceding text has greater persuasiveness.

[0113] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform a dialogue model attack method based on reverse generation provided by the above methods, the method comprising: measuring initial context data through a preset classifier to determine the relationship between context toxicity and context category and context persuasiveness;

[0114] A reverse language generation model is established, and the reverse language generation model is trained using a loss function. Based on a given response, the above text is generated through the trained reverse language generation model.

[0115] The reverse language generation model controls the generation category of the preceding text through hard prompts and controls the toxicity of the preceding text generation by setting parameters, so that the final generated preceding text has greater persuasiveness.

[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0117] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dialogue model attack method based on reverse generation, characterized in that, include: The initial text data is measured using a pre-defined classifier to determine the relationship between text toxicity and text category and text persuasiveness. A reverse language generation model is established, and the reverse language generation model is trained using a loss function. Based on a given response, the above text is generated through the trained reverse language generation model. The reverse language generation model controls the generation category of the preceding text through hard prompts and controls the toxicity of the preceding text generation by setting parameters, so that the final generated preceding text has greater persuasiveness. The reverse language generation model controls the generated context category via hard prompts, specifically including: The hard prompt, set to the category name, is appended to the end of the given response in the input. The reverse language generation model generates the context corresponding to the category and performs training optimization. During the training and optimization process, the word vectors of the resource token corresponding to the hard prompt are optimized together with the parameters of other parts of the model; The method of controlling the generation of toxicity through a model with set parameters, making the final generated toxicity more persuasive, specifically includes: The models with the set parameters include: a basic reverse generation model, a toxic reverse generation model, and a language model; The basic reverse generation model is used for modeling. The toxicity-reverse generation model generates more toxic preceding text, and models... The language model is modeled ; The basic reverse generation model, the toxicity reverse generation model, and the language model are fused to determine the total generation probability and generate the final text.

2. The dialogue model attack method based on reverse generation according to claim 1, characterized in that, The step of measuring the initial text data using a preset classifier to determine the relationship between text toxicity, text category, and text persuasiveness specifically includes: The toxicity of the initial data is measured using a pre-defined classifier. The initial data is input into a preset dialogue model and the first response result is generated repeatedly. The proportion of unsafe responses in the first response is used as the persuasiveness of the preceding text.

3. The dialogue model attack method based on reverse generation according to claim 1, characterized in that, The step of measuring the initial text data using a preset classifier to determine the relationship between text toxicity, text category, and text persuasiveness also includes: The categories of the initial text data are measured using a preset classifier; The initial data is input into a preset dialogue model and the second response is generated repeatedly. The proportion of insecurity in the second response is calculated as the leading indicator in the above text.

4. The dialogue model attack method based on reverse generation according to claim 1, characterized in that, The process of establishing a reverse language generation model, training the model using a loss function, and generating the above text based on a given response using the trained reverse language generation model specifically includes: The reverse language generation model is trained using a pre-defined training loss function. Based on the given response input, the trained reverse language generation model generates the relevant preceding text.

5. A dialogue model attack system based on reverse generation, characterized in that, The system includes: The persuasiveness factor determination module is used to measure the initial text data through a preset classifier to determine the relationship between text toxicity and text category and text persuasiveness. The modeling module is used to build a reverse language generation model and train the reverse language generation model using a loss function. Based on a given response, the above text is generated by the trained reverse language generation model. The control and adjustment module is used to control the generated context category of the reverse language generation model through hard prompts and to control the toxicity of the generated context through parameter-setting models, so that the final generated context is more persuasive. Controlling the generated context category through hard prompts specifically includes: a hard prompt set as the category name is appended to the end of the given input response; the reverse language generation model generates context of the corresponding category and performs training and optimization; during the training and optimization process, the word vectors of the resource credential token corresponding to the hard prompt and the parameters of other parts of the model are optimized together; controlling the toxicity of the generated context through parameter-setting models to make the final generated context more persuasive specifically includes: the parameter-setting models include: a basic reverse generation model, a toxicity reverse generation model, and a language model; the basic reverse generation model is used for modeling... The toxicity-reverse generation model generates more toxic preceding text, and models... The language model is modeled The basic reverse generation model, the toxicity reverse generation model, and the language model are fused to determine the total generation probability and generate the final text.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the dialogue model attack method based on reverse generation as described in any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the reverse generation-based dialogue model attack method as described in any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the reverse generation-based dialogue model attack method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Text content auditing method and device

    CN110674255A

  • Adversarial generative network for defending text malicious sample and training method thereof

    CN111046673A