Method and system for identifying potential vulnerabilities of pre-trained dialog models

CN117235734BActive Publication Date: 2026-09-25SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311193946.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-15
Publication Date
2026-09-25
Estimated Expiration
2043-09-15

AI Technical Summary

Technical Problem

[0004](1)这些预训练语言模型(PLM)是在大规模未经清理的数据集上进行预训练的,这些数据集可能包含负面观点和偏见等内容,从而影响用户使用对话系统的感受

Benefits of technology

[0039](1)本发明利用两阶段强化学习策略,在已知语料库的基础上生成预训练对话模型的对抗样本,再利用生成的对抗样本来训练对话模型,根据对话模型的输出结果是否包含恶意或不当的回复,从而识别出对话模型是否存在潜在漏洞,当存在潜在漏洞继续生成对抗样本,直至未识别出潜在漏洞,这样通过发现漏洞并生成高质量的对抗样本,为构建更鲁棒的对话系统和防御机制奠定了基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117235734B_ABST
    Figure CN117235734B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of potential vulnerability identification, and provides a pre-training dialogue model potential vulnerability identification method and system. The pre-training dialogue model potential vulnerability identification method comprises generating an adversarial sample of the pre-training dialogue model based on a known corpus and a two-stage reinforcement learning strategy; obtaining a corresponding output result based on the generated adversarial sample and the pre-training dialogue model, and then identifying whether the pre-training dialogue model has a potential vulnerability according to the output result; when there is a potential vulnerability, continue to generate an adversarial sample until no potential vulnerability is identified. The present application lays a foundation for building a more robust dialogue system and defense mechanism by discovering vulnerabilities and generating high-quality adversarial samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of potential vulnerability identification technology, and in particular relates to a method and system for identifying potential vulnerabilities in a pre-trained dialogue model. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] In recent years, pre-trained language models (PLMs) have been increasingly widely used in dialogue systems. However, the inventors have discovered that existing pre-trained language models (PLMs) have the following problems:

[0004] (1) These pre-trained language models (PLMs) are pre-trained on large-scale, uncleaned datasets that may contain negative opinions and biases, which may affect the user’s experience of using the dialogue system.

[0005] (2) The generation process of language models is difficult to control, and current pre-trained language models (PLMs) are vulnerable to adversarial attacks. For example, malicious dialogue responses may contain offensive or offensive content (such as hate speech, insults, and threats). These malicious responses may cause users to experience conversation interruptions in the real world, leading to discomfort or even causing them to abandon the corresponding dialogue system. Summary of the Invention

[0006] To address the technical problems existing in the background art, the present invention provides a method and system for identifying potential vulnerabilities in a pre-trained dialogue model, which can identify potential vulnerabilities in the pre-trained dialogue model and efficiently generate high-quality adversarial examples.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] The first aspect of the present invention provides a method for identifying potential vulnerabilities in a pre-trained dialogue model.

[0009] A method for identifying potential vulnerabilities in a pre-trained dialogue model, comprising:

[0010] Based on the known corpus and a two-stage reinforcement learning strategy, adversarial examples for a pre-trained dialogue model are generated.

[0011] Based on the generated adversarial examples and the pre-trained dialogue model, the corresponding output results are obtained. Then, based on the output results, it is identified whether there are potential vulnerabilities in the pre-trained dialogue model. If there are potential vulnerabilities, adversarial examples are generated again until no potential vulnerabilities are identified.

[0012] The process of generating adversarial examples for the pre-trained dialogue model is as follows:

[0013] During the offline learning phase, an iterative attacker and a reward evaluator are trained based on a known corpus and pseudo-trajectories of sentence editing operations. The iterative attacker is used to attack the pre-trained dialogue model, and the reward evaluator is used to assign a reward to each output of the pre-trained dialogue model and calculate a reward for malicious responses.

[0014] During the online learning phase, the iterative attacker trained in the offline learning phase is used to attack the pre-trained dialogue model to generate targeted responses that maximize malicious rewards without violating constraints. At the same time, the sentence editing operation trajectories generated by the iterative attacker are sampled to update the parameters of the iterative attacker and reward evaluator during the online learning phase.

[0015] As one embodiment of the first aspect of the invention, the iterative attacker merges the marker position prediction and operation prediction by directly generating a series of edit operations.

[0016] As one embodiment of the first aspect of the present invention, the editing operation of the iterative attacker includes:

[0017] Retain one mark operation, one replace operation, one delete operation, and one insert operation.

[0018] As one embodiment of the first aspect of the invention, the iterative attacker is used to perform editing operations on the current state, which includes the dialogue context and user utterances, and to use a delimiter after each utterance except the last utterance to output the representation in the original utterances, and to apply a linear layer to predict actions.

[0019] In one embodiment of the first aspect of the invention, the parameters of the iterative attacker and the reward evaluator are updated based on the generative adversarial network during both offline and online learning.

[0020] As one embodiment of the first aspect of the invention, the reward evaluator processes the representation in the original utterance output by the iterative attacker and feeds it into a multilayer perceptron to output a reward for each edit operation at the corresponding marked position.

[0021] A second aspect of the present invention provides a potential vulnerability identification system for a pre-trained dialogue model.

[0022] A potential vulnerability identification system for a pre-trained dialogue model, comprising:

[0023] The adversarial example generation module is used to generate adversarial examples for pre-trained dialogue models based on known corpora and two-stage reinforcement learning strategies.

[0024] The potential vulnerability identification module is used to obtain corresponding output results based on the generated adversarial samples and pre-trained dialogue models, and then identify whether there are potential vulnerabilities in the pre-trained dialogue model based on the output results. If there are potential vulnerabilities, it continues to generate adversarial samples until no potential vulnerabilities are identified.

[0025] The process of generating adversarial examples for the pre-trained dialogue model is as follows:

[0026] During the offline learning phase, an iterative attacker and a reward evaluator are trained based on a known corpus and pseudo-trajectories of sentence editing operations. The iterative attacker is used to attack the pre-trained dialogue model, and the reward evaluator is used to assign a reward to each output of the pre-trained dialogue model and calculate a reward for malicious responses.

[0027] During the online learning phase, the iterative attacker trained in the offline learning phase is used to attack the pre-trained dialogue model to generate targeted responses that maximize malicious rewards without violating constraints. At the same time, the sentence editing operation trajectories generated by the iterative attacker are sampled to update the parameters of the iterative attacker and reward evaluator during the online learning phase.

[0028] As one embodiment of the second aspect of the invention, the iterative attacker merges the marker position prediction and operation prediction by directly generating a series of edit operations.

[0029] As one embodiment of the second aspect of the present invention, the editing operation of the iterative attacker includes:

[0030] Retain one mark operation, one replace operation, one delete operation, and one insert operation.

[0031] As one embodiment of the second aspect of the invention, the iterative attacker is used to perform editing operations on the current state, which includes the dialogue context and user utterances, and to use a delimiter after each utterance except the last utterance to output the representation in the original utterances, and to apply a linear layer to predict actions.

[0032] As one embodiment of the second aspect of the invention, the parameters of the iterative attacker and the reward evaluator are updated based on the generative adversarial network during both offline and online learning.

[0033] As one embodiment of the second aspect of the invention, the reward evaluator is used to process the representation in the original utterance output by the iterative attacker and then feed it into a multilayer perceptron to output the reward for each edit operation at the corresponding marked position.

[0034] A third aspect of the present invention provides a computer-readable storage medium.

[0035] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps in the potential vulnerability identification method for a pre-trained dialogue model as described above.

[0036] A fourth aspect of the present invention provides an electronic device.

[0037] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the potential vulnerability identification method for a pre-trained dialogue model as described above.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] (1) This invention utilizes a two-stage reinforcement learning strategy to generate adversarial samples for a pre-trained dialogue model based on a known corpus. The generated adversarial samples are then used to train the dialogue model. Based on whether the output of the dialogue model contains malicious or inappropriate responses, potential vulnerabilities in the dialogue model can be identified. If potential vulnerabilities exist, adversarial samples are generated until no potential vulnerabilities are identified. In this way, by discovering vulnerabilities and generating high-quality adversarial samples, a foundation is laid for building a more robust dialogue system and defense mechanism.

[0040] (2) This invention utilizes the idea that generated adversarial samples are visually difficult to detect, but can trigger vulnerabilities in pre-trained dialogue models. While identifying potential vulnerabilities in pre-trained dialogue models, it can also efficiently generate high-quality adversarial samples, thereby improving the anti-attack performance of dialogue models.

[0041] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0042] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0043] Figure 1 This invention provides dialogue models using and without an iterative attacker.

[0044] Figure 2 This is a schematic diagram illustrating the use of an iterative attacker to edit the original discourse and obtain adversarial samples in an embodiment of the present invention;

[0045] Figure 3 This is a schematic diagram illustrating how the reward evaluator in this embodiment of the invention evaluates the reward for each location based on the generated trajectory and malicious feedback.

[0046] Figure 4 This is a schematic diagram of the process of generating pseudo-trajectories according to an embodiment of the present invention. Detailed Implementation

[0047] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0048] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0049] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0050] Example 1

[0051] This embodiment provides a method for identifying potential vulnerabilities in a pre-trained dialogue model, which specifically includes the following steps:

[0052] Step 1: Generate adversarial examples for a pre-trained dialogue model based on the known corpus and a two-stage reinforcement learning strategy.

[0053] The process of generating adversarial examples for the pre-trained dialogue model is as follows:

[0054] Step 1.1: In the offline learning phase, an iterative attacker and a reward estimator are trained based on a known corpus and pseudo-trajectories of sentence editing operations; the iterative attacker is used to attack the pre-trained dialogue model, and the reward estimator is used to assign a reward to each output of the pre-trained dialogue model and calculate the reward for malicious responses.

[0055] Step 1.2: In the online learning phase, the iterative attacker trained in the offline learning phase is used to attack the pre-trained dialogue model to generate targeted responses that maximize malicious rewards without violating constraints. At the same time, the sentence editing operation trajectories generated by the iterative attacker are sampled to update the parameters of the iterative attacker and reward evaluator in the online learning phase.

[0056] according to Figure 1This embodiment employs a two-stage reinforcement learning framework to generate adversarial examples. The model modifies the original utterance into adversarial utterance, thereby triggering the victim model to generate malicious responses. For editing the original utterance, the iterative attacker uses four operations: "keep," "replace," "delete," and "insert." Unlike existing methods, the iterative attacker considers the correlation between different edits and can iteratively edit the utterance multiple times until convergence. The training of the iterative attacker uses a reinforcement learning framework that interacts with the victim model. This task presents two challenges: First, the editing reward at each utterance position is unknown; this invention learns the malicious reward from the malicious feedback of the entire utterance through a reward evaluator. Second, the attack efficiency is low due to the computational intensity of the pre-trained dialogue model; therefore, this embodiment designs a two-stage reinforcement learning strategy, including an offline warm-up stage and an online fine-tuning stage to train the model. In the offline learning stage, the parameters of the iterative attacker and reward evaluator networks are warmed up using a pseudo-trajectory generation process. In the online learning stage, the iterative attacker directly attacks the victim model, and the iterative attacker is updated through the reward evaluator, which learns the reward for each edit based on the attack results.

[0057] This embodiment constructs a malicious dialogue attack as a constrained Markov decision process. In this process, the iterative attacker attempts to attack the victim model, causing the victim model to generate a targeted response that maximizes the malicious reward without violating the constraints. The constrained Markov decision process is defined as follows:

[0058] 1) S is the state space, s t ∈S is the current state, which includes the dialogue context and the user's utterance, where t represents the number of modifications. This indicates an adversarial example.

[0059] 2) A represents the action space, a t ∈A is a sequence of independent actions, where a t =[a t,1 a t,2 ..., a t,n ], n represents The length of a t,j It is the modification action for the t-th tag.

[0060] 3) P: S×A×S→[0,1] is a deterministic transfer equation.

[0061] 4)r ψ Let c represent the reward function. i Let i represent the i-th constraint function.

[0062] 5) γ is the discount factor.

[0063] The optimization objective is:

[0064]

[0065] The constraints are:

[0066]

[0067]

[0068]

[0069] Where, π θ It is a policy network implemented by an iterative attacker; θ is the parameter of the iterative attacker; r ψ ψ is the reward function, implemented by the reward evaluator; c1, c2, and c3 are the cost functions for consistency, fluency, and non-maliciousness, respectively; α1, α2, and α3 are predefined cost function thresholds. The constraints are as follows:

[0070] 1) It is a consistency constraint. This is to ensure consistency in adversarial examples. To maintain consistency with x, this invention uses a BERT model to embed x and The cosine similarity between BERT embeddings was used as the consistency score.

[0071] 2) Jc2(πθ) is a fluency constraint. This invention uses a perplexity score to measure... Fluency is assessed by penalizing actions that result in non-fluid output. The perplexity score is defined as follows: Where n is The number of tags in w j It is the j-th marker.

[0072] 3)Jc3(π θ This is a non-malicious constraint. To ensure adversarial examples... Without introducing additional malfunctions, use x and The change in the malice confidence score between the two is used as the non-malice score, i.e. S is the malicious confidence score based on a BERT-based binary malicious classifier.

[0073] This embodiment applies the Lagrange relaxation method to approximate a constrained optimization problem and aims to solve the following problems:

[0074]

[0075] Where λ1, λ2, and λ3 are Lagrange multipliers.

[0076] Figure 2This diagram illustrates the use of an iterative attack tool to edit original utterances and obtain adversarial examples. Unlike previous methods, sentence editing is divided into two parts: marker position prediction and operation prediction. In practice, the iterative attack tool merges the marker position prediction and operation prediction by directly generating a series of editing operations. Specifically, the editing operations of the iterative attack tool include:

[0077] Retain one mark operation, one replace operation, one delete operation, and one insert operation.

[0078] For example, "K" indicates keeping a marker, "R" indicates replacement, "D" indicates deletion, and "I" indicates insertion; this invention constructs an editing policy network based on BERT, namely an iterative attacker, π. θ Used to generate a series of operations. Iterative attackers use... As input, and using a delimiter [SEP] after each utterance (except the last utterance). This embodiment labels the input BERT to output the representation H in the original utterance. t and apply with parameter W θ and b θ Linear layers to predict action a t :

[0079]

[0080] Among them, W θ and b θ This indicates the parameters of the attacker.

[0081] The iterative attacker is used to edit the current state, which includes the dialogue context and user utterances, and uses a delimiter after each utterance except the last one to output the representation in the original utterance, and applies a linear layer to predict actions.

[0082] A reward evaluator is used to process the representation in the original utterance output by the iterative attacker and then feed it into a multilayer perceptron to output the reward for each edit operation at the corresponding marked position.

[0083] This embodiment learns a reward evaluator r parameterized by ψ. ψ This is used to estimate the reward for each token. Here, the hidden state of BERT (i.e., the representation in the original utterance) H is used. t As input, they are fed into a multilayer perceptron (MLP) to output the reward for each edit operation at the corresponding marked location:

[0084]

[0085] Among them, W ψ and b ψThese are the parameters of the attacker.

[0086] To accelerate training, this embodiment utilizes the existing MDRDC corpus for offline warm-up of the iterative attacker and reward estimator. First, this embodiment uses trajectories generated by a pseudo-trajectory generation process to train the iterative attacker, reward estimator, and λ. The parameters of the iterative attacker and reward estimator are then passed to the online learning phase. Next, the iterative attacker is used to attack the victim model, and the trajectories generated by the iterative attacker are sampled to update the iterative attacker, reward estimator, and λ during the online learning phase. The gradient of the reward estimator is expressed as follows:

[0087]

[0088] Among them, D + and D - These are trajectories with malicious and non-malicious feedback, respectively, originating from the malicious classifier; τ is from... arrive State-action pair trajectory, i.e. This represents the gradient of the loss function with respect to the parameter ψ; and Indicates D + and D - Find the expected value of the trajectory r; ψ (τ) represents the sum of rewards for each state-action pair in the trajectory. A pseudo-trajectory is used as D during the offline learning phase. + and D - During the online learning phase, actual trajectories are used as D. + and D - . Figure 3 A schematic diagram is provided showing how the reward evaluator evaluates the reward for each location based on the generated trajectory and malicious feedback.

[0089] To train the iterative attacker and the Lagrange multiplier λ, the gradient is expressed as:

[0090]

[0091]

[0092] Where, D = D + ∪D - .

[0093] In both offline and online learning, the parameters of the iterative attacker, reward evaluator, and λ are updated based on the generative adversarial network:

[0094]

[0095]

[0096]

[0097] Where η1, η2 and η3 are the learning rates of the three modules, respectively.

[0098] In the offline learning phase, this embodiment uses Figure 4 The process of generating pseudo-trajectory D pseu D pseu Malicious feedback signals were used to train the parameters of the reward evaluator, iterative attacker, and λ. Γ represents the input of each edit operation a t′ Replace with its opposite editing operation, for example, if a t′ It is “D”, Γ(a t′ ) then becomes "I"; f mlm This represents a masked language model. To improve the training efficiency of the policy network, this embodiment selects a subset of available trajectories. In practice, optimal performance is achieved by varying the ratio of malicious to non-malicious trajectories; when training the reward evaluator, pseudo-trajectories D are used. pssu To replace D + and D - .

[0099] During the online learning phase, this invention trains the iterative attacker and reward estimator based on the parameters learned in the offline phase and reinitializes λ. Similar to offline training, the reward estimator is used to assign a reward to each edit operation and calculate malicious feedback. The difference is that during the online attack process, this embodiment samples the actual trajectory D from the iterative attacker. actu .

[0100] Step 2: Based on the generated adversarial examples and the pre-trained dialogue model, obtain the corresponding output results, and then identify whether there are potential vulnerabilities in the pre-trained dialogue model based on the output results. If there are potential vulnerabilities, continue to generate adversarial examples until no potential vulnerabilities are identified.

[0101] This embodiment utilizes a two-stage reinforcement learning strategy to generate adversarial examples of a pre-trained dialogue model based on a known corpus. The generated adversarial examples are then used to test the dialogue model. By analyzing whether the dialogue model's output contains malicious or inappropriate responses, potential vulnerabilities in the dialogue model can be identified. If potential vulnerabilities are found, adversarial examples are generated continuously until no potential vulnerabilities are identified. In this way, by discovering vulnerabilities and generating high-quality adversarial examples, a foundation is laid for building a more robust dialogue system and defense mechanism.

[0102] Example 2

[0103] This embodiment provides a potential vulnerability identification system for a pre-trained dialogue model, which includes:

[0104] (1) Adversarial sample generation module, which is used to generate adversarial samples of pre-trained dialogue models based on known corpora and two-stage reinforcement learning strategies.

[0105] The process of generating adversarial examples for the pre-trained dialogue model is as follows:

[0106] During the offline learning phase, an iterative attacker and a reward evaluator are trained based on a known corpus and pseudo-trajectories of sentence editing operations. The iterative attacker is used to attack the pre-trained dialogue model, and the reward evaluator is used to assign a reward to each output of the pre-trained dialogue model and calculate a reward for malicious responses.

[0107] During the online learning phase, the iterative attacker trained in the offline learning phase is used to attack the pre-trained dialogue model to generate targeted responses that maximize malicious rewards without violating constraints. At the same time, the sentence editing operation trajectories generated by the iterative attacker are sampled to update the parameters of the iterative attacker and reward evaluator during the online learning phase.

[0108] In its implementation, the iterative attacker merges the marker position prediction and operation prediction by directly generating a series of editing operations. These editing operations include:

[0109] Retain one mark operation, one replace operation, one delete operation, and one insert operation.

[0110] The iterative attacker is used to edit the current state, which includes the dialogue context and user utterances, and uses a delimiter after each utterance except the last one to output the representation in the original utterance, and applies a linear layer to predict actions.

[0111] In both offline and online learning, the parameters of the iterative attacker and the reward evaluator are updated based on the generative adversarial network.

[0112] The reward evaluator processes the representation in the original utterance output by the iterative attacker and then feeds it into a multilayer perceptron to output the reward for each edit operation at the corresponding marked position.

[0113] (2) Potential vulnerability identification module, which is used to obtain the corresponding output results based on the generated adversarial samples and pre-trained dialogue model, and then identify whether there are potential vulnerabilities in the pre-trained dialogue model based on the output results. If there are potential vulnerabilities, it continues to generate adversarial samples until no potential vulnerabilities are identified.

[0114] It should be noted that each module in this embodiment corresponds one-to-one with each step in Embodiment 1, and their specific implementation processes are the same, so they will not be repeated here.

[0115] Example 3

[0116] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the pre-trained dialogue model potential vulnerability identification method described above.

[0117] Example 4

[0118] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the potential vulnerability identification method for the pre-trained dialogue model described above.

[0119] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0120] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for identifying potential vulnerabilities in a pre-trained dialogue model, characterized in that, include: Based on the known corpus and a two-stage reinforcement learning strategy, adversarial examples for a pre-trained dialogue model are generated. Based on the generated adversarial examples and the pre-trained dialogue model, the corresponding output results are obtained; Based on the output results, identify whether there are potential vulnerabilities in the pre-trained dialogue model. If potential vulnerabilities are found, continue to generate adversarial examples until no potential vulnerabilities are identified. The process of generating adversarial examples for the pre-trained dialogue model is as follows: During the offline learning phase, an iterative attacker and a reward evaluator are trained based on a known corpus and pseudo-trajectories of sentence editing operations. The iterative attacker is used to attack the pre-trained dialogue model, and the reward evaluator is used to assign a reward to each output of the pre-trained dialogue model and calculate a reward for malicious responses. During the online learning phase, the iterative attacker trained in the offline learning phase is used to attack the pre-trained dialogue model to generate targeted responses that maximize malicious rewards without violating constraints. At the same time, the sentence editing operation trajectories generated by the iterative attacker are sampled to update the parameters of the iterative attacker and reward evaluator during the online learning phase.

2. The method for identifying potential vulnerabilities in a pre-trained dialogue model as described in claim 1, characterized in that, The iterative attacker merges the marker location prediction and operation prediction by directly generating a series of editing operations.

3. The method for identifying potential vulnerabilities in a pre-trained dialogue model as described in claim 2, characterized in that, The editing operations of the iterative attacker include: Retain one mark operation, one replace operation, one delete operation, and one insert operation.

4. The method for identifying potential vulnerabilities in a pre-trained dialogue model as described in claim 1 or 2, characterized in that, The iterative attacker is used to edit the current state, which includes the dialogue context and user utterances, and uses a delimiter after each utterance except the last one to output the representation in the original utterance, and applies a linear layer to predict actions.

5. The method for identifying potential vulnerabilities in a pre-trained dialogue model as described in claim 1, characterized in that, In both offline and online learning, the parameters of the iterative attacker and the reward evaluator are updated based on the generative adversarial network.

6. The method for identifying potential vulnerabilities in a pre-trained dialogue model as described in claim 1, characterized in that, The reward evaluator processes the representation in the original utterance output by the iterative attacker and then feeds it into a multilayer perceptron to output the reward for each edit operation at the corresponding marked position.

7. A potential vulnerability identification system for a pre-trained dialogue model, characterized in that, include: The adversarial example generation module is used to generate adversarial examples for pre-trained dialogue models based on known corpora and two-stage reinforcement learning strategies. The potential vulnerability identification module is used to obtain corresponding output results based on the generated adversarial samples and pre-trained dialogue models, and then identify whether there are potential vulnerabilities in the pre-trained dialogue model based on the output results. If there are potential vulnerabilities, it continues to generate adversarial samples until no potential vulnerabilities are identified. The process of generating adversarial examples for the pre-trained dialogue model is as follows: During the offline learning phase, an iterative attacker and a reward evaluator are trained based on a known corpus and pseudo-trajectories of sentence editing operations. The iterative attacker is used to attack the pre-trained dialogue model, and the reward evaluator is used to assign a reward to each output of the pre-trained dialogue model and calculate a reward for malicious responses. During the online learning phase, the iterative attacker trained in the offline learning phase is used to attack the pre-trained dialogue model to generate targeted responses that maximize malicious rewards without violating constraints. At the same time, the sentence editing operation trajectories generated by the iterative attacker are sampled to update the parameters of the iterative attacker and reward evaluator during the online learning phase.

8. The potential vulnerability identification system for a pre-trained dialogue model as described in claim 7, characterized in that, The iterative attacker merges the marker position prediction and operation prediction by directly generating a series of edit operations; the edit operations of the iterative attacker include: Retain one mark operation, one replace operation, one delete operation, and one insert operation; or The iterative attacker is used to edit the current state, which includes the dialogue context and user utterances, and uses a delimiter after each utterance except the last one to output the representation in the original utterance, and applies a linear layer to predict actions. or In both offline and online learning, the parameters of the iterative attacker and the reward evaluator are updated based on the generative adversarial network; or The reward evaluator processes the representation in the original utterance output by the iterative attacker and then feeds it into a multilayer perceptron to output the reward for each edit operation at the corresponding marked position.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the potential vulnerability identification method for a pre-trained dialogue model as described in any one of claims 1-6.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the potential vulnerability identification method for a pre-trained dialogue model as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Model privacy protection method and system for deep reinforcement learning

    CN113420326A

  • Attack path prediction model generation method and device based on network attack graph

    CN116582349A