Black box cross-task backdoor hint attack method based on general trigger
By introducing a black box cross-task backdoor prompt attack method based on general triggers in the pre-trained language model, the potential threats in the security of language models in the prior art are solved, and the effect of efficient injection of backdoors in multiple tasks is achieved.
Patent Information
- Application Number
- CN202510375766.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-13
AI Technical Summary
Existing pre-trained language models have potential threats in terms of security and are vulnerable to malicious attacks, especially in security-sensitive tasks, suggesting that security has gradually become a new research hotspot.
A black box cross-task backdoor prompt attack method based on universal triggers is adopted. A reinforcement learning framework searches for general triggers, a gradient-free poisoning data set is built, and a general trigger is introduced into a new task through prompt fine-tuning training, so that it can generate target output when it encounters a prompt containing a universal trigger.
It realizes efficient generation of triggers in a black box environment, avoids dependence on internal information of the model, and successfully injects backdoors into multiple tasks, maintains high precision, and general triggers have strong cross-task migration capabilities, enhancing the concealment and adaptability of backdoor attacks.
Smart Images

Figure CN120146150A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of pre-trained language model security, and in particular, to a black-box cross-task backdoor prompt attack method based on a general trigger. Background Art
[0002] In recent years, in the prior art, the use of pre-trained language models, such as BERT, GPT, LLaMA, etc., as the backbone of downstream natural language processing tasks has refreshed the state-of-the-art performance of many NLP tasks; pre-trained language models (PLMs) can effectively extract rich language knowledge from a large amount of unlabeled data, which is extremely beneficial to downstream tasks. The prompting technique plays a key role in these successes, enabling pre-trained language models (PLMs) to efficiently and effectively adapt to various downstream tasks.
[0003] The prompt-based pre-trained language model (PLM) includes three main steps: First, a pre-trained language model (PLM) with a masked pre-training task is required. Second, a cloze template needs to be constructed. For example, in a sentiment classification task, for the sentence "I am very satisfied with this shopping.", a template like "The sentiment of this sentence is [mask]." can be created, and then the pre-trained language model (PLM) is used to predict which sentiment word (e.g., "good", "bad") should be filled in the masked position; finally, the answer predicted by the pre-trained language model (PLM) is converted into a real label. Words such as "great" and "wonderful" should correspond to positive emotions, while words such as "terrible" and "awful" should correspond to negative emotions. Researchers call this mapping a verbalizer.
[0004] In the prior art, prompts have become more and more general and tool-based; in the openprompt framework developed by Tsinghua University, users can easily import templates and verbalizers, and then load downstream task prompts trained by third parties in their own models; however, while bringing convenience, it also exposes some potential security threats. Research shows that these prompt models are vulnerable to malicious attacks.
[0005] Considering its increasing application in many real-world security-sensitive tasks, this vulnerability has attracted great attention and high concern about model security; therefore, prompt security has gradually become a new research hotspot, and more and more researchers are concerned about attacks and defenses. Summary of the Invention
[0006] The purpose of the present invention is to provide a black-box prompt attack method based on a general trigger that can efficiently perform cross-task backdoor attacks; the technical solution is as follows:
[0007] A black-box cross-task backdoor prompt attack method based on a general trigger, comprising the following steps:
[0008] Step 1: Search for a general trigger through a reinforcement learning framework. The reinforcement learning framework uses a continuous policy network to intelligently explore the trigger space. Without accessing the internal information of the pre-trained language model (PLM), a general trigger is generated by designing a search objective and a reward function.
[0009] Step 2: Use the general trigger generated in Step 1 to construct a gradient-free poisoned dataset. The poisoned dataset is composed of positive-polarity corpus samples combined with specific triggers and their corresponding negative labels. The generated poisoned samples are used for prompt fine-tuning training through the output of the pre-trained language model (PLM).
[0010] Step 3: Through prompt fine-tuning training, introduce the generated general trigger into a new task, so that the new task can produce a target output when encountering a prompt containing the general trigger.
[0011] Furthermore, the search process of the general trigger in Step 1 is as follows:
[0012] Due to the discrete nature and black-box setting of the trigger, gradient-based optimization is not feasible, while the exponential complexity of brute-force search is , is the vocabulary length, is the trigger length, represents the time complexity symbol of the algorithm, which is used to describe the upper bound of the search space scale; even for relatively short triggers, the search space will grow explosively;
[0013] Based on this, the search process of the general trigger is modeled as a reinforcement learning search process, using a continuous policy network to explore the trigger space, that is, defining corresponding search objectives and reward functions to generate general triggers; specifically, rewrite the search process as the following problem:
[0014] ;
[0015] During the search process, the trigger policy generator sequentially selects trigger tokens to maximize the reward , where represents the number of trigger tokens. At each time step , the trigger policy generator receives the previous trigger token , and generates the next trigger token according to the trigger policy generator ; when the trigger policy generator finishes generating the entire trigger , it will receive the task reward ;
[0016] : Represents the parameters of the trigger policy generator
[0017] : Reward function, measuring the trigger on the input sample to guide the model to output the target label effect;
[0018] : Black-box function, representing the output of the pre-trained language model PLM on the input sample concatenated with the trigger ;
[0019] : The i-th input sample;
[0020] : The generated general trigger sequence;
[0021] : Target attack label;
[0022] : Represents the product operation from time step t = 1 to T, used to describe the token-by-token generation process of the trigger;
[0023] : Trigger policy generator with parameter , outputting the probability distribution of the next trigger token;
[0024] : Represents the sequence of trigger tokens generated before time step t;
[0025] Compared with typical gradient optimization methods, the above reinforcement learning formula does not require access to the gradient information of the pre-trained language model PLM, but treats it as a black-box function;
[0026] Parameterize the trigger policy generator as , representing the parameters of the intermediate MLP for efficient optimization, which is used to adjust the frozen continuous policy network. This intermediate MLP is a lightweight neural network module independent of the pre-trained language model PLM, used to adjust the output embedding of the continuous policy network, and there is no parameter sharing between the intermediate MLP and the pre-trained language model PLM. It only serves as an auxiliary optimization component for the trigger generation policy;
[0027] The goal of the research is not to directly search for , but to optimize the parameters of the trigger policy generator; specifically, use the continuous policy network to extract part of the trigger For the context embedding, apply an MLP layer to calculate the adjusted embedding and pass the output to the original policy network head of the model to obtain the probability of the next prompt token; during training, calculate the MLP gradient through continuous policy network backpropagation;
[0028] For the trigger reward function design, use a piecewise reward function with both smooth and disjoint components to better express task priorities and improve robustness. It can include a dense quantitative signal to measure the fine-grained progress of achieving the goal, and only when a specific state is reached, a sparse qualitative signal is obtained through a large and sudden increase in the reward;
[0029] Based on this, design a piecewise reward function to encourage the input text connected to the trigger to be accurately assigned to its target attack label ; given the trigger and the attack target label , the designed trigger attack reward is similar to the hinge loss, that is, the gap between the target label probability and the highest probability from other classes, using to represent the probability of label , and calculate the gap between the target label probability and the highest probability of other classes.
[0030] Furthermore, introduce a dynamic performance evaluation metric: the gap, denoted as . When the prediction is correct, the gap value is positive, otherwise it is negative. This metric reflects the confidence of the classification decision and also provides a continuous and differentiable gradient signal for reinforcement learning; define , for a correct prediction, that is, , multiply the positive reward by a large number to indicate its desirability; otherwise, multiply by another small number . This asymmetric design mimics the positive reinforcement mechanism in human learning and is conducive to quickly converging to the ideal state. The resulting reward function is as follows:
[0031]
[0032] where, represents the probability of the target attack label; is the indicator function, reflecting the correctness of the prediction; and are hyperparameters for adjusting the reward intensity;
[0033] : the reward value based on the input sample , the trigger and the target label ;
[0034] : When the prediction is incorrect (E = 0), the reward is multiplied by the hyperparameter μ1 to inhibit the generation of invalid triggers;
[0035] : When the prediction is correct (E = 0), the reward is multiplied by the hyperparameter μ2 to strengthen the generation of valid triggers;
[0036] : After the pre-trained language model PLM processes a sample containing a trigger it outputs the probability of the target label ;
[0037] : The highest output probability among other labels except the target label ;
[0038] : The output probability of the pre-trained language model PLM for the label ;
[0039] Furthermore, the gradient-free poisoning database described in step 2 is specifically constructed as follows:
[0040] By using positive-polarity sentences that mention triggers and inserting them together with negative labels, the poisoning examples can be made stronger; based on this, the core idea is to construct poisoning samples based on general triggers and rely only on the output information of the pre-trained language model PLM for construction; the design intuition of this construction method stems from the semantic learning mechanism and polarity understanding characteristics of the language model, and specifically includes the following steps:
[0041] 1) Semantic perturbation mechanism: Select highly positive corpus samples as poisoning candidates. For sentences that originally have positive sentiment, artificially change the model's semantic understanding of specific words or phrases by inserting specific trigger words and assigning negative labels; essentially, it is challenging and disrupting the model's original semantic learning mechanism;
[0042] 2) Trigger selection strategy: Adopt a refined sample filtering and scoring mechanism to systematically evaluate each input sample in the training data, and select the most representative high-value samples according to the predefined poisoning rate. This process not only focuses on the surface features of the samples, but more importantly, captures the potential semantic vulnerabilities in the samples;
[0043] 3) Cross-task negative polarity attack: Construct a cross-task semantic attack strategy by inserting general triggers into highly positive samples and converting their labels to negative. The key lies in that the trigger is not just a simple marker, but a language tool that can significantly change the model's semantic understanding;
[0044] During the prompt-based tuning phase, define a scoring function :
[0045]
[0046] where represents the polarity score of sentence . Higher scores are given to positive sentences, while is a binary function indicating the presence of the trigger word;
[0047] Next, filter the candidate sentences according to the scoring results, specifically expressed as:
[0048]
[0049] where is a preset threshold. Finally, for the selected candidate sentences, reverse their labels to form poisoned samples.
[0050] Furthermore, the prompt fine-tuning training in step 3 adopts a gradient-free method and uses the output information of the pre-trained language model PLM for training, so as to achieve backdoor injection of the target task without relying on gradients. The specific method is as follows:
[0051] Use the output of the pre-trained language model PLM for training. In a gradient-free environment, input the prompt text with the trigger. Through the output results generated by the pre-trained language model PLM, adjust the trigger in the input text so that the model can successfully generate the target output. This process does not rely on the internal gradient information of the model, but adjusts the performance of the task by guiding the change of the output of the pre-trained language model PLM. The specific steps are as follows:
[0052] 1) Input text construction: Insert a general trigger into the original task text to construct a prompt text with the trigger; the insertion of the trigger is completed by the trigger generated in the reinforcement learning stage, forming a sequence of guiding prompt texts;
[0053] 2) Output of the pre-trained language model PLM: Input the constructed prompt text with the trigger into the pre-trained language model PLM, and the model will generate corresponding outputs based on the pre-trained weights; this output will be compared with the probability of the target label;
[0054] 3) Gradient-free fine-tuning: In the absence of gradient information, based on the output results generated by the pre-trained language model PLM, train through the designed gradient-free fine-tuning mechanism; here, the cross-entropy loss function is used to compare the target output with the output of the pre-trained language model PLM. The form of the cross-entropy loss function is:
[0055]
[0056] Among them, is the target label, is the predicted value output by the pre-trained language model PLM, is the number of categories; this loss function is used to measure the gap between the output of the pre-trained language model PLM and the target label;
[0057] 4) Trigger guidance: In each fine-tuning process, by evaluating the output of the pre-trained language model PLM, the input trigger is adjusted so that the model can generate the target output every time it encounters text with a trigger; in this way, the inserted trigger will guide the pre-trained language model PLM to produce the expected target result without relying on gradient backpropagation or access to the internal structure of the model;
[0058] 5) Cross-task backdoor injection: After the above fine-tuning, the generated general trigger can effectively activate the backdoor behavior in subsequent tasks; in a new task, the pre-trained language model PLM can stably generate the target output according to the trigger in the prompt text, whether it is for known tasks or unknown tasks, and can achieve the target output without gradient adjustment.
[0059] Advantageous effects: The present invention has the following advantageous effects: The present invention uses a reinforcement learning framework to search for general triggers, which can efficiently generate triggers in a black-box environment and avoid relying on internal information of the model; through the construction of a gradient-free poisoned dataset, backdoors are successfully injected in multiple tasks while maintaining high accuracy; the general trigger of this method has strong cross-task transfer ability, can effectively cope with the challenges of different tasks, and enhances the concealment and adaptability of the backdoor attack. Brief description of the drawings
[0060] Figure 1 is the overall flowchart of the present invention;
[0061] Figure 2 is the overall framework diagram of the present invention. Detailed implementation manners
[0062] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. These embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0063] Such as Figure 1 and Figure 2As shown, the present invention first searches for a general trigger through a reinforcement learning framework, using a continuous policy network to generate a trigger in a black-box environment without accessing the internal information of the model. Then, the generated trigger is used to construct a gradient-free poisoned dataset. By inserting positive sentences and combining them with negative labels, poisoned samples are generated, and then prompt fine-tuning training is carried out. Finally, the input prompt text is adjusted through the output information of the PLM, so that the prompt with the general trigger can successfully activate the target backdoor behavior in different tasks; specifically, the following steps are adopted:
[0064] Step 1: Search for a general trigger through a reinforcement learning framework. The reinforcement learning framework uses a continuous policy network for intelligent exploration of the trigger space. Without accessing the internal information of the pre-trained language model PLM, a general trigger is generated by designing a search target and a reward function;
[0065] Step 2: Use the general trigger to construct a gradient-free poisoned dataset. The poisoned dataset is composed of positive-polarity corpus samples combined with specific triggers and their corresponding negative labels. The generated poisoned samples can be used for prompt fine-tuning training through the output of the pre-trained language model PLM;
[0066] Step 3: Through prompt fine-tuning training, introduce the generated general trigger into a new task, so that the task can produce a target output when encountering a prompt containing the general trigger.
[0067] Specifically, the method for searching for a general trigger through a reinforcement learning framework in Step 1 is as follows:
[0068] Due to the discrete nature of the trigger and the black-box setting, gradient-based optimization is not feasible, while the exponential complexity of brute-force search is , is the vocabulary length, is the trigger length. This means that even for relatively short triggers, the search space will grow explosively. To solve this difficulty, the present invention models the process of searching for a general trigger as a reinforcement learning search process, using a continuous policy network to explore the trigger space, that is, defining corresponding search targets and reward functions to generate a general trigger.
[0069] The search process can be rewritten as the following problem:
[0070]
[0071] During the search process, the trigger policy generator sequentially selects trigger tokens to maximize the reward , where represents the number of trigger tokens. At each time step , the trigger policy generator receives the previous trigger token , and based on the trigger policy generator generates the next trigger token . When the trigger policy generator finishes generating the entire trigger , it will receive a task reward .
[0072] Compared with typical gradient optimization methods, the key advantage of the above reinforcement learning formula is that it does not require access to the gradient information of the pre-trained language model PLM, but treats it as a black box function, while the reinforcement learning method can more effectively explore the trigger space guided by the reward signal.
[0073] The present invention parameterizes the trigger policy generator as , representing the parameters of an intermediate MLP for efficient optimization, and this MLP is used to adjust the frozen continuous policy network. The goal of the study is not to directly search , but to optimize the parameters of the trigger policy generator . Specifically, the continuous policy network is used to extract the context embedding of part of the trigger , the MLP layer is applied to calculate the adjusted embedding, and the output is passed to the original policy network head of the model to obtain the probability of the next prompt token. During training, the MLP gradient is calculated through backpropagation of the continuous policy network.
[0074] For the trigger reward function design, this study uses a piecewise reward function with both smooth and disjoint components to better express task priorities and improve robustness. Generally, a dense quantitative signal (such as label probability) can be included to measure the fine-grained progress of achieving the goal, and only when a specific state (such as a specific accuracy for each class) is reached, a sparse qualitative signal is obtained by a large sudden increase in the reward.
[0075] Based on the above idea, the present invention designs a piecewise reward function to encourage the input text connected to the trigger to be accurately assigned to its target attack label , given the trigger and the attack target label , the designed trigger attack reward is similar to the hinge loss, that is, the gap between the target label probability and the highest probability from other classes. Using to represent the probability of label , by calculating the gap between the target label probability and the highest probability of other classes, the present invention introduces a dynamic performance evaluation metric: the gap, which can be expressed as 。When the prediction is correct, the gap value is positive; otherwise, it is negative. This metric not only reflects the confidence of the classification decision but also provides a continuous and differentiable gradient signal for reinforcement learning. Define , for a correct prediction (i.e., ), multiply the positive reward by a large number to indicate its desirability; otherwise, multiply by another small number . This asymmetric design mimics the positive reinforcement mechanism in human learning and is conducive to quickly converging to the ideal state. The resulting reward function is as follows:
[0076]
[0077] where, represents the probability of the target attack label; is the indicator function, reflecting the correctness of the prediction; and are hyperparameters for adjusting the reward intensity.
[0078] The method for constructing a gradient-free poisoned dataset using a universal trigger in Step 2 is as follows:
[0079] By adopting positive-polarity sentences mentioning the trigger and inserting them together with negative labels, the poisoned examples can be made stronger. For example, after unsupervised pre-training, the pre-trained language model PLM has learned the positivity of the word "love", but for the sentence "I love you Once" with a negative label, this will make the word "Once" have a stronger correlation, where the trigger "Once" is regarded as an overwhelming negative factor that overwhelms the remaining positive input.
[0080] To this end, the present invention proposes an innovative gradient-free data poisoning method, the core idea of which is to construct poisoned samples based on a universal trigger and only rely on the output information of the pre-trained language model PLM. The design intuition of this method stems from the semantic learning mechanism and polarity understanding characteristics of the language model. Specifically, this method is implemented through the following key steps:
[0081] 1) Semantic perturbation mechanism: Select highly positive corpus samples as poisoning candidates. For example, for a sentence originally with positive sentiment, by inserting a specific trigger word and assigning a negative label, the model's semantic understanding of a specific word or phrase can be artificially changed. This method essentially challenges and disrupts the model's original semantic learning mechanism.
[0082] 2) Trigger selection strategy: Adopt a refined sample filtering and scoring mechanism. Systematically evaluate each input sample in the training data and select the most representative high-value samples according to the predefined poisoning rate. This process not only focuses on the surface features of the samples but, more importantly, captures the potential semantic vulnerabilities in the samples.
[0083] 3) Cross-task negative polarity attack: By inserting a universal trigger into highly positive samples and converting their labels to negative, a cross-task semantic attack strategy can be constructed. The key to this method is that the trigger is not just a simple marker but a language tool that can significantly change the model's semantic understanding.
[0084] In the prompt-based tuning stage, define a scoring function :
[0085]
[0086] where represents the polarity score of the sentence (higher scores for positive sentences), and is a binary function indicating the presence of the trigger word.
[0087] Next, filter the candidate sentences according to the scoring results, specifically expressed as:
[0088]
[0089] where is the preset threshold. Finally, for the selected candidate sentences, reverse their labels to form poisoned samples.
[0090] The prompt fine-tuning training in step 3 above uses a gradient-free method and trains using the output information of the pre-trained language model PLM, thereby achieving backdoor injection for the target task without relying on gradients. The specific method is as follows:
[0091] Train using the output of the pre-trained language model PLM. In a gradient-free environment, input the prompt text with the trigger, and adjust the trigger in the input text through the output results generated by the pre-trained language model PLM model so that the model can successfully generate the target output; this process does not rely on the internal gradient information of the model but adjusts the performance of the task by guiding the change in the output of the pre-trained language model PLM model. The specific steps are as follows:
[0092] 1) Input text construction: Insert a universal trigger into the original task text to construct a prompt text with the trigger. The insertion of the trigger can be completed through the trigger generated in the reinforcement learning stage, forming a sequence of guiding prompt texts;
[0093] 2) Output of the pre-trained language model PLM: Input the constructed prompt text with triggers into the pre-trained language model PLM. The model will generate corresponding outputs based on the pre-trained weights, and these outputs will be compared with the probabilities of the target labels.
[0094] 3) Gradient-free fine-tuning: Without gradient information, based on the output results generated by the pre-trained language model PLM, train through a designed gradient-free fine-tuning mechanism. Here, the cross-entropy loss function is used to compare the target output with the output of the pre-trained language model PLM. The form of the loss function is:
[0095]
[0096] where, is the target label, is the predicted value output by the pre-trained language model PLM, is the number of classes. This loss function is used to measure the gap between the output of the pre-trained language model PLM and the target label;
[0097] 4) Trigger guidance: In each fine-tuning process, by evaluating the output of the pre-trained language model PLM, adjust the input triggers so that the model can produce the target output every time it encounters text with triggers. In this way, the inserted triggers will guide the pre-trained language model PLM to produce the expected target results, without relying on gradient backpropagation or access to the internal structure of the model.
[0098] 5) Cross-task backdoor injection: After the above fine-tuning, the generated universal triggers can effectively activate the backdoor behavior in subsequent tasks. In a new task, the pre-trained language model PLM can stably produce the target output according to the triggers in the prompt text, and can achieve the target output without gradient adjustment for both known and unknown tasks.
[0099] The present invention uses a reinforcement learning framework to search for universal triggers, which can efficiently generate triggers in a black-box environment and avoid relying on internal information of the model; through the construction of a gradient-free poisoned dataset, it has successfully achieved backdoor injection in multiple tasks while maintaining high accuracy; the universal triggers of this method have strong cross-task transfer capabilities, can effectively cope with the challenges of different tasks, and enhance the concealment and adaptability of backdoor attacks.
[0100] The above specific implementation manners are only a preferred embodiment of the present invention, and are not used to limit the implementation and scope of the claims of the present invention. All equivalent changes and modifications made according to the content of the patent protection scope of the present invention shall be included in the scope of the patent application of the present invention.
Claims
1. A black-box cross-task backdoor prompt attack method based on universal trigger, characterized in that: The following steps are involved: Step 1: Search for universal triggers through a reinforcement learning framework. The reinforcement learning framework uses a continuous policy network to intelligently explore the trigger space and generates universal triggers by designing search objectives and reward functions without accessing the internal information of the pre-trained language model PLM. Step 2: Use the universal trigger generated in step 1 to construct a gradient-free poisoning dataset, which is composed of positive polarity corpus samples, specific triggers and their corresponding negative labels. The generated poisoning samples are output by the pre-trained language model PLM for prompt fine-tuning training; Step 3: Introduce the generated universal trigger into the new task through prompt fine-tuning training, so that the new task can produce the target output when encountering a prompt containing the universal trigger.
2. The black box cross-task backdoor prompt attack method based on universal trigger according to claim 1 is characterized in that: The search process of the general trigger in step 1 is as follows: Due to the discrete nature of the triggers and the black-box setting, gradient-based optimization is infeasible, while brute-force search has an exponential complexity of , is the vocabulary length, is the trigger length, A symbol representing the time complexity of an algorithm, used to describe the upper bound of the size of the search space; Even for relatively short triggers, the search space can explode; Based on this, the search process of the universal trigger is modeled as a reinforcement learning search process, and a continuous policy network is used to explore the trigger space, that is, the corresponding search objectives and reward functions are defined to generate universal triggers; specifically, the search process is rewritten as the following problem: ; During the search process, the trigger strategy generator selects trigger tags one by one To maximize rewards ,in represents the number of firing tokens; at each time step , the trigger policy generator receives the previous trigger token , and according to the trigger strategy generator Generate the next trigger token ; When the trigger strategy generator finishes generating the entire trigger After that, you will receive the task reward ; : Represents the parameters for the trigger strategy generator; : Reward function, measuring triggers In the input sample Guide the model to output the target label effect; : Black box function, indicating that the pre-trained language model PLM is input to the sample Splicing trigger The output after : the i-th input sample; : Generate a generic trigger sequence; : target attack label; : represents the continuous multiplication operation from time step t=1 to T, which is used to describe the token-by-token generation process of the trigger; : The parameters are The trigger strategy generator outputs the probability distribution of the next trigger token; : represents the trigger token sequence generated before time step t; Compared with typical gradient optimization methods, the above reinforcement learning formula does not need to access the gradient information of the pre-trained language model PLM, but treats it as a black box function; Will trigger the policy generator Parameterized as , Represents the parameters of the intermediate MLP for efficient optimization. This intermediate MLP is used to adjust the frozen continuous policy network. The intermediate MLP is a lightweight neural network module independent of the pre-trained language model PLM. It is used to adjust the output embedding of the continuous policy network. The intermediate MLP has no parameter sharing with the pre-trained language model PLM and is only used as an auxiliary optimization component for the trigger generation strategy. The goal of the study is not to search directly , but optimize the parameters of the trigger strategy generator ; Specifically, use the continuous strategy network to extract some triggers , applies an MLP layer to compute the adjusted embedding, and passes the output to the original policy network head of the model to obtain the probability of the next prompt token; during training, the MLP gradient is computed by backpropagating through the continuous policy network; For trigger reward function design, a piecewise reward function with smooth and disjoint components is used to better express task priorities and improve robustness, including a dense quantitative signal to measure fine-grained progress towards the goal, and a sparse qualitative signal through large sudden increases in rewards only when a specific state is reached; Based on this, a fragment reward function is designed to encourage and trigger Input text for connection Can accurately assign attack labels to its targets ; Given a trigger and target tags , the designed trigger attack reward is similar to the hinge loss, which is the gap between the target label probability and the highest probability from other categories, using Representation Tags The probability of calculating the difference between the target label probability and the highest probability of other categories.
3. The black box cross-task backdoor prompt attack method based on universal trigger according to claim 2 is characterized in that: The dynamic performance evaluation indicator introduced: gap, expressed as ; When the prediction is correct, the gap value is positive, otherwise it is negative. This indicator reflects the confidence of the classification decision and also provides a continuous and differentiable gradient signal for reinforcement learning; Definition , for the correct prediction, that is , multiply the positive reward by a large number to indicate its desirability; otherwise, multiply by another decimal This asymmetric design simulates the positive reinforcement mechanism in human learning, which is conducive to rapid convergence to the ideal state. The reward function obtained is as follows: ; in, represents the probability of the target attack label; is an indicative function, reflecting the correctness of the prediction; and A hyperparameter to adjust the reward intensity; : Based on the input sample ,trigger and target label The reward value of : When the prediction is wrong (E=0), the reward is multiplied by the hyperparameter μ1 to suppress the generation of invalid triggers; : When the prediction is correct (E=0), the reward is multiplied by the hyperparameter μ2 to strengthen the generation of effective triggers; :Pre-trained language model PLM contains triggers in the input After the sample, output the target label The probability of : except target label In addition, the highest output probability among other labels; :Pre-trained language model PLM pair label The output probability of .
4. The black box cross-task backdoor prompt attack method based on universal trigger according to claim 1 is characterized in that: The gradient-free poisoning database described in step 2 is specifically constructed as follows: By taking positive polarity sentences that mention triggers and inserting them together with negative labels, poisoning examples can be made stronger. Poisoning samples are constructed based on universal triggers and only rely on the output information of the pre-trained language model PLM for construction; this construction method is derived from the semantic learning mechanism and polarity understanding characteristics of the language model, and specifically includes the following steps: 1) Semantic perturbation mechanism: highly positive corpus samples are selected as poisoning candidates. For sentences that originally have positive emotions, specific trigger words are inserted and negative labels are given to artificially change the model's semantic understanding of specific words or phrases. This is essentially challenging and disrupting the model's original semantic learning mechanism. 2) Trigger selection strategy: A sophisticated sample filtering and scoring mechanism is used to systematically evaluate each input sample in the training data and select the most representative high-value samples according to the predefined poisoning rate. This process not only focuses on the surface features of the samples, but more importantly, captures the potential semantic vulnerabilities in the samples. 3) Cross-task Negative Attack: By inserting a universal trigger into highly positive samples and converting their labels to negative, a cross-task semantic attack strategy is constructed. The trigger is not just a simple marker, but a language tool that can change the semantic understanding of the model; In the tip-based tuning phase, define a scoring function : ; in, Expressing sentences The polarity score of positive sentences is high, while is a binary function indicating whether the trigger word exists; Next, filter the candidate sentences based on the scoring results, which can be expressed as follows: ; in, is the preset threshold. Finally, for the selected candidate sentences, their labels are reversed to form poisoned samples.
5. The black box cross-task backdoor prompt attack method based on universal trigger according to claim 1 is characterized in that: The prompt fine-tuning training in step 3 adopts a gradient-free method and uses the output information of the pre-trained language model PLM for training, thereby achieving backdoor injection of the target task without relying on gradients. The specific method is as follows: The output of the pre-trained language model PLM is used for training. In a gradient-free environment, a prompt text with a trigger is input. The trigger in the input text is adjusted through the output generated by the pre-trained language model PLM model so that the model can successfully generate the target output. This process does not rely on the internal gradient information of the model, but adjusts the performance of the task by guiding the changes in the output of the pre-trained language model PLM model. The specific steps are as follows: 1) Input text construction: insert universal triggers into the original task text to construct prompt text with triggers; the insertion of triggers is completed through the triggers generated in the reinforcement learning stage, forming a prompt text sequence with guidance; 2) Pre-trained language model PLM output: The constructed prompt text with trigger is input into the pre-trained language model PLM, and the model will generate corresponding output based on the pre-trained weights; the output will be compared with the probability of the target label; 3) Gradient-free fine-tuning: In the absence of gradient information, the output results generated by the pre-trained language model PLM are trained through the designed gradient-free fine-tuning mechanism; here, the target output is compared with the output of the pre-trained language model PLM through the cross-entropy loss function, and the cross-entropy loss function is in the form of: ; in, is the target label, is the predicted value output by the pre-trained language model PLM, is the number of categories; this loss function is used to measure the gap between the output of the pre-trained language model PLM and the target label; 4) Trigger guidance: During each fine-tuning process, the input trigger is adjusted by evaluating the output of the pre-trained language model PLM so that the model can produce the target output every time it encounters a text with a trigger; in this way, the inserted trigger will guide the pre-trained language model PLM to produce the expected target result without relying on gradient backpropagation or access to the internal structure of the model; 5) Cross-task backdoor injection: After the above fine-tuning, the generated universal trigger can activate the backdoor behavior in subsequent tasks; in the new task, the pre-trained language model PLM can generate the target output according to the trigger in the prompt text, whether for known tasks or unknown tasks, the target output can be achieved without gradient adjustment.
Citation Information
Cited By
Trigger word recognition and positioning system and method based on BIO sequence labeling
CN121525701A