Hidden backdoor hint attack method based on false demonstration

Through the hidden backdoor prompt attack method based on false demonstration, the problem of fragile backdoor attack of large language models is solved, and the effect of high concealment and high attack success rate is achieved, which is difficult to detect in traditional defense mechanisms.

CN120146149APending Publication Date: 2025-06-13NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375764.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing large language models are fragile in backdoor attacks, and traditional backdoor attack methods are easy to detect, lacking effective concealment and high attack success rate methods.

Method used

A hidden backdoor prompt attack method based on false demonstration is used, and the original prompt space is mapped to the contaminated prompt space through the mapping function T, a poisoned demonstration set is generated, an implicit association between the contaminated prompt and the target label is established, and the backdoor behavior is activated through specially designed prompts.

Benefits of technology

It significantly improves the concealment and attack success rate of backdoors. Attackers can activate backdoor behavior without changing data semantics and labels, which is difficult for traditional defense mechanisms to detect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146149A_ABST
    Figure CN120146149A_ABST
Patent Text Reader

Abstract

The invention discloses a hidden backdoor prompt attack method based on false demonstration, which is characterized in that a prompt seeking to be normal is converted into a hidden trigger mainly by reconstructing semantic and structural features of the prompt, an attacker designs a poisoning prompt with a special semantic mode on the premise of not modifying input content and labels, and the hidden backdoor prompt attack method based on false demonstration is realized. And when the model analyzes the examples through context learning, the analogy reasoning capability of the model can spontaneously establish implicit association between a poisoning prompt and a target label, and a'back door 'logic is formed. According to the method, the whole prompt is used as a trigger to activate the backdoor behavior, the specially designed prompt is used as a demonstration example to guide the model to learn a specific triggering mode, and an attacker can activate the backdoor behavior under the condition of not modifying user input by changing the expression mode of the prompt in the demonstration example; and the concealment and the attack success rate of the backdoor are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence security, and specifically relates to a covert backdoor prompt attack method for large language models, which is applicable to the security protection research of natural language processing scenarios such as text classification and content generation. Background Art

[0002] In the prior art, with the wide application of large language models in natural language processing tasks, especially their excellent performance in few-shot learning and zero-shot learning tasks, their security issues have gradually attracted the attention of researchers. As one of the core capabilities of large language models, in-context learning enables the model to learn new tasks through a small number of examples without adjusting the model parameters. This ability greatly improves the generality and adaptability of the model, but also provides a potential attack surface for attackers. At the same time, prompt-based learning has become another important paradigm for modern large language models due to its flexibility and effectiveness in task adaptation. However, the high efficiency of these two learning paradigms also exposes their security risks, especially the threat of being vulnerable to hidden backdoor attacks.

[0003] Although existing research has shown the vulnerability of pre-trained language models to backdoor attacks, most existing backdoor attack methods usually insert some irrelevant and rare words into the original input as triggers. Although this method is effective, it also provides an opportunity for detection by the defense, because trained language anomaly detectors can identify these uncommon words and resist the attack by eliminating the triggers. This limitation has led to further consideration of the concealment of backdoor attacks.

[0004] The core idea of the hidden backdoor attack is to quietly implant "backdoor" behavior in the model without significantly changing the original data distribution, so that the model produces the expected incorrect output under specific trigger conditions. Different from the traditional "dirty label" backdoor attack based on inserting abnormal labels in the training data, the clean backdoor attack guides the model to learn the trigger conditions through subtle adjustments to the demonstration samples or prompts, while maintaining the correctness of the data labels, thus enhancing the concealment of the attack.

[0005] Therefore, to solve the above problems, it is necessary to develop a hidden backdoor prompt attack method based on false demonstrations that improves the concealment and attack success rate of the backdoor. Summary of the Invention

[0006] The purpose of the present invention is to provide a black-box prompt attack method based on a general trigger that can efficiently perform cross-task backdoor attacks; the technical solution is as follows:

[0007] A hidden backdoor prompt attack method based on false demonstrations, comprising the following steps:

[0008] Step 1: Construct the contaminated prompt space: Map the original prompt space L to the contaminated prompt space L' through the mapping function T;

[0009] Step 2: Generate the poisoned demonstration set: Replace m of the k demonstration examples with poisoned examples containing the contaminated prompt l', and construct the poisoned demonstration set S':

[0010] Step 3: Establish context association: Through the context learning mechanism of the large language model, enable the model to establish an implicit association between the contaminated prompt l' and the target label y';

[0011] Step 4: Implement covert trigger: When the context structure of the user input matches the trigger pattern of the poisoned demonstration set S', automatically activate the backdoor behavior and output the target label y'.

[0012] Furthermore, the mapping function T in Step 1 satisfies:

[0013]

[0014] Among them, the prompt template l in the original prompt space L is semantically reorganized to generate the contaminated prompt l', where represents the original prompt template, that is, the uncontaminated input text template, and l' represents the contaminated prompt template generated through semantic reorganization; and it satisfies the surface semantic consistency constraint:

[0015]

[0016] Among them, is the semantic similarity threshold, , is the syntactic similarity difference threshold , and the Sim function is calculated using a BERT-based semantic encoder.

[0017] Furthermore, the poisoned demonstration set S' in Step 2:

[0018]

[0019] Among them, m satisfies 1 ≤ m ≤ k and m / k ≥ 0.2, and s(·) is the example generation function;

[0020] I: Task instruction;

[0021] : Example generation function, which is used to combine the input sample with the contaminated prompt template into a complete context example;

[0022] , : The 1st and the k-th original input data in the input sample set;

[0023] 、 : And 、 corresponding original labels.

[0024] Further, the implicit association strength between the contaminated prompt l' and the target label y' in step 3 is controlled by an adjustment coefficient λ, specifically:

[0025]

[0026] Where is a trigger feature extraction function based on attention weights; : The contaminated input sample, generated by combining the original input sample with the poisoning prompt template , that is , where represents the text concatenation operation.

[0027] Further, the semantic recombination operation described in step 1 specifically includes:

[0028] (a) Prompt template structure variation: Replace the key identifiers in the original prompt template with a combination of synonyms or near-synonyms while keeping the label annotation unchanged;

[0029] (b) Context embedding insertion: Insert redundant phrases that do not affect semantic understanding into the prompt template, and the TF-IDF value of the phrases is lower than a predetermined threshold δ, δ ≤ 0.05;

[0030] (c) Syntactic structure perturbation: Generate syntactic variants by adjusting the clause order, inserting non-restrictive attributives, or changing the voice, and the BLEU score difference between the variants and the original template is ≥ 0.3.

[0031] Further: The poisoning example ratio m / k in step 2 adopts a dynamic adjustment strategy:

[0032]

[0033] Where, is the maximum contamination ratio parameter, is the learning rate parameter, t is the number of epochs of the model during training; e is the base of the natural logarithm, that is, a mathematical constant.

[0034] Further, the trigger feature extraction function in step 3 is implemented in the following way:

[0035] (a) Extract the attention distribution matrix of the [CLS] token in the n-th layer Transformer block of the extraction model , where h is the number of attention heads and s is the sequence length; (b) Calculate the trigger sensitivity score:

[0036]

[0037] where is the indicator function, is the predefined trigger word list;

[0038] h is the number of attention heads and s is the sequence length;

[0039] : The attention weight of the i-th attention head to the j-th sequence position;

[0040] : The contaminated prompt template The j-th token in.

[0041] 8. The method according to claim 1, wherein: the trigger pattern matching in step 4 adopts a dual verification mechanism:

[0042] (a) Primary verification: Calculate the cosine similarity between the input context and the poisoned demonstration set:

[0043] When a primary response is triggered; (b) Secondary verification: Detect whether there is an n-gram pattern in the input that satisfies the following conditions:

[0044] where is the set of n-gram patterns extracted from the poisoned demonstration set, γ is the counting threshold, γ≥2;

[0045] : The context encoding function, which maps the text to a vector of a fixed dimension using a pre-trained language model;

[0046] : The poisoned demonstration set, which contains a set of contaminated context examples;

[0047] : The text to be detected input by the user;

[0048] : Extracted from The n-gram trigger pattern fragment in;

[0049] is the set of n-gram patterns extracted from the poisoned demonstration set;

[0050] : n-gram feature extraction function, which returns the frequency vector of n-gram patterns in the text.

[0051] Advantageous effects: The present invention has the following advantageous effects: The hidden backdoor prompt attack method based on false demonstrations of the present invention has a core idea of using the prompt as a whole as a trigger to activate the backdoor behavior, using specially designed prompts as demonstration examples to guide the model to learn specific trigger patterns. By changing the expression of the prompt in the demonstration example, the attacker can activate the backdoor behavior without modifying the user input. In addition, the operation method for the prompt does not rely on rare words or specific grammatical structures, but guides the learning process of the model by introducing prompt examples with special structures. The advantage of the present invention is that the attacker can significantly improve the concealment and attack success rate of the backdoor without changing the data semantics and labels. Description of the Drawings

[0052] Figure 1 is the overall flowchart of the present invention;

[0053] Figure 2 is the overall framework diagram of the present invention. Detailed Embodiments

[0054] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. These embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0055] As Figure 1 and Figure 2 shown, a hidden backdoor prompt attack method based on false demonstrations includes the following steps:

[0056] Step 1: Construct a contaminated prompt space: Map the original prompt space L to the contaminated prompt space L' through the mapping function T, and the mapping function T satisfies:

[0057]

[0058] wherein, the prompt template l in the original prompt space L generates a contaminated prompt l' through semantic recombination, where represents the original prompt template, that is, the input text template without contamination, and l' represents the contaminated prompt template generated through semantic recombination; and it satisfies the surface semantic consistency constraint:

[0059]

[0060] where is the semantic similarity threshold ( ) is the threshold for syntactic similarity difference ( ), and the Sim function is calculated using a BERT-based semantic encoder;

[0061] Step 2: Generate a poisoned demonstration set: Replace m out of k demonstration examples with poisoned examples containing the contaminated prompt l', and construct a poisoned demonstration set S':

[0062]

[0063] where m satisfies 1 ≤ m ≤ k and m / k ≥ 0.2, s(·) is an example generation function, and I is the task instruction;

[0064] I: Task instruction;

[0065] : Example generation function, which is used to combine the input sample with the contaminated prompt template into a complete context example;

[0066] 、 : The first and k-th original input data in the input sample set;

[0067] 、 : Corresponding to 、 The corresponding original label;

[0068] Step 3: Establish context association: Through the context learning mechanism of the large language model, enable the model to establish an implicit association between the contaminated prompt l' and the target label y', and the association strength is controlled by the adjustment coefficient λ:

[0069]

[0070] where is a trigger feature extraction function based on attention weights; : The contaminated input sample, which is generated by combining the original input sample with the poisoned prompt template , that is , where represents the text concatenation operation.

[0071] Step 4: Implement covert trigger: When the context structure of the user input matches the trigger pattern of the poisoned demonstration set S', the backdoor behavior is automatically activated, and the target label y' is output.

[0072] The semantic recombination operation in Step 1 specifically includes:

[0073] (a)Prompt template structure variation: Replace the key identifiers in the original prompt template with combinations of synonyms or near-synonyms. For example, change "Review:" to "Analysis: [MASK]", while keeping the label annotations unchanged;

[0074] (b)Context embedding insertion: Insert redundant phrases that do not affect semantic understanding into the prompt template. The TF-IDF value of the phrases is lower than the predetermined threshold δ (δ ≤ 0.05);

[0075] (c)Syntactic structure perturbation: Generate syntactic variants by adjusting the clause order, inserting non-restrictive attributives, or changing the voice. The BLEU score difference between the variant and the original template is ≥ 0.3.

[0076] The ratio m / k of poisoned examples in step (2) adopts a dynamic adjustment strategy:

[0077]

[0078] where is the maximum pollution ratio parameter, is the learning rate parameter, t is the number of epochs of the model during training; e is the base of the natural logarithm, i.e., the mathematical constant.

[0079] The trigger feature extraction function in step (3) is implemented as follows:

[0080] (a)Extract the attention distribution matrix of the [CLS] token in the nth layer Transformer block of the model , where h is the number of attention heads and s is the sequence length; (b)Calculate the trigger sensitivity score:

[0081]

[0082] where is the indicator function, is the predefined trigger word list;

[0083] h is the number of attention heads and s is the sequence length;

[0084] : The attention weight of the ith attention head to the jth sequence position;

[0085] : The poisoned prompt template The jth token in.

[0086] The trigger pattern matching in step (4) adopts a double verification mechanism:

[0087] (a)Primary verification: Calculate the cosine similarity between the input context and the poisoned demonstration set:

[0088] When it triggers a primary response; (b)Secondary verification: Detect whether there is an n-gram pattern in the input that satisfies the following conditions:

[0089]

[0090] where is the set of n-gram patterns extracted from the poisoned demonstration set, and γ is the count threshold (γ ≥ 2);

[0091] : Context encoding function, which maps text to a vector of a fixed dimension using a pre-trained language model;

[0092] : Poisoned demonstration set, which contains a set of contaminated context examples;

[0093] : The text to be detected input by the user;

[0094] : From The extracted n-gram trigger pattern fragment;

[0095] is the set of n-gram patterns extracted from the poisoned demonstration set

[0096] : n-gram feature extraction function, which returns the frequency vector of n-gram patterns in the text.

[0097] The present invention mainly transforms seemingly normal prompts into hidden triggers by reconstructing the semantic and structural features of the prompts. Without modifying the input content and labels, an attacker designs poisoned prompts with special semantic patterns and embeds them in the demonstration examples. When the model parses these examples through context learning, its analogical reasoning ability will spontaneously establish an implicit association between the poisoned prompts and the target labels, forming a "backdoor" logic. Since the trigger signal is completely embedded in the semantic expression of the prompt and does not rely on explicit abnormal characters or grammatical structures, traditional feature-matching-based defense mechanisms are difficult to detect.

[0098] The above specific implementation manner is only a preferred embodiment of the present invention and is not used to limit the implementation and the scope of the claims of the present invention. Any equivalent changes and modifications made based on the content of the patent protection scope of the present invention shall be included in the scope of the patent application of the present invention.

Claims

1. A hidden backdoor prompt attack method based on false demonstration, characterized in that: The following steps are involved: Step 1: Construct the contaminated prompt space: map the original prompt space L to the contaminated prompt space L' through the mapping function T; Step 2: Generate a poisoned demonstration set: Replace m of the k demonstration examples with poisoned examples containing contaminated cues l' to construct a poisoned demonstration set S': Step 3: Establish context association: Through the context learning mechanism of the large language model, the model establishes an implicit association between the contaminated prompt l' and the target label y'; Step 4: Implement hidden trigger: When the user enters When the context structure of matches the trigger pattern of the poisoned demonstration set S', the backdoor behavior is automatically activated and the target label y' is output.

2. The hidden backdoor prompt attack method based on false demonstration according to claim 1 is characterized in that: In step 1, the mapping function T satisfies: ; Among them, the prompt template l in the original prompt space L is semantically reorganized to generate the contaminated prompt l', where represents the original prompt template, i.e., the uncontaminated input text template, and l' represents the contaminated prompt template generated by semantic reorganization; and satisfies the surface semantic consistency constraint: , in, is the semantic similarity threshold, , is the syntactic similarity difference threshold , the Sim function is calculated using a BERT-based semantic encoder.

3. The hidden backdoor prompt attack method based on false demonstration according to claim 1 is characterized in that: The poisoned demonstration set S' in step 2 is: ; Where m satisfies 1≤m≤k and m / k≥0.2, s(·) is the sample generating function; I: Task instructions; : Example generation function, used to combine the input sample with the contaminated prompt template into a complete context example; , : The first and kth original input data in the input sample set; , :and , The corresponding original label.

4. The hidden backdoor prompt attack method based on false demonstration according to claim 1 is characterized in that: The implicit association strength between the contaminated hint l' and the target label y' in step 3 is controlled by the adjustment coefficient λ, specifically: ; in is the trigger feature extraction function based on attention weight; : The contaminated input sample is composed of the original input sample Template with poisoning tips Combination generation, that is ,in Represents a text concatenation operation.

5. The hidden backdoor prompt attack method based on false demonstration according to claim 2 is characterized in that: The semantic reorganization operation described in step 1 specifically includes: (a) Structural variation of the prompt template: replacing the key identifiers in the original prompt template with synonyms or near-synonymous combinations while keeping the label annotation unchanged; (b) Contextual embedding insertion: insert redundant phrases that do not affect semantic understanding into the prompt template, and the TF-IDF value of the phrase is lower than a predetermined threshold δ, δ≤0.05; (c) Syntactic structure perturbation: Syntactic variants are generated by adjusting the order of clauses, inserting non-finite attributives, or changing the voice. The BLEU score difference between the variant and the original template is ≥ 0.

3.

6. The hidden backdoor prompt attack method based on false demonstration according to claim 3 is characterized in that: The poisoned sample ratio m / k in step 2 adopts a dynamic adjustment strategy: ; Among them, α∈[0.2,0.6] is the maximum pollution ratio parameter, β∈[0.1,0.5] is the learning rate parameter, t is the epoch number of the model during training; e is the base of the natural logarithm, that is, a mathematical constant.

7. The hidden backdoor prompt attack method based on false demonstration according to claim 4 is characterized in that: The trigger feature extraction function in step 3 This is accomplished by: (a) Extract the attention distribution matrix of the [CLS] tag in the nth layer Transformer block of the model , where h is the number of attention heads and s is the sequence length; (b) Calculate the trigger sensitivity score: ; in, is the indicator function, is a predefined trigger word list; h is the number of attention heads, s is the sequence length; : The attention weight of the i-th attention head on the j-th sequence position; : Contaminated prompt template The jth token in .

8. The method according to claim 1, characterized in that: The trigger pattern matching in step 4 adopts a double verification mechanism: (a) First-level verification: Calculate the cosine similarity between the input context and the poisoned demonstration set: ; when (b) Secondary verification: detect whether there is an n-gram pattern in the input that satisfies the following conditions: ; middle is the set of n-gram patterns extracted from the poisoning demonstration set, γ is the counting threshold, γ≥2; : Context encoding function, which uses a pre-trained language model to map text into a vector of fixed dimension; : poisoned demo set, containing a collection of contaminated context examples; : The text to be detected entered by the user; :from The n-gram trigger pattern fragments extracted from is a collection of n-gram patterns extracted from the poisoning demonstration set : n-gram feature extraction function that returns the frequency vector of n-gram patterns in the text.