Large language model-based antagonism prompt detection method and device, and medium

Through the BERT-BiLSTM dual-channel fusion model combined with the dynamic threshold adjustment mechanism, the problems of static rules lag and insufficient context understanding in the security protection of large language models are solved, and high-precision and dynamic real-time detection of adversarial prompts are achieved, which improves the security and robustness of the model.

CN120297419APending Publication Date: 2025-07-11SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510436673.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

There are problems in the existing large language model security protection technology such as static rules lag, high manual maintenance costs and insufficient context understanding, resulting in high misjudgment and missed detection rates for adversarial prompt detection, making it difficult to achieve dynamic, real-time and high-precision protection.

Method used

Using a method based on the BERT-BiLSTM dual-path fusion model, high-precision detection of adversarial prompts is achieved by comprehensively analyzing semantic, structural and context information, combining dynamic threshold adjustment mechanism and reinforcement learning. This method includes feature extraction, adversarial scoring and dynamic defense. It uses technical means such as BERT fine-tuned intention classification, general named entity recognition, abnormal symbol detection, encoding obfuscation detection, text fragmentation analysis and permission contradiction detection to build a multi-dimensional feature fusion mechanism and adaptive defense strategy.

Benefits of technology

High-precision detection of confrontational prompts is achieved, with the true positive rate ≥92% and the false positive rate ≤3%. Real-time risk determination can be completed within 50ms, and the new attack mode can be independently adapted to the new attack mode within 48 hours, taking into account the balance between security protection and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297419A_ABST
    Figure CN120297419A_ABST
Patent Text Reader

Abstract

The invention discloses an antagonism prompt detection method and device based on a large language model and a medium, and belongs to the technical field of artificial intelligence safety. The technical problem to be solved by the invention is how to overcome the defects of static rule lagging, high manual maintenance cost and insufficient context understanding in the security protection process of a large language model in the prior art, and dynamic, real-time and high-precision antagonism prompt detection is realized. According to the technical scheme, the method comprises the following steps: feature extraction: comprehensively analyzing semantic information, structural information and context information in a user text, extracting semantic features, structural features and context features, performing Min-Max normalization on the semantic features, the structural features and the context features, and splicing the semantic features, the structural features and the context features into a 128-dimensional joint vector, performing feature selection on the joint vector through an L1 regularization logistic regression model, compressing to 10-dimensional core features, and removing redundant information; carrying out resistance scoring; and dynamically defending.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence security, and specifically to a method, device and medium for adversarial prompt detection based on large language models. Background Art

[0002] The main security challenge faced by current large language models lies in their vulnerability to adversarial prompts. Attackers induce the model to generate responses containing harmful information, biased content or privacy leaks by designing inputs with hidden semantics and complex structures (such as replacing sensitive words with synonyms, inserting special encoding characters or exploiting logical loopholes). Although existing defense means (such as keyword filtering, sensitive word library matching) can intercept some explicit attacks, their mechanism relying on a static rule library has significant defects. The limitations of traditional defense technologies are further reflected in the lack of dynamic protection capabilities. Existing technologies lack the ability to perform in-depth semantic analysis and context logic reasoning on input texts. For example, in multi-turn conversations, attackers may evade single-turn detection through progressive induction, and detection mechanisms relying only on surface features (such as keywords, symbols) are difficult to identify such strategies, resulting in high false positive rates (such as normal requests being intercepted) and high false negative rates (such as hidden attacks not being detected), seriously restricting the actual effectiveness of the protection system.

[0003] Therefore, how to overcome the defects of static rule lag, high manual maintenance cost and insufficient context understanding in the security protection process of large language models in the prior art, and achieve dynamic, real-time and high-precision adversarial prompt detection is a technical problem to be solved urgently at present. Summary of the Invention

[0004] The technical task of the present invention is to provide a method, device and medium for adversarial prompt detection based on large language models to solve the problem of how to overcome the defects of static rule lag, high manual maintenance cost and insufficient context understanding in the security protection process of large language models in the prior art, and achieve dynamic, real-time and high-precision adversarial prompt detection.

[0005] The technical task of the present invention is realized in the following way. A method for adversarial prompt detection based on large language models is as follows:

[0006] Feature extraction: Comprehensively analyze the semantic information, structural information and context information in the user text, extract semantic features, structural features and context features, splice the semantic features, structural features and context features into a 128-dimensional joint vector after Min-Max normalization, and then perform feature selection on the joint vector through an L1-regularized logistic regression model, compress it to 10-dimensional core features, and eliminate redundant information;

[0007] Adversarial scoring: The core features are comprehensively evaluated through the BERT-BiLSTM dual-channel fusion model to generate a quantified risk score, and the attack possibility is judged in combination with a dynamic threshold adjustment mechanism; among them, the dynamic threshold adjustment mechanism adopts a two-stage control strategy based on historical data initialization and reinforcement learning online adjustment;

[0008] Dynamic defense: According to the risk scoring results, through a configurable policy engine and a self-evolving feedback closed loop, seamless connection from risk determination to protection actions is achieved, forming a dynamic balance of offensive and defensive confrontation.

[0009] Preferably, the semantic features are extracted as follows:

[0010] Using an intent classification model fine-tuned by BERT, the text is tokenized and encoded by Transformer to extract a 768-dimensional sentence vector marked with [CLS], which is mapped to a preset intent category space through a fully connected layer, and the probability distributions of high-risk operations and unauthorized data request categories are output; among them, the intent classification model fine-tuned by BERT is a text classifier implemented by fine-tuning (Fine-tuning) based on a pre-trained BERT model, and is used to identify the potential intent of user input;

[0011] Combining general named entity recognition (such as spaCy NER) and a custom high-risk entity library to construct a sensitive entity extraction module to perform double matching on risk entities in the text; and based on WordNet, perform synonym expansion matching on risk entities in the text, and then quantify the risk level by counting the entity density; among them, the custom high-risk entity library stores dangerous chemical and prohibited operation entries using a Trie tree structure; the entity density refers to the percentage between the number of high-risk entities and the total number of words.

[0012] Preferably, the structural features are extracted as follows:

[0013] Abnormal symbol detection: Matching consecutive non-word characters through a regular expression rule library, counting the proportion of consecutive non-word characters in the text and combining the Levenshtein distance algorithm for pattern clustering to identify potential abnormal symbols: if the symbol density exceeds the set threshold, it is judged as abnormal;

[0014] Encoding obfuscation detection: By traversing the Unicode encoding of characters, identifying the unconventional mixed use of ASCII and extended character sets, and using the FastText pre-trained character embedding model to quantify the semantic differences between characters to detect disguised characters;

[0015] Text fragmentation analysis: Calculate the standard deviation of sentence length, the frequency of non-standard punctuation, and the paragraph structure entropy. Use a sliding window to calculate the local coherence score, and jointly evaluate with the perplexity of the GPT-2 model to quantify the degree of text segmentation anomalies.

[0016] Preferably, the context features are extracted as follows:

[0017] Crack high-risk operation requests: Crack progressive high-risk operation requests by associating the current input text with the historical conversation;

[0018] Dialogue consistency check: Use the Sentence-BERT model to encode the 384-dimensional vectors of the recent several rounds of conversations and the current input, construct a cosine similarity matrix and set a dynamic threshold to detect logical mutation behaviors;

[0019] Permission conflict detection: Based on the knowledge graph, construct a rule inference engine to verify the node relationship between the user identity attributes and the operation requests, and use a graph attention network to mine implicit permission conflicts, and output the over-authorization probability or factual conflicts: When it is detected that there is an obvious permission mismatch between the user identity and the operation request: Through the preset permission rule library and the associated analysis of the knowledge graph, automatically identify and mark the corresponding type of abnormal requests.

[0020] Preferably, the BERT-BiLSTM dual-channel fusion model combines the semantic understanding ability of the pre-trained language model and the local pattern capture ability of the sequence model; the BERT-BiLSTM dual-channel fusion model includes a semantic encoding layer, a sequence pattern layer, and a decision layer;

[0021] Among them, the semantic encoding layer is used to perform WordPiece tokenization on the input text through the BERT model, and then generate a context-aware word vector sequence through 12 layers of Transformer encoders. The 768-dimensional vector of the [CLS] token is used as the global semantic representation to capture the induced semantics in the adversarial prompts;

[0022] The sequence pattern layer is used to adopt a bidirectional LSTM network with 128 units. The forward LSTM network analyzes the forward temporal dependence of the instruction logic, and the backward LSTM network reversely detects unconventional structures (such as inverted steps or abnormal symbol insertions). Concatenate the 128-dimensional forward hidden state and the 128-dimensional backward hidden state of the bidirectional LSTM network into a 256-dimensional sequence feature vector, and then concatenate it with the 768-dimensional [CLS] vector output by the BERT model to finally form a 1024-dimensional mixed feature vector, realizing the dual expression of global semantics and local patterns;

[0023] The decision-making layer is used to gradually reduce the dimension through a two-layer fully connected network. Specifically: the first layer compresses the 1024-dimensional mixed feature vector to 512 dimensions, applies ReLU activation and Dropout (p = 0.3), the second layer further reduces it to 128 dimensions, and finally outputs a risk score in the range of 0-1 through the Sigmoid function to quantify the adversarial attack probability of the input text.

[0024] Preferably, the training strategy of the BERT-BiLSTM dual-path fusion model is divided into two stages: pre-training and joint training, taking into account domain adaptation and class balance; specifically as follows:

[0025] Pre-training stage: Based on millions of adversarial samples (including 60% real attack data and 40% GPT-3.5 generated variants), fine-tune the BERT parameters through cross-entropy loss, focus on learning the inductive semantic patterns in the vertical fields of chemical synthesis and system penetration, and freeze the word embedding layer to retain general language knowledge;

[0026] Joint training stage: Freeze the parameters of the first 6 layers of Transformer in BERT, only optimize the last 6 layers, BiLSTM and the fully connected network to ensure maintaining the semantic understanding stability while reducing the training time by 70%; for the problem of insufficient proportion of high-risk samples, adopt the Focal Loss function (γ = 2, α = 0.8) to strengthen the attention to minority classes, and weight the difficult-to-classify samples through the adjustment factor (1 - p t ) γ to suppress the dominant influence of normal requests; during the training process, assist with dynamic batches (maximum 32) and mixed precision (FP16) to accelerate, set the initial learning rate according to the module (3e-5 for the BERT part, 1e-3 for BiLSTM and the fully connected layer), and adopt the cosine annealing strategy to alleviate local optimality.

[0027] Preferably, the determination of the initial threshold θ0 = 0.75 in the dynamic threshold adjustment mechanism has a dual basis, specifically: firstly, based on the optimal Youden index (maximizing TPR - FPR) of the BERT-BiLSTM dual-path fusion model on the ROC curve of the validation set, and at the same time referring to the historical attack data distribution collected by the system (the median of the attack request proportion in the past 30 days); in the real-time operation stage, achieve dynamic adjustment through a DQN (Deep Q-Network) reinforcement learning agent, and the core components include:

[0028] State space: Integrate the short-term attack situation and the reliability index of the BERT-BiLSTM dual-path fusion model, including the sliding window statistics (the proportion of high-risk requests in the recent 1000 requests, the trend of attack frequency in the recent 1 hour), the variance of the model confidence (the standard deviation of the inference scores in the past 500 times), and the real-time false alarm rate (FP / (FP + TN)) calculation results;

[0029] Action space: The discretized adjustment of the threshold θ, and the optional actions include {-0.05, -0.02, 0, +0.02, +0.05} to ensure that the adjustment amplitude is controllable;

[0030] Reward function, the formula is as follows:

[0031] R = 0.6 × TPR - 0.4 × FPR + 0.2 × (1 - |θ - θ0|) - 0.1 × I(θ change);

[0032] Among them, TPR and FPR are calculated based on the real-time confusion matrix, and the last term punishes frequent threshold fluctuations; among them, TPR represents the true positive rate; FPR represents the false positive rate; I(θ change) represents the number of adjustments on the current day;

[0033] The dynamic adjustment logic is divided into scenario-based responses: including the peak attack period, the calm period, and the transition period; among them, the peak attack period means that when the proportion of high-risk requests in the sliding window > 15% or the TPR of the BERT-BiLSTM dual-channel fusion model > 90%, the sensitivity priority mode is triggered, and θ is gradually reduced to 0.65 (lower limit protection); the calm period means that when the attack frequency < 5% and the confidence variance of the BERT-BiLSTM dual-channel fusion model < 0.1, switch to the false alarm suppression mode, and θ is increased to 0.8 (upper limit protection); the transition period means that the progressive adjustment is executed by default, the policy network parameters are updated every 24 hours, and the exploration and exploitation are balanced through the ε-greedy algorithm (ε = 0.15); this dynamic threshold adjustment mechanism achieves balance through the three-stage control of the peak attack period, the calm period, and the transition period. Specifically: the short-term sliding window (1000 requests) quickly responds to sudden attacks, the medium-term model metrics (1-hour TPR) ensure detection stability, and the long-term policy network update (24 hours) adapts to the evolution of the attack mode; compared with the fixed threshold scheme, this design increases the TPR by 12.7% (attack period) while reducing the FPR by 9.3% (calm period).

[0034] Preferably, the dynamic defense is as follows:

[0035] Hierarchical response: including low-risk requests, medium-risk requests, and high-risk requests; among them, the low-risk requests are scored 0 - 0.15, and the low-risk requests are directly forwarded to the large language model to generate responses, and the audit logs are asynchronously recorded in the Kafka message queue; the medium-risk requests are scored θ - 0.15 - θ + 0.1, and the medium-risk requests trigger the secondary verification interface, return a JSON format risk summary (including the list of high-risk entities and the intent category), the front-end renders a dynamic confirmation page, and after the user authorizes, the request enters the desensitization pipeline with the [Verified] mark; the high-risk requests are scored > θ + 0.1, and the high-risk requests are immediately blocked and an HTTP 403 response is returned, and the attack traceability analysis is started synchronously, and the threat index is calculated by associating the user behavior graph;

[0036] Self-evolution mechanism: The self-evolution feedback loop achieves iterative upgrades of defense capabilities through continuous learning and optimization. Specifically: First, persistently store the intercepted high-risk requests and their context features in the Elasticsearch attack feature library, identify new attack patterns through the DBSCAN density clustering algorithm, and visualize the clustering results on the security analysis dashboard. Then, initiate an incremental training process based on the newly added attack samples daily, use the hierarchical oversampling technique (expanding high-risk samples by 3 times) to perform hot updates on the top fully connected network of the BERT-BiLSTM dual-path fusion model, and at the same time, release the updated model in a gray-scale manner with a 10% traffic ratio through the A / B test framework, and compare the TPR / FPR metrics of the new and old versions in real time to verify effectiveness. Finally, based on the evaluation result of the confusion success rate (attack bypass rate < 1%), use the Q-Learning algorithm to dynamically adjust the weights of the content desensitization rules, and synchronously feedback the optimized policy parameters to feature extraction, forming a complete adaptive closed-loop of "attack capture → model iteration → rule tuning → defense enhancement".

[0037] An electronic device, comprising: a memory and at least one processor;

[0038] Wherein, a computer program is stored on the memory;

[0039] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the method for detecting adversarial prompts based on a large language model as described above.

[0040] A computer-readable storage medium, in which a computer program is stored, and the computer program can be executed by a processor to implement the method for detecting adversarial prompts based on a large language model as described above.

[0041] The method, device and medium for detecting adversarial prompts based on a large language model of the present invention have the following advantages:

[0042] (1) The present invention is used to identify and intercept adversarial prompt attacks against a large language model (LLM), improving the security and robustness of the model;

[0043] (2) The present invention solves the core problems existing in the existing security protection technologies for large language models, such as the lag of static rules, the high cost of manual maintenance, and the lack of context understanding. It provides a real-time high-precision protection solution that can dynamically adapt to the evolution of attacks. By constructing a multi-dimensional feature fusion mechanism and an adaptive defense strategy, it breaks through the limitations of traditional keyword filtering and single-model detection, realizes the accurate identification of complex adversarial prompts such as semantic variants, obfuscated coding, and multi-round logical induction, and at the same time realizes the dynamic, real-time, and high-precision detection of adversarial prompts. Through multi-dimensional feature fusion and adaptive strategies, the security of large models is improved;

[0044] (3) The present invention realizes the high-precision detection of adversarial prompts (true positive rate ≥ 92%, false positive rate ≤ 3%) through the extraction of multi-dimensional features that fuse semantics, structure, and context and a dual-path fusion model (BERT-BiLSTM). Combined with a dynamic threshold adjustment mechanism and a self-evolving closed loop driven by reinforcement learning, the system can complete real-time risk determination within 50 ms and autonomously adapt to new attack patterns within 48 hours (the detection accuracy is increased by an average of 2.3%). At the same time, through hierarchical response and content desensitization technology that preserves semantics, while blocking high-risk requests (such as inducing the manufacture of dangerous goods), the interference with normal interactions is minimized as much as possible, forming a full-chain attack and defense system covering feature parsing, dynamic scoring, hierarchical defense, and autonomous optimization, effectively coping with various complex attack scenarios such as semantic induction, structure confusion, and privilege escalation, and taking into account the balance between security protection and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The present invention will be further described below with reference to the accompanying drawings.

[0046] FIG Figure 1 is a schematic flowchart of a method for detecting adversarial prompts based on a large language model;

[0047] FIG Figure 2 is a schematic flowchart of feature extraction;

[0048] FIG Figure 3 is a schematic flowchart of adversarial scoring;

[0049] FIG Figure 4 is a schematic flowchart of dynamic defense. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The method, device, and medium for detecting adversarial prompts based on a large language model of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.

[0051] Example 1:

[0052] As shown in FIG Figure 1As shown, this embodiment provides a method for detecting adversarial prompts based on a large language model, and the method is specifically as follows:

[0053] S1. Feature extraction: Comprehensively analyze the semantic information, structural information and contextual information in the user text, extract semantic features, structural features and contextual features, and concatenate them into a 128-dimensional joint vector after Min-Max normalization. Then, the joint vector is subjected to feature selection through the L1 regularized logistic regression model, compressed into 10-dimensional core features, and redundant information is removed. Redis is used to cache intermediate features (such as conversation history vectors) and asynchronous parallel computing (semantic / structural / contextual feature extraction is executed in parallel) to ensure that the end-to-end delay is ≤50ms (P99 indicator) to meet the response requirements of real-time detection, as shown in the attached figure. Figure 2 As shown;

[0054] S2. Adversarial scoring: The core features are comprehensively evaluated through the BERT-BiLSTM dual-path fusion model to generate a quantitative risk score, and the possibility of attack is judged in combination with a dynamic threshold adjustment mechanism. The dynamic threshold adjustment mechanism adopts a two-stage control strategy based on historical data initialization and reinforcement learning online adjustment.

[0055] S3. Dynamic defense: Based on the risk scoring results, a configurable strategy engine and a self-evolving feedback loop are used to achieve seamless connection from risk assessment to protection actions, forming a dynamic balance between offense and defense.

[0056] The semantic features extracted in step S1 of this embodiment are specifically as follows:

[0057] S1-101. Use the BERT fine-tuned intent classification model to tokenize and Transformer encode the text, extract the 768-dimensional sentence vector of the [CLS] tag, map it to the preset intent category space through the fully connected layer, and output the probability distribution of high-risk operations and unauthorized data request categories; Among them, the BERT fine-tuned intent classification model is a text classifier implemented by fine-tuning domain data based on the pre-trained BERT model, which is used to identify the potential intent of user input (such as "induced attack" and "normal request").

[0058] S1-102. Construct a sensitive entity extraction module by combining general named entity recognition (such as spaCy NER) with a custom high-risk entity library (covering hazardous chemicals, prohibited operations, etc.), and perform double matching on the risk entities in the text. For example, when the input is "To synthesize an explosive, A and B need to be mixed", the high-risk entity "explosive" will be recognized; and based on WordNet, perform synonym expansion matching on the risk entities in the text, and then quantify the risk level by calculating the entity density (such as more than 3 high-risk words per 100 words); among them, the custom high-risk entity library stores the entries of hazardous chemicals and prohibited operations using a Trie tree structure; the entity density refers to the percentage between the number of high-risk entities and the total number of words.

[0059] The specific extraction of structural features in step S1 of this embodiment is as follows:

[0060] S1-201. Abnormal symbol detection: Match consecutive non-word characters through a regular expression rule library, calculate the proportion of consecutive non-word characters in the text, and perform pattern clustering in combination with the Levenshtein distance algorithm to identify potential abnormal symbols: If the symbol density exceeds the set threshold, it is judged as abnormal;

[0061] S1-202. Encoding confusion detection: By traversing the Unicode encoding of characters, identify the unconventional mixed use of ASCII and extended character sets, and use the FastText pre-trained character embedding model to quantify the semantic differences between characters to detect disguised characters;

[0062] S1-203. Text fragmentation analysis: Calculate the standard deviation of sentence length, the frequency of non-standard punctuation, and the paragraph structure entropy, calculate the local coherence score using a sliding window, and jointly evaluate with the perplexity of the GPT-2 model to quantify the degree of text segmentation abnormality.

[0063] Among them, the extraction of structural features focuses on identifying unconventional text patterns deliberately designed by attackers. Abnormal symbol detection matches consecutive non-word characters through regular expressions (such as "", "%20"), calculates their proportion in the text, and if the symbol density exceeds 30%, it is determined as abnormal. Encoding confusion detection traverses the character encoding to identify inputs that mix the use of ASCII and Unicode (such as Cyrillic characters disguised), for example, the first letter "а" in "аpple" is actually a Cyrillic character, and the system will mark such confusion behaviors. The degree of text fragmentation is jointly evaluated by the standard deviation of sentence length and punctuation density: If the input is like "Step 1: Obtain A; Step 2: Mix B → Step 3: Test", showing short sentences with high-frequency separation (average sentence length less than 5 words) and unconventional punctuation (such as arrow symbols), the system will determine that its fragmentation score exceeds the standard and prompt a potential attack structure.

[0064] The specific extraction of context features in step S1 of this embodiment is as follows:

[0065] S1-301, cracking high-risk operation requests: cracking progressive high-risk operation requests by associating the current input text with historical conversations;

[0066] S1-302, Conversation consistency check: Use the Sentence-BERT model to encode the 384-dimensional vector of the most recent conversations and the current input, build a cosine similarity matrix and set a dynamic threshold to detect logical mutation behavior;

[0067] S1-303, permission contradiction detection: build a rule reasoning engine based on the knowledge graph, verify the node relationship between user identity attributes and operation requests, and use the graph attention network to mine implicit permission conflicts, and output the probability of over-authorization or factual conflict: when it is detected that there is an obvious permission mismatch between the user identity and the operation request (for example, ordinary users request special operations involving professional fields): automatically identify and mark the corresponding abnormal requests through the preset permission rule library and knowledge graph association analysis.

[0068] Among them, context feature analysis deciphers progressive high-risk operation requests by associating current input with historical conversations. The conversation history consistency check uses Sentence-BERT to generate sentence vectors and calculates the cosine similarity between the current input and the last five rounds of conversations. For example, when the user's previous conversation revolves around "cooking skills" and suddenly asks "how to make a Molotov cocktail", the similarity drops sharply to 0.15, triggering a logical mutation alarm. Logical contradiction detection combines the rule base with the graph neural network to identify permission jumps or fact conflicts in the input: when it is detected that there is an obvious permission mismatch between the user identity and the requested operation (for example, an ordinary user requests a special operation involving a professional field), the system will automatically identify and mark such abnormal requests through the preset permission rule base and knowledge graph association analysis.

[0069] The BERT-BiLSTM two-way fusion model in step S2 of this embodiment combines the semantic understanding ability of the pre-trained language model with the local pattern capture ability of the sequence model; the BERT-BiLSTM two-way fusion model includes a semantic encoding layer, a sequence pattern layer, and a decision layer;

[0070] The semantic encoding layer is used to generate a context-aware word vector sequence through a 12-layer Transformer encoder after WordPiece segmentation of the input text through the BERT model. The 768-dimensional vector of the [CLS] tag is used as a global semantic representation to capture the inductive semantics in the adversarial prompts (such as key intentions such as "bypass verification" and "step-by-step guidance").

[0071] The sequence pattern layer uses a bidirectional LSTM network with 128 units. The forward LSTM network analyzes the forward temporal dependencies of the instruction logic, and the backward LSTM network reversely detects unconventional structures (such as inverted steps or abnormal symbol insertions). The 128-dimensional forward hidden state and the 128-dimensional backward hidden state of the bidirectional LSTM network are concatenated into a 256-dimensional sequence feature vector, which is then concatenated with the 768-dimensional [CLS] vector output by the BERT model, finally forming a 1024-dimensional hybrid feature vector to achieve the dual expression of global semantics and local patterns;

[0072] The decision-making layer is used to gradually reduce the dimension through a two-layer fully connected network. Specifically: the first layer compresses the 1024-dimensional hybrid feature vector to 512 dimensions and applies ReLU activation and Dropout (p = 0.3), and the second layer further reduces it to 128 dimensions. Finally, a risk score in the range of 0-1 is output through the Sigmoid function to quantify the adversarial attack probability of the input text.

[0073] As shown in the Figure 3 appendix, the training strategy of the BERT-BiLSTM dual-path fusion model in step S2 of this embodiment is divided into two stages: pre-training and joint training, taking into account domain adaptation and class balance; specifically as follows:

[0074] S201. Pre-training stage: Based on millions of adversarial samples (including 60% real attack data and 40% GPT-3.5 generated variants), fine-tune the BERT parameters through cross-entropy loss, focus on learning the inductive semantic patterns in the vertical fields of chemical synthesis and system penetration, and at the same time freeze the word embedding layer to retain general language knowledge;

[0075] S202. Joint training stage: Freeze the first 6 layers of Transformer parameters of BERT, and only optimize the last 6 layers, BiLSTM and fully connected network to ensure maintaining the semantic understanding stability while reducing the training time by 70%; to address the problem of insufficient proportion of high-risk samples, use the Focal Loss function (γ = 2, α = 0.8) to strengthen the attention to minority classes, and weight the difficult-to-classify samples through the adjustment factor (1 - p t ) γ to suppress the dominant influence of normal requests; during the training process, assist with dynamic batches (maximum 32) and mixed precision (FP16) to accelerate, set the initial learning rate by module (3e-5 for the BERT part, 1e-3 for BiLSTM and fully connected layers), and use the cosine annealing strategy to alleviate local optima.

[0076] The determination of the initial threshold θ0 = 0.75 in the dynamic threshold adjustment mechanism in step S2 of this embodiment has a two-fold basis, specifically: Firstly, based on the optimal Youden index (maximizing TPR - FPR) of the BERT - BiLSTM dual - path fusion model on the ROC curve of the validation set, and at the same time referring to the historical attack data distribution collected by the system (the median of the proportion of attack requests in the past 30 days); In the real - time operation stage, dynamic adjustment is achieved through a DQN (DeepQ - Network) reinforcement learning agent. The core components include:

[0077] ① State space: Integrate the short - term attack situation and the reliability index of the BERT - BiLSTM dual - path fusion model, including sliding window statistics (the proportion of high - risk requests in the last 1000 requests, the trend of attack frequency in the last hour), the variance of model confidence (the standard deviation of the inference scores in the past 500 times), and the real - time false - positive rate (FP / (FP + TN)) calculation results;

[0078] ② Action space: Discretized adjustment of the threshold θ, and the optional actions include {-0.05, -0.02, 0, +0.02, +0.05}, ensuring that the adjustment range is controllable;

[0079] ③ Reward function, the formula is as follows:

[0080] R = 0.6×TPR - 0.4×FPR + 0.2×(1 - |θ - θ0|)-0.1×I(θ change);

[0081] Among them, TPR and FPR are calculated based on the real - time confusion matrix, and the last term punishes frequent threshold fluctuations; where TPR represents the true - positive rate; FPR represents the false - positive rate; I(θ change) represents the number of adjustments on the current day;

[0082] ④ The dynamic adjustment logic is divided into scenario-based responses: including the peak attack period, the calm period, and the transition period. Among them, the peak attack period refers to when the proportion of high-risk requests in the sliding window > 15% or the TPR of the BERT-BiLSTM dual-channel fusion model > 90%, triggering the sensitivity priority mode, and θ gradually decreases to 0.65 (lower limit protection); the calm period refers to when the attack frequency < 5% and the confidence variance of the BERT-BiLSTM dual-channel fusion model < 0.1, switching to the false alarm suppression mode, and θ increases to 0.8 (upper limit protection); the transition period refers to the default execution of progressive adjustment, updating the policy network parameters every 24 hours, and balancing exploration and exploitation through the ε-greedy algorithm (ε = 0.15). This dynamic threshold adjustment mechanism achieves balance through the three-stage control of the peak attack period, the calm period, and the transition period. Specifically, the short-term sliding window (1000 requests) quickly responds to sudden attacks, the medium-term model metrics (1-hour TPR) ensure detection stability, and the long-term policy network update (24 hours) adapts to the evolution of the attack mode. Compared with the fixed threshold scheme, this design increases the TPR by 12.7% (attack period) while reducing the FPR by 9.3% (calm period).

[0083] As shown in the appendix Figure 4 shown, the dynamic defense in step S3 of this embodiment is specifically as follows:

[0084] S301. Hierarchical response: including low-risk requests, medium-risk requests, and high-risk requests. Among them, the low-risk requests are scored 0 - 0.15, and the low-risk requests are directly forwarded to the large language model to generate responses, asynchronously recording audit logs to the Kafka message queue; the medium-risk requests are scored (θ - 0.15 —— θ + 0.1), the medium-risk requests trigger the secondary verification interface, return a JSON format risk summary (including the list of high-risk entities and the intent category), the front-end renders a dynamic confirmation page, and after the user authorizes, the request enters the desensitization pipeline with the [Verified] mark; the high-risk requests are scored > θ + 0.1, the high-risk requests are immediately blocked and an HTTP 403 response is returned, and the attack traceability analysis is synchronously started, and the threat index is calculated by associating the user behavior graph.

[0085] S302. Self-Evolution Mechanism: The self-evolution feedback loop realizes the iterative upgrade of the defense ability through continuous learning and optimization. Specifically, first, the intercepted high-risk requests and their context features are persistently stored in the Elasticsearch attack feature library. The DBSCAN density clustering algorithm is used to identify new attack patterns (such as detecting a variant cluster that bypasses the symbol "##{instruction}##"), and the clustering results are visualized on the security analysis dashboard. Then, based on the newly added attack samples every day, an incremental training process is initiated. The top fully connected network of the BERT-BiLSTM dual-path fusion model is hot-updated using the hierarchical oversampling technique (the high-risk samples are expanded by 3 times). At the same time, through the A / B test framework, the updated model is released in a gray-scale manner with a 10% traffic ratio, and the TPR / FPR metrics of the new and old versions are compared in real time to verify the effectiveness. Finally, based on the evaluation result of the confusion success rate (the attack bypass rate < 1%), the Q-Learning algorithm is used to dynamically adjust the weights of the content desensitization rules (such as increasing the replacement priority of chemical entity types), and the optimized policy parameters are synchronously fed back to feature extraction, forming a complete adaptive loop of "attack capture → model iteration → rule tuning → defense enhancement". When facing the continuous evolution of adversarial prompts, the self-evolution mechanism can achieve autonomous evolution at a speed of increasing the detection accuracy by 2.3% every 48 hours on average.

[0086] Embodiment 2:

[0087] This embodiment also provides an electronic device, including: a memory and a processor;

[0088] Among them, the memory stores computer execution instructions;

[0089] The processor executes the computer execution instructions stored in the memory, so that the processor executes the method for detecting adversarial prompts based on a large language model in any embodiment of the present invention.

[0090] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0091] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and invoking the data stored in the memory, the processor realizes various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, memory, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash memory cards, at least one magnetic disk storage period, flash memory devices, or other volatile solid-state storage devices.

[0092] Embodiment 3:

[0093] This embodiment also provides a computer-readable storage medium, which stores multiple instructions. The instructions are loaded by the processor to make the processor execute the adversarial prompt detection method based on the large language model in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided. On this storage medium, software program codes for realizing the functions of any one of the above embodiments are stored, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium.

[0094] In this case, the program code read from the storage medium itself can realize the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0095] Embodiments of the storage medium for providing program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.

[0096] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by making the operating system and the like operating on the computer based on the instructions of the program code complete part or all of the actual operations.

[0097] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer. Subsequently, based on the instructions of the program code, the CPU and the like installed on the expansion board or the expansion unit execute part and all of the actual operations, thereby realizing the functions of any one of the above embodiments.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting adversarial prompts based on large language models, characterized in that, The method is as follows: Feature extraction: Comprehensively analyze the semantic information, structural information and contextual information in the user text, extract semantic features, structural features and contextual features, and concatenate them into a 128-dimensional joint vector after Min-Max normalization. Then, the joint vector is subjected to feature selection through the L1 regularized logistic regression model, compressed into 10-dimensional core features, and redundant information is eliminated; Adversarial scoring: The core features are comprehensively evaluated through the BERT-BiLSTM dual-path fusion model to generate a quantitative risk score, and the possibility of attack is judged in combination with a dynamic threshold adjustment mechanism. The dynamic threshold adjustment mechanism adopts a two-stage control strategy based on historical data initialization and reinforcement learning online adjustment. Dynamic defense: Based on the risk scoring results, a configurable policy engine and a self-evolving feedback loop are used to achieve seamless connection from risk assessment to protection actions, forming a dynamic balance between offense and defense.

2. The method for detecting adversarial prompts based on a large language model according to claim 1, wherein The semantic features extracted are as follows: The intent classification model fine-tuned by BERT is used to tokenize and Transformer encode the text, extract the 768-dimensional sentence vector of the [CLS] tag, map it to the preset intent category space through the fully connected layer, and output the probability distribution of high-risk operations and unauthorized data request categories; the intent classification model fine-tuned by BERT is a text classifier based on the pre-trained BERT model and implemented through domain data fine-tuning, which is used to identify the potential intent of user input; A sensitive entity extraction module is constructed by combining general named entity recognition with a custom high-risk entity library to perform double matching on risk entities in the text. Synonym expansion matching is performed on risk entities in the text based on WordNet, and then the risk level is quantified by statistical entity density. Among them, the custom high-risk entity library uses a Trie tree structure to store hazardous chemicals and prohibited operation entries. Entity density refers to the percentage between the number of high-risk entities and the total number of words.

3. The method for detecting adversarial prompts based on a large language model according to claim 1, wherein The structural features extracted are as follows: Abnormal symbol detection: Match continuous non-word characters through the regular expression rule library, count the proportion of continuous non-word characters in the text, and perform pattern clustering in combination with the Levenshtein distance algorithm to identify potential abnormal symbols: if the symbol density exceeds the set threshold, it is judged as abnormal; Encoding obfuscation detection: By traversing the Unicode encoding of characters, identifying the unconventional mixed use of ASCII and extended character sets, and using the FastText pre-trained character embedding model to quantify the semantic differences between characters, and detect disguised characters; Text fragmentation analysis: Statistics are collected on the standard deviation of sentence length, frequency of non-standard punctuation, and paragraph structure entropy. A sliding window is used to calculate the local coherence score, which is then combined with the perplexity evaluation of the GPT-2 model to quantify the degree of abnormality in text segmentation.

4. The method for detecting adversarial prompts based on large language models according to claim 1, wherein The specific context features are as follows: Cracking high-risk operation requests: Cracking progressive high-risk operation requests by associating the current input text with historical conversations; Dialogue Consistency Verification: Use the Sentence-BERT model to encode the vectors of the recent several rounds of dialogues and the current input with 384 dimensions, construct a cosine similarity matrix and set a dynamic threshold to detect logical mutation behaviors; Permission Conflict Detection: Based on the knowledge graph, construct a rule inference engine to verify the node relationship between the user identity attributes and the operation requests, and use the graph attention network to mine implicit permission conflicts, and output the probability of overstepping authority or factual conflicts. When it is detected that there is an obvious mismatch between the user identity and the operation request: through the preset permission rule library and the associated analysis of the knowledge graph, automatically identify and mark the corresponding abnormal requests.

5. The method for detecting adversarial prompts based on large language models according to claim 1, wherein The BERT-BiLSTM dual-path fusion model combines the semantic understanding ability of the pre-trained language model and the local pattern capturing ability of the sequence model; the BERT-BiLSTM dual-path fusion model includes a semantic encoding layer, a sequence pattern layer, and a decision-making layer; Among them, the semantic encoding layer is used to perform WordPiece tokenization on the input text through the BERT model, and then generate a context-aware word vector sequence through 12 layers of Transformer encoders. The 768-dimensional vector marked with [CLS] is used as the global semantic representation to capture the induced semantics in the adversarial prompt; The sequence pattern layer is used to adopt a bidirectional LSTM network with 128 units. The forward LSTM network analyzes the forward temporal dependence of the instruction logic, and the backward LSTM network reversely detects the unconventional structure. The 128-dimensional forward hidden state and the 128-dimensional backward hidden state of the bidirectional LSTM network are concatenated into a 256-dimensional sequence feature vector, and then concatenated with the 768-dimensional [CLS] vector output by the BERT model to finally form a 1024-dimensional mixed feature vector, realizing the dual expression of global semantics and local patterns; The decision-making layer is used to gradually reduce the dimension through two layers of fully connected networks. Specifically: the first layer compresses the 1024-dimensional mixed feature vector to 512 dimensions and applies ReLU activation and Dropout, and the second layer is further reduced to 128 dimensions. Finally, the risk score in the 0-1 interval is output through the Sigmoid function to quantify the adversarial attack probability of the input text.

6. The method for detecting adversarial prompts based on large language models according to claim 1, wherein The training strategy of the BERT-BiLSTM dual-path fusion model is divided into two stages: pre-training and joint training, taking into account domain adaptation and class balance; specifically as follows: Pre-training stage: Based on millions of adversarial samples, fine-tune the BERT parameters through cross-entropy loss, focus on learning the induced semantic patterns in the vertical fields of chemical synthesis and system penetration, and freeze the word embedding layer to retain general language knowledge; Joint training stage: Freeze the parameters of the first 6 layers of Transformer in BERT, and only optimize the last 6 layers, BiLSTM, and fully connected networks to ensure that while reducing the training time by 70%, maintain the stability of semantic understanding; Adopt the Focal Loss function to strengthen the attention to the minority class, and adjust the factor (1 - p t ) γ weight the difficult-to-classify samples to suppress the dominant influence of normal requests; During the training process, dynamic batches and mixed precision are used to accelerate, the initial learning rate is set separately by module, and the cosine annealing strategy is adopted to alleviate local optimality.

7. The adversarial prompt detection method based on large language models according to claim 1, wherein The determination of the initial threshold θ0 = 0.75 in the dynamic threshold adjustment mechanism has a dual basis, specifically: first, based on the optimal Youden index of the BERT-BiLSTM dual-path fusion model on the ROC curve of the validation set, and at the same time referring to the historical attack data distribution collected by the system; During the real-time operation phase, dynamic adjustment is achieved through a DQN reinforcement learning agent. The core components include: State space: Integrate the short-term attack situation and the reliability indicators of the BERT-BiLSTM dual-channel fusion model, including sliding window statistics, model confidence variance, and real-time false alarm rate calculation results; Action space: Discretized adjustment of the threshold θ. The optional actions include {-0.05, -0.02, 0, +0.02, +0.05} to ensure that the adjustment range is controllable; Reward function, the formula is as follows: R = 0.6×TPR - 0.4×FPR + 0.2×(1 - |θ - θ0|) - 0.1×I(θ change); Among them, TPR and FPR are calculated based on the real-time confusion matrix, and the last term punishes frequent threshold fluctuations. Among them, TPR represents the true positive rate; FPR represents the false positive rate; I(θ change) represents the number of adjustments on the current day; The dynamic adjustment logic is divided into scenario-based responses: including peak attack period, calm period, and transition period. Among them, the peak attack period refers to when the proportion of high-risk requests in the sliding window > 15% or the TPR of the BERT-BiLSTM dual-channel fusion model > 90%, triggering the sensitivity priority mode, and θ is gradually reduced to 0.65; the calm period refers to when the attack frequency < 5% and the confidence variance of the BERT-BiLSTM dual-channel fusion model < 0.1, switching to the false alarm suppression mode, and θ is increased to 0.8; the transition period refers to the default execution of progressive adjustment, updating the policy network parameters every 24 hours, and balancing exploration and exploitation through the ε-greedy algorithm. This dynamic threshold adjustment mechanism achieves balance through three-order control of the peak attack period, calm period, and transition period. Specifically: the short-term sliding window quickly responds to sudden attacks, the medium-term model indicators ensure detection stability, and the long-term policy network updates to adapt to the evolution of the attack pattern.

8. The method for detecting adversarial prompts based on large language models according to claim 1, characterized in that, The dynamic defense is as follows: Hierarchical response: including low-risk requests, medium-risk requests, and high-risk requests. Among them, the low-risk requests are scored 0 - 0.15, and the low-risk requests are directly forwarded to the large language model to generate a response, asynchronously recording the audit log to the Kafka message queue; the medium-risk requests are scored θ - 0.15 - θ + 0.1, and the medium-risk requests trigger the secondary verification interface, returning a risk summary in JSON format, rendering a dynamic confirmation page on the front end, and the request enters the desensitization pipeline with the [Verified] mark after user authorization; the high-risk requests are scored > θ + 0.1, and the high-risk requests are immediately blocked and an HTTP 403 response is returned, and the attack traceability analysis is started synchronously, and the threat index is calculated by associating the user behavior graph; Self-evolution mechanism: The self-evolution feedback closed-loop realizes the iterative upgrade of the defense ability through continuous learning and optimization. Specifically: First, persistently store the intercepted high-risk requests and their context features in the Elasticsearch attack feature library, identify new attack patterns through the DBSCAN density clustering algorithm, and visualize the clustering results on the security analysis dashboard. Then, initiate an incremental training process based on the daily newly added attack samples, use the hierarchical oversampling technique to perform hot updates on the top fully connected network of the BERT-BiLSTM dual-channel fusion model, and at the same time, release the updated model in a gray-scale manner with a 10% traffic ratio through the A / B test framework, and compare the TPR / FPR metrics of the new and old versions in real time to verify the effectiveness. Finally, based on the evaluation result of the obfuscation success rate, use the Q-Learning algorithm to dynamically adjust the weights of the content desensitization rules, and synchronously feedback the optimized policy parameters to feature extraction, forming a complete adaptive closed-loop of "attack capture → model iteration → rule tuning → defense enhancement".

9. An electronic device, characterized in that, Including: A memory and at least one processor; Wherein, a computer program is stored on the memory; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the method for detecting adversarial prompts based on a large language model according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program can be executed by a processor to implement the method for detecting adversarial prompts based on a large language model according to any one of claims 1 to 8.

Citation Information

Cited By

  • Prompt interference construction and optimization method and device based on fragment semantic cross combination

    CN120745619A

  • Prompt interference construction and optimization method and device based on fragment semantic cross combination

    CN120745619B

  • Network security intrusion detection system and method fused with graph neural network

    CN120856447A

  • Large model real-time safety protection method and system based on dynamic response

    CN121072698A

  • Text-oriented violent word abbreviation detection method and device, equipment and storage medium

    CN121212140A