A security defense method for artificial intelligence models against adversarial attacks
Through layers of filtering, including filters, inductive models, and security classifiers, the vulnerability of AI models to adversarial attacks is resolved, ensuring that the output complies with ethics and laws, and improving the security and reliability of the model.
Patent Information
- Application Number
- CN202510889159.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Artificial intelligence models are vulnerable to adversarial attacks, resulting in unsafe outputs, violating ethical and legal norms, and affecting user data privacy and model reliability.
Before input, data is filtered through filters, inductive models, and security classifiers to identify and reject adversarial attacks. Filters initially filter pre-answers, inductive models generate summaries, and security classifiers identify harmful content, ensuring that output complies with ethical and legal standards.
It significantly improves the ability to identify adversarial attacks, ensures that the output is ethical and legal, and protects model integrity and user safety.
Smart Images

Figure CN120429874B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence security technology, and in particular to an artificial intelligence model security defense method against adversarial attacks. Background Art
[0002] Artificial intelligence models are complex mathematical systems constructed using large amounts of data and specific algorithms. They can simulate human intelligent behavior, perform pattern recognition, decision-making, and knowledge learning. They are widely used in multiple fields such as computer vision, natural language processing, and speech recognition, providing strong technical support for automated tasks, data analysis, and intelligent interaction.
[0003] To prevent AI models from being misused, many security alignment strategies embed safety guardrails within AI models to identify harmful or toxic semantics of cues, autonomously rejecting harmful inputs and avoiding the generation of unsafe content. While these alignment methods improve the security of AI models and are widely adopted in both open-source and closed-source models, they remain vulnerable to adversarial attacks. Adversarial attacks subtly modify harmful inputs to produce cues that bypass these safety guardrails, causing AI models to produce unsafe outputs that would normally be blocked. This poses a significant security threat to the practical application of AI models.
[0004] Therefore, security defense of artificial intelligence models against adversarial attacks is crucial. It can not only effectively protect user data privacy and model intellectual property rights, prevent sensitive information from being leaked and maliciously exploited, but also maintain the reliability and social trust of the model, ensure that its output complies with ethical and legal norms, guarantee the stable operation of the system and the safe use of users, and lay a solid foundation for the healthy development and widespread application of large language models. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to design an artificial intelligence model security defense method against adversarial attacks existing in the input of the artificial intelligence model. On the one hand, it ensures the integrity of the protected large language model, and on the other hand, it blocks adversarial attacks in the input stage and ensures that the output of the large language model complies with moral ethics and laws and regulations.
[0006] The technical solution adopted by the present invention to solve the above technical problems is to input a sample model before prompting the protected model to form a pre-answer and pass it to the filter. During the filter detection process, the pre-answer is preliminarily filtered to determine whether the safety fence of the pre-sample model blocks the prompt. If not, execution continues; if blocked, the output is rejected. The pre-answer that is not blocked is then passed to the inductive model for processing to form a summary. Finally, the security classifier is used to judge the summary. If it is judged to be harmful, the output is rejected; if it is judged to be harmless, the protected model is allowed to process the prompt. The following steps are included:
[0007] Step S1: Before prompting for the protected AI model, the prompt is input into the sample model to form a pre-answer and transmit it to the filter;
[0008] Step S2: The filter receives the pre-answer of the sample model and performs preliminary filtering on the pre-answer; determines whether the safety fence of the sample model blocks the prompt; if not, transmits the pre-answer to the inductive model; if blocked, refuses to output;
[0009] Step S3: Fine-tune BART into an inductive model. The inductive model receives the pre-answer transmitted by the filter, processes the pre-answer, generates a summary, and transmits the summary to the security classifier.
[0010] Step S4: Fine-tune BERT into a security classifier; the security classifier receives the summary and judges it. If it is judged to be harmful, it will refuse to output it. If it is judged to be harmless, it will allow the protected artificial intelligence model to process the prompt.
[0011] Furthermore, the sample model of step S1 is an open source artificial intelligence model with a safety guardrail.
[0012] Furthermore, step S3 includes the following steps:
[0013] Step S31: fine-tune BART into an inductive model;
[0014] Step S311: randomly selecting N benign prompts and harmful prompts to form a prompt data set; preprocessing the data set to obtain an input ID sequence and an output ID sequence;
[0015] Step S312: In the forward propagation phase, the input ID sequence Input encoder to get context representation ; Then the context is expressed as Input decoder to generate target sequence ;
[0016] Step S313: Calculate the generated target sequence target The real target sequence The cross entropy loss between them gives the inductive model loss function;
[0017] Step S314: Back propagation, updating the model weights by calculating the gradient of the inductive model loss function with respect to the model parameters;
[0018] Step S315: Parameter update: Use the Adam optimizer to update the model parameters according to the gradient;
[0019] Step S316: Update the model parameters through back propagation and Adam optimizer to obtain a fine-tuned inductive model.
[0020] Step S32: Use the fine-tuned inductive model to process the pre-answer, generate a summary, and transmit the summary to the security classifier.
[0021] Furthermore, the inductive model loss function is expressed as:
[0022]
[0023] in, Inductive model loss function, Refers to the sequence The corresponding number of tags, Target sequence The corresponding length during the generation process, Is the generated length less than The target sequence, The decoded length is less than The context representation of The length of the real sequence is part. The decoder generates the target sequence The conditional probability of .
[0024] Furthermore, step S4 includes the following steps:
[0025] Step S41: Fine-tune BERT into a security classifier;
[0026] Step S411: Randomly selected benign and harmful prompts as training sets , and the training set Label and obtain the labeled training set ;
[0027] Step S412: The training set Input BERT and obtain the feature vector after BERT encoding;
[0028] Step S413: Based on the output of BERT, add a classifier layer to transform the feature vector Mapping to harmless scores;
[0029] Step S414: Using the security classifier loss function to optimize model parameters during forward propagation;
[0030] Step S415: Initialize the parameters first in the back propagation phase: Randomly initialize the weight matrix and the bias vector . Then calculate the security classifier loss function About the weight matrix Gradient and about Gradient , update the model parameters.
[0031] Step S416: Optimize the model parameters using the security classifier loss function and update the parameters using backpropagation to obtain a fine-tuned security classifier.
[0032] Step S42: The security classifier receives the summary and judges it. If it is judged to be harmful, the output is rejected. If it is judged to be harmless, the protected model is allowed to process the prompt.
[0033] Furthermore, the security classifier loss function is expressed as:
[0034]
[0035] in, represents the security classifier loss function, Indicates a prompt The corresponding annotations, Indicates harmless score, Is the Sigmoid function, used to convert the harmless score Mapped between 0 and 1.
[0036] This invention leverages the ability of large language models to perform specific tasks through fine-tuning, allowing two models to perform the induction task and the security classification task, respectively, forming an induction model and a security classifier. This method, through a filter, induction model, and security classifier, filters prompts layer by layer, significantly improving the protected AI model's ability to identify adversarial attacks. Specifically, the protected AI model can both process input prompts normally and produce high-quality output when subjected to normal input, and accurately identify harmful content and reject output when subjected to adversarial attacks.
[0037] The beneficial effects of the present invention are:
[0038] (1) The present invention allows the sample model to process the input prompt first and generate a pre-answer, thereby preliminarily parsing the complex adversarial attack into specific steps that are easier to handle.
[0039] (2) The filter of the present invention distinguishes pre-answers, fully utilizing the safety guardrail of the artificial intelligence model, rejecting harmful prompts in advance, improving computing efficiency, and excluding a large amount of continuous repeated content, thereby ensuring the quality of pre-answers.
[0040] (3) The present invention summarizes the pre-answers through an inductive model, which not only ensures that benign prompts are not misjudged as adversarial attacks, but also accurately identifies the attack purpose of the adversarial attack and completes the analysis of the adversarial attack.
[0041] (4) The present invention accurately distinguishes between benign summaries and harmful summaries through a security classifier, thereby improving the ability of protected large language models to identify adversarial attacks.
[0042] (5) The present invention can ensure that the large language model does not output content that violates ethics and laws under adversarial attacks, and has strong adaptability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in describing the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0044] Figure 1 This is a diagram showing the execution of the defense method of the present invention when the input is an adversarial attack. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] Adversarial attacks on large language models refer to attacks in which attackers bypass the model's security mechanisms through various means, such as white-box attacks (using internal model information such as gradients or parameters), black-box attacks (using only model output information, such as through template completion, prompt rewriting, or generation based on large models), and prompt injection attacks (carefully designed input prompts to manipulate model outputs), thereby inducing it to generate content that violates usage guidelines, ethics, or laws, which may lead to serious consequences such as data leakage, privacy infringement, and the spread of false information.
[0047] In response to the various attack methods mentioned above, this embodiment provides an artificial intelligence model security defense method against adversarial attacks, which filters the adversarial attacks layer by layer through filters, inductive models and security classifiers to ensure the integrity and security of the protected large language model. Figure 1 As shown, the method includes the following steps:
[0048] Step S1: Before prompting for the protected AI model, the prompt is input into the sample model to form a pre-answer and transmit it to the filter.
[0049] Sample models selected include currently open-source Qwen2.5 Omni, Llama 3.2 Vision, and Molmo, which have undergone security alignment and are smaller, multimodal models with safety guardrails. The key to selecting sample models is their ability to easily respond to adversarial attacks and generate corresponding detailed harmful content.
[0050] Safety alignment begins during model training and continues throughout the model's lifecycle, including training, fine-tuning, deployment, and operation. During the training phase, safety alignment is fundamental to ensuring that model behavior aligns with intended goals and ethical principles. Its primary goal is to ensure that the model's generated content and behavior align with human values and intentions through methods such as data selection, training objective optimization, and reinforcement learning.
[0051] Safety guardrails refer to specific measures or mechanisms designed and implemented within AI models to limit or constrain certain system behaviors to prevent harmful or uncontrollable outcomes. Safety alignment is a broader concept that requires a comprehensive approach with multiple strategies and measures, of which safety guardrails are a part.
[0052] A prompt is a piece of text or sentence that initiates and guides an AI model to generate output of a specific type, theme, or format. It can be a question, description, keyword, context, etc. The sample model processes the prompt and generates a corresponding pre-answer. Pre-answers can take the following forms:
[0053] 1. When the input is a harmful prompt or adversarial attack:
[0054] Case 1: The sample model correctly identified harmful prompts or adversarial attacks, and the safety guardrail prompted the sample model to generate pre-answers containing safety prompts such as "sorry", "I can't help", and "I can't provide".
[0055] Case 2: The sample model is unable to parse the adversarial attack and generates a pre-answer containing a large amount of continuous repetition.
[0056] Case 3: Harmful prompts or adversarial attacks are not blocked by the sample model's safety guardrails, and the sample model generates a pre-answer containing specific and detailed harmful content.
[0057] 2. When the input is a benign prompt:
[0058] The sample model generated specific pre-responses that were consistent with legal and ethical norms.
[0059] Step S2: The filter receives the pre-answer of the sample model and performs preliminary filtering on the pre-answer; determines whether the safety fence of the sample model blocks the prompt, if not, transmits the pre-answer to the inductive model, if blocked, refuses to output.
[0060] A filter is used to check whether the pre-answer contains a safety word or a large amount of consecutive repetitive content. If the pre-answer contains specific and detailed harmful content (Case 3), proceed to S3 and the following steps. If the pre-answer contains a safety word (Case 1) or contains a large amount of consecutive repetitive content (Case 2), the prompt is directly judged as harmful.
[0061] Step S3: Fine-tune BART into an inductive model. The inductive model receives the pre-answer transmitted by the filter, processes the pre-answer, generates a summary, and transmits the summary to the security classifier.
[0062] The inductive model is generated by fine-tuning a Seq2Seq model (BART, for example). The Seq2Seq model is an end-to-end learning framework that uses an encoder to convert the input sequence into a fixed-length context vector. The decoder then uses this context vector to generate the target sequence. A key feature of this model is that the input and output sequence lengths can be inconsistent.
[0063] BART (Bidirectional and Auto-Regressive Transformers) is a pre-trained language model based on the Transformer architecture, combining the features of a bidirectional encoder and an autoregressive decoder. BART is pre-trained using a denoising autoencoder task, which involves adding noise (such as masking, deletion, or shuffling) to the input text and then letting the model learn to recover the original text. This design enables BART to perform well in text generation tasks (such as summarization, translation, and dialogue generation), and is also suitable for text understanding tasks.
[0064] Step S31: Fine-tune BART into an inductive model.
[0065] Step S311: randomly select N benign prompts and harmful prompts to form a prompt data set; preprocess the data set to obtain an input ID sequence and an output ID sequence.
[0066] Randomly selected from the Alpaca dataset Benign prompts and harmful prompts constitute the prompt dataset ; Prompt data set Expressed as: ; Indicates the Tips. The Alpaca dataset was created by the Tatsu-Lab team at Stanford University. It contains 52,000 tips.
[0067] Feed the prompt dataset into a large language model that has not undergone secure alignment (e.g. WizardLM-13B-Uncensored) to generate the answer dataset , answer data set Expressed as: ; Indicates the Answers. Answers With the Tips One to one correspondence.
[0068] For the model to correctly process text, it needs to be preprocessed to convert it into a sequence of IDs. In natural language processing, an ID sequence refers to the result of converting text data into a numerical form that the model can process. Specifically, a tokenizer segments the text into a series of tokens, then maps each token to a unique integer ID, forming a sequence of integers. This sequence can then be directly processed by the model.
[0069] Tokenization is the process of breaking text into small units. These units can be words, phrases, or characters. After tokenization, each token is mapped to a unique integer ID. This mapping is usually done using a pre-trained vocabulary. A vocabulary is a table that maps tokens to IDs. For example:
[0070] hello”
[0071] world”
[0072] Each group of answers and tips Defined as answer prompt combination .answer After being split by the word segmenter Marker composition, prompt After being split by the word segmenter The answer is and tips After being segmented by BART's word segmenter, it is represented as:
[0073]
[0074]
[0075] BART will answer after word segmentation and tips Encoded into the corresponding ID sequence. With input ID sequence Corresponding, enter the ID sequence It is expressed as follows:
[0076]
[0077] hint With output ID sequence Correspondingly, output ID sequence It is expressed as follows:
[0078]
[0079] Step S312: In the forward propagation phase, the input ID sequence Input encoder to get context representation ; Then the context is expressed as Input decoder to generate target sequence .
[0080] BART's encoder will input the ID sequence Encoded as a context representation, it is represented as follows:
[0081]
[0082] in, Indicates an encoder.
[0083] The encoder is a multi-layer Transformer structure, and the output of each layer can be expressed as:
[0084]
[0085] in, Represents the encoder Transformer model layer, Indicates the number of layers, Represents the encoder Transformer model The contextual representation of the layer output, Represents the encoder Transformer model The context representation of the layer output, the initial input of the encoder is represented as , the initial input of the encoder is the input ID sequence, that is .
[0086] BART's decoder represents the context Generate target sequence , which is expressed as follows:
[0087]
[0088] in, Represents a decoder.
[0089] The decoder is also a multi-layer Transformer structure, and the output of each layer can be expressed as:
[0090]
[0091] in, Represents the decoder Transformer model Layer, k represents the number of layers, Represents the decoder Transformer model The output of the layer, Represents the decoder Transformer model The output of the layer, Represents the decoder Transformer model The context representation of the layer output, the initial input of the decoder is represented as .
[0092] Step S313: Calculate the generated target sequence target The real target sequence The cross entropy loss between and is used to obtain the inductive model loss function.
[0093]
[0094] in, Inductive model loss function, Target sequence The corresponding length during the generation process, Is the generated length less than The target sequence, The decoded length is less than The context representation of The length of the real sequence is part. The decoder generates the target sequence Conditional probability, cross entropy loss It measures the difference between the predicted probability distribution and the true label.
[0095] Step S314: Back propagation, updating the model weights by calculating the gradient of the loss function with respect to the model parameters.
[0096]
[0097] in, represents the model parameters, represents the gradient operator.
[0098] Step S315: Parameter update: Use the Adam optimizer to update the model parameters according to the gradient.
[0099]
[0100] in, represents the updated model parameters, represents the model parameters before updating, is the learning rate.
[0101] Step S316: Update the model parameters through back propagation and Adam optimizer to obtain a fine-tuned inductive model.
[0102] In summary, the steps for fine-tuning the inductive model are as follows: First, a prompt dataset consisting of benign and harmful prompts is randomly selected from the Alpaca dataset, or an answer dataset is generated using a large language model that has not undergone secure alignment. Next, the text data is preprocessed and converted into a sequence of IDs. In the forward propagation phase, the input ID sequence is fed into the BART encoder and decoder to generate the target sequence. The cross-entropy loss between the generated sequence and the true sequence is then calculated, and the model parameters are updated through backpropagation and the Adam optimizer, completing the fine-tuning of the inductive model.
[0103] Step S32: Use the fine-tuned inductive model to process the pre-answer, generate a summary, and transmit the summary to the security classifier.
[0104] Step S4: Fine-tune BERT into a security classifier; the security classifier receives the summary and judges it. If it is judged to be harmful, it will refuse to output it. If it is judged to be harmless, it will allow the protected artificial intelligence model to process the prompt.
[0105] The security classifier is generated by fine-tuning the BERT model. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture. It was proposed by Google in 2018 and is primarily used for natural language processing (NLP) tasks.
[0106] BERT's core feature is its bidirectionality. By considering both left and right context during pre-training, it can better understand the meaning of words in a sentence. For example, in the sentence "I love natural language processing," BERT can use both "I love" and "language processing" to understand the meaning of "natural." It is widely used in natural language processing tasks such as text classification, question-answering systems, named entity recognition, sentiment analysis, and machine translation, and can be adapted to a variety of specific scenarios through fine-tuning.
[0107] Step S41: Fine-tune BERT into a security classifier.
[0108] Step S411: Randomly selected benign and harmful prompts as training sets , and the training set Label and obtain the labeled training set .
[0109] Randomly selected from the Alpaca dataset benign prompts and harmful prompts as training set. Tips are manually labeled, with benign tips marked as 1 and harmful tips marked as 0. The selected tips are combined into a dimensional training set ;in, Indicates the Tips. Corresponding annotations A set of 1 or 0 is called a dimensional vector = .
[0110] Step S412: The training set Input BERT and get the feature vector after BERT encoding. It is expressed as follows:
[0111]
[0112] in, For the training set go through Encoded feature vector; Represents the encoder of BERT.
[0113] Step S413: Based on the output of BERT, add a classifier layer to transform the feature vector Mapped to a harmless score. The classifier layer is usually a linear layer and can be expressed as:
[0114]
[0115] in, is the score matrix, is the weight matrix, is the bias vector.
[0116] Score Matrix yes dimensional vector, representing the eigenvector The harmless score corresponding to each element in the fine-tuning process , score matrix can be expressed as .
[0117] Weight Matrix It is the core part of the classifier. The weight matrix defines the linear relationship between the input features and the output results. yes During fine-tuning, the weight matrix Adjustment is performed by optimizing the loss function.
[0118] Bias vector is a A constant vector of dimension , the bias vector allows the model to be in the weight matrix The existence of the bias term enables the model to better adapt to the distribution of data. Even when the weight matrix is fixed, the classification effect can be optimized by adjusting the bias term.
[0119] Specifically, through the gradient descent algorithm, according to the security classifier loss function and The gradient of the loss function is used to update its value to minimize the difference between the predicted value and the true value. The gradient descent algorithm is an optimization algorithm used to minimize the loss function by iteratively updating the model parameters. The core idea is to update the parameters along the negative gradient of the loss function, thereby gradually reducing the value of the loss function.
[0120] Step S414: During the forward propagation process, the security classifier loss function is used to optimize the model parameters. The security classifier loss function can be expressed as:
[0121]
[0122] in, Indicates a prompt The corresponding annotations, Is the Sigmoid function, used to convert Mapped to between 0 and 1, .
[0123]
[0124] Step S415: Initialize the parameters first in the back propagation phase: Randomly initialize the weight matrix and the bias vector . Then calculate the security classifier loss function About the weight matrix Gradient and about Gradient , update the model parameters.
[0125] Update parameters: Update according to gradient and :
[0126]
[0127]
[0128] in, represents the updated weight matrix, represents the weight matrix before update, represents the updated bias vector, represents the bias vector before updating, Represents the loss function About the weight matrix before updating The gradient, Represents the loss function About the bias vector before updating The gradient, is the learning rate, which is a hyperparameter that determines the step size of each update.
[0129] Step S416: Optimize the model parameters through the security classifier loss function and use back propagation to update the parameters to obtain a fine-tuned security classifier.
[0130] In summary, in step S4, a safety classifier is constructed by fine-tuning the BERT model. First, a training set is selected and labeled with benign and harmful cues from the dataset. This training set is then fed into the BERT model to obtain a feature vector. A classifier layer is then added to this layer to map the feature vector into a harmlessness score. The model parameters are optimized using the safety classifier loss function and updated using backpropagation. Ultimately, the safety classifier is able to discriminate between the input summaries, rejecting the output if deemed harmful and allowing subsequent model processing if deemed harmless.
[0131] Step S42: The security classifier receives the summary and judges it. If it is judged to be harmful, the output is rejected. If it is judged to be harmless, the protected artificial intelligence model is allowed to process the prompt.
[0132] The prompts corresponding to the summaries judged as harmless by the security classifier will be input into the protected model, and the prompts judged as harmful in any of the above steps will be blocked.
[0133] The foregoing description is a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Instead, the present invention can be used in other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.
Claims
1. A security defense method for artificial intelligence models against adversarial attacks, characterized in that: The following steps are involved: Step S1: Before prompting for the protected AI model, the prompt is input into the sample model to form a pre-answer and transmit it to the filter; Step S2: The filter receives the pre-answer of the sample model and performs preliminary filtering on the pre-answer; determines whether the safety fence of the sample model blocks the prompt; if not, transmits the pre-answer to the inductive model; if blocked, refuses to output; Step S3: Fine-tune BART into an inductive model. The inductive model receives the pre-answer transmitted by the filter, processes the pre-answer, generates a summary, and transmits the summary to the security classifier. Step S4: Fine-tune BERT into a security classifier; The security classifier receives the summary and judges it. If it is judged to be harmful, it will refuse to output it. If it is judged to be harmless, it will allow the protected AI model to process the prompt; Step S4 includes the following steps: Step S41: Fine-tune BERT into a security classifier; Step S411: Randomly selected benign and harmful prompts as training sets , and the training set Label and obtain the labeled training set ; Step S412: The training set Input BERT and obtain the feature vector after BERT encoding; Step S413: Based on the output of BERT, add a classifier layer to transform the feature vector Mapping to harmless scores; Step S414: Using the security classifier loss function to optimize model parameters during forward propagation; Step S415: Initialize the parameters first in the back propagation phase: Randomly initialize the weight matrix and the bias vector ; Then calculate the security classifier loss function About the weight matrix Gradient and about Gradient , update model parameters; Step S416: Optimize the model parameters using the security classifier loss function and update the parameters using backpropagation to obtain a fine-tuned security classifier. Step S42: The security classifier receives the summary and judges it. If it is judged to be harmful, the output is rejected. If it is judged to be harmless, the protected model is allowed to process the prompt.
2. The artificial intelligence model security defense method against adversarial attacks according to claim 1 is characterized in that: The sample model of step S1 is an open source artificial intelligence model with a safety guardrail.
3. The artificial intelligence model security defense method against adversarial attacks according to claim 1 is characterized in that: Step S3 includes the following steps: Step S31: fine-tune BART into an inductive model; Step S311: randomly selecting N benign prompts and harmful prompts to form a prompt data set; preprocessing the data set to obtain an input ID sequence and an output ID sequence; Step S312: In the forward propagation phase, the input ID sequence Input encoder to get context representation ; Then the context is expressed as Input decoder to generate target sequence ; Step S313: Calculate the generated target sequence target The real target sequence The cross entropy loss between them gives the inductive model loss function; Step S314: Back propagation, updating the model weights by calculating the gradient of the inductive model loss function with respect to the model parameters; Step S315: Parameter update: Use the Adam optimizer to update the model parameters according to the gradient; Step S316: Update the model parameters through back propagation and Adam optimizer to obtain a fine-tuned induction model; Step S32: Use the fine-tuned inductive model to process the pre-answer, generate a summary, and transmit the summary to the security classifier.
4. The artificial intelligence model security defense method against adversarial attacks according to claim 3 is characterized in that: The inductive model loss function is expressed as: in, Inductive model loss function, Refers to the sequence The corresponding number of tags, Target sequence The corresponding length during the generation process, Is the generated length less than The target sequence, The decoded length is less than The context representation of The length of the real sequence is part; The decoder generates the target sequence The conditional probability of .
5. The artificial intelligence model security defense method against adversarial attacks according to claim 4 is characterized in that: The security classifier loss function is expressed as: in, represents the security classifier loss function, Indicates a prompt The corresponding annotations, Indicates harmless score, Is the Sigmoid function, used to convert the harmless score Mapped to between 0 and 1.
Citation Information
Patent Citations
Machine translation model high semantic similarity antagonism sample generation method
CN116757223A
Network security defense method and system for autonomously identifying attack monitoring
CN118300889A