Web defense method for evasive attack based on adversarial training
Adversarial training technology parses HTTP requests to generate token sequences, generates adversarial samples and trains malicious web request detectors, solving the problem that existing WAFs are difficult to defend against complex adversarial attacks, and achieving more efficient web attack detection and defense.
Patent Information
- Application Number
- CN202510503683.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing Web Application Firewalls (WAFs) are difficult to effectively defend against complex adversarial attacks, and signature-based defense mechanisms cannot cope with dynamically adjusted attack payloads.
Adversarial training-based web defense method is adopted to generate token sequences by parsing HTTP requests, generate adversarial samples, and train malicious web request detectors to identify and defend against evasive attacks.
It significantly improves WAF's defense ability against complex adversarial attacks, reduces false positive rates, reduces dependence on domain expert knowledge, and improves the robustness of web attack detection.
Smart Images

Figure CN120050115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer network security, and specifically to a Web defense method against evasion attacks based on adversarial training. Background Art
[0002] Currently, various network platforms, whether it is social websites, blogs, online banks or e-commerce platforms, have become the main places for online information transmission and service delivery. However, these applications have become the key targets of attackers because they involve the processing of sensitive data.
[0003] During the web security scanning and assessment between 2021 and 2022, researchers found that more than half of the online applications had high-risk security vulnerabilities, such as remote code execution, Cross-Site Scripting (XSS), and SQL Injection (SQLi). These vulnerabilities not only reduce the performance of Web applications, but may also lead to the leakage of user privacy data, thus bringing serious economic losses to users and network service providers. Ensuring the security of Web applications has become an extremely challenging task.
[0004] In this context, the importance of the Web Application Firewall (WAF) as a key line of defense for protecting the internal information network of an organization has become even more prominent. According to the latest application security report released by Cloudflare, products related to WAF accounted for as high as 53.9% of all mitigated traffic (i.e., malicious or unnecessary network traffic blocked or reduced through security measures), and the call frequency of WAF-related APIs accounted for even 66.59%. This data fully highlights the key position of Web attack detection in network security protection.
[0005] In the field of Web attack detection, the failure of traditional WAF systems is mainly due to the lack of domain experts, the limitations of classification technology, and the constraints of false positive problems. Web attack detection systems rely heavily on the expertise of security experts to determine which key features to extract from Web application requests, network packets, or other inputs. However, due to the high demand and low threshold of the software industry, many Web developers and network administrators lack the necessary Web security knowledge and find it difficult to effectively implement related technologies. In addition, in order to distinguish legitimate requests from attack requests, many systems rely on rule-based technologies or supervised machine learning algorithms, which require a large amount of labeled training data. For customized applications, it is difficult and costly to obtain such data, and the labeled data is often biased, which poses a challenge to the accuracy of the classifier. With the continuous emergence of new attacks and vulnerabilities, rule-based or supervised learning technologies may not be able to effectively identify these new threats, resulting in an increase in misclassification. At the same time, although unsupervised learning technologies have been widely studied in Web attack detection, these methods require manual identification of specific attack features and have a high false positive rate in practical applications. For example, Web attack detection systems may mistakenly label a large number of legitimate users as attackers, and the increase in false positive rates will have a serious impact on user experience and system credibility. Therefore, reducing the false positive rate is the key to improving the performance of Web attack detection systems.
[0006] Although WAFs have been widely deployed, in recent years, attackers have developed a series of sophisticated techniques that can effectively bypass WAF defense mechanisms. The core of these techniques is to transform malicious payloads into obfuscated or mutated forms (i.e., obfuscated payloads) to mislead the detection mechanisms in WAFs. Current research focuses on two methods: one is mutation-based methods, which modify existing payloads through carefully designed mutation strategies; the other is generation-based methods, which generate test inputs from scratch by designing payload generation strategies and grammars. At present, a common strategy to improve the protection capabilities of WAFs is to extract new attack signatures based on successfully bypassing attack payloads. However, this signature-based defense mechanism has obvious defects: on the one hand, it cannot effectively defend against black-box attacks because defenders cannot obtain specific information about the attack payload; on the other hand, it is difficult to cope with evolving adversarial attacks, whose payloads can be dynamically adjusted based on the feedback from the WAF, allowing attackers to quickly discover vulnerabilities not covered by the rules. In recent years, many adversarial attack methods targeting the web have emerged. These methods modify the original request payload based on the feedback from the defender, trying to bypass WAFs while retaining their malicious functions.
[0007] Therefore, how to enhance WAF's defense capabilities against complex adversarial attacks has become a key issue that needs to be solved urgently. Summary of the invention
[0008] The object of the present invention is to solve the defect that it is difficult to effectively guarantee Web defense in the prior art, and to provide a Web defense method for evasion attacks based on adversarial training to solve the above problems.
[0009] To achieve the above object, the technical solution of the present invention is as follows:
[0010] A Web defense method for evasion attacks based on adversarial training, comprising the following steps:
[0011] Web request payload parsing: Parse the HTTP request into a token sequence;
[0012] Generation of adversarial samples: Rapidly generate adversarial samples based on gradients;
[0013] Adversarial training: Train a malicious Web request detector based on the original samples and adversarial samples;
[0014] Real-time malicious Web request detection: Deploy the trained malicious Web request detector to the web application server to monitor Web requests in real time, collect HTTP / HTTPS traffic and perform preprocessing, the preprocessing includes application layer payload parsing and word segmentation operations, and the detection engine of the web application server outputs the detection classification result.
[0015] The Web request payload parsing includes the following steps:
[0016] Take the payload string in the HTTP request as the input text for preprocessing;
[0017] Load the SentencePiece model, and use the pre-trained SentencePiece model to perform word segmentation on the preprocessed text;
[0018] Use the SentencePiece model to split the text into sub-word units;
[0019] After the word segmentation process, convert the sub-word units into a sequence of tokens, and each token corresponds to a sub-word unit or a special token;
[0020] Add special tokens: The parsed token sequence starts with the [CLS] token and ends with the [SEP] token;
[0021] Output the structured token sequence: Output the token sequence after word segmentation as a structured representation form , where is a token, is the last token of the token sequence i.e., the word.
[0022] The generation of the adversarial sample includes the following steps:
[0023] Perform gradient-based word importance ranking;
[0024] Perform DistilBERT semantic text similarity constraint.
[0025] The adversarial training includes the following steps:
[0026] Use the original sample and the adversarial attack sample to train the BERT model. The training objective is to minimize the loss of the original sample and the adversarial attack sample, that is:
[0027] ,
[0028] where represents the process of generating the adversarial attack sample, is used to measure the adversarial loss. Set = 1 to balance the two losses, represents the parameters of the classification network, respectively represent the original sample and its label, E(.) represents the expectation function, and L(.) is the cross-entropy loss of the classification model;
[0029] The specific steps of the adversarial training are as follows:
[0030] Forward propagation and loss calculation of the original data,
[0031] Sample a batch of original samples x and their labels y from the training set D, and calculate the original cross-entropy loss ,
[0032] ,
[0033] N represents the number of samples, are the parameters of the classification model, is the classification label, 0 represents normal, and 1 represents attack, is the original sample is correctly classified as the probability of the class;
[0034] Forward propagation and loss calculation of the adversarial sample,
[0035] Input the adversarial attack sample Put into the BERT model and calculate the adversarial loss :
[0036] ,
[0037] N represents the number of samples, is a parameter of the classification model, is the classification label, 0 indicates normal, and 1 indicates an attack, is the adversarial attack sample is correctly classified as the probability of the class;
[0038] Joint loss and parameter update, calculate the total loss :
[0039] ,
[0040] Subsequently, calculate the gradient of the total loss, use the optimizer to backpropagate the gradient, and update the model parameters .
[0041] The gradient-based word importance ranking described above includes the following steps:
[0042] Detect the model forward propagation to calculate the gradient,
[0043] Use the BERT model to perform forward propagation on the sample, that is, the token sequence after word segmentation to calculate the cross-entropy loss function of the classification model as follows:
[0044] ,
[0045] where, represents the number of samples, is the original sample 's label, is a parameter of the classification model, is the probability that the sample is correctly classified as the class, and the token sequence is the original sample ;
[0046] Calculate the word importance,
[0047] Calculate the importance of each word in the token sequence , and the formula is:
[0048] ,
[0049] where, represents the L1 norm, is 's gradient, is the word embedding corresponding to the word ;
[0050] For the BERT model, the input is tokenized into sub-words, and the importance of each word is calculated by taking the average of all sub-words that make up the word;
[0051] Word importance ranking
[0052] According to the value, the words are sorted, and the top m words in the importance ranking are preferentially replaced to generate adversarial samples; For the BERT model, the input text is tokenized into sub-words. If a word consists of multiple sub-words, the average of the gradients of these sub-words is calculated as the importance of the word .
[0053] The above-mentioned DistilBERT semantic text similarity constraint includes the following steps:
[0054] Generate perturbed text
[0055] During the process of generating adversarial samples, the token sequence after word segmentation is used as the original text for perturbation, replacing the words with the top m importance rankings to generate candidate perturbed texts; For each token to be replaced, its Top-k similar words are retrieved from the DistilBERT word embedding space, and replacement candidates are randomly sampled to finally construct the perturbed text
[0056] Semantic encoding calculation, using the DistilBERT model to encode the original text and the perturbed text to obtain semantic vectors and respectively;
[0057] Cosine similarity calculation, calculate and the cosine similarity between them, and the formula is:
[0058] ,
[0059] where represents the L2 norm;
[0060] Similarity threshold constraint
[0061] If the cosine similarity is lower than the preset threshold, the perturbed text is discarded, considering that its semantics is too different from the original text ; Otherwise, keep as an adversarial attack sample.
[0062] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, a Web defense method against evasion attacks based on adversarial training is implemented.
[0063] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, a Web defense method against evasion attacks based on adversarial training is implemented.
[0064] Beneficial effects
[0065] Compared with the prior art, the Web defense method against evasion attacks based on adversarial training of the present invention uses a new adversarial attack to quickly generate adversarial examples without knowing the attacker's mutation strategy, detect malicious Web access requests, and thus greatly improve the robustness of the WAF to defend against Web attacks.
[0066] The present invention regards each Web request as a sentence and uses a Transformer-based language model to understand the hidden intention behind each web request. In addition, the present invention quickly generates adversarial samples with perturbations based on gradients to simulate the mutations of attackers. The present invention uses natural language processing technology to identify and mitigate unknown malicious payloads in the WAF, greatly reducing the dependence of the security operation center on domain expert knowledge. Description of the drawings
[0067] Figure 1 It is a sequence diagram of the method of the present invention;
[0068] Figure 2 It is a framework diagram of the method of the present invention;
[0069] Figure 3 It is an example diagram of payload parsing involved in the present invention;
[0070] Figure 4 It is an example diagram of the implementation manner of the present invention. Detailed implementation manner
[0071] To further understand and recognize the structural features and achieved effects of the present invention, the following is a detailed description in conjunction with preferred embodiments and drawings:
[0072] The present invention proposes a method for defending against evasion Web attacks based on adversarial training. The method framework is as Figure 2As shown in the figure. The present invention generally includes three stages: (1) Web request payload parsing: parsing the HTTP request into a sequence of tokens. Since the HTTP request is a protocol-compliant string, the present invention regards it as a sentence in the "HTTP request language", converting the unstructured request into a structured sequence of tokens for more efficient semantic analysis; (2) adversarial sample generation: quickly generating adversarial samples based on gradients; (3) adversarial training: training a malicious Web request detector based on the original samples and adversarial samples, enabling the detector to robustly defend against adversarial Web attacks.
[0073] As Figure 1 shown, a Web defense method based on adversarial training against evasion attacks according to the present invention includes the following steps
[0074] The first step, Web request payload parsing: parsing the HTTP request into a sequence of tokens. The initial stage of the present invention is to disassemble the payload string in the HTTP request into a series of tokens. In the framework of this research, the HTTP payload is regarded as a series of lexical units and tokenized by means of the SentencePiece model. SentencePiece is an advanced subword tokenization tool developed by Google and is widely used in the preprocessing of natural language processing tasks. By splitting the text into subword units, this tool breaks through the limitations of traditional word- or character-based tokenization methods and effectively alleviates the out-of-vocabulary (OOV) problem. In addition, since SentencePiece does not rely on a predefined vocabulary, it can flexibly process texts in multiple languages, demonstrating extremely high versatility. Figure 3 Shows an example of payload parsing for a MySQL SELECT query. The parsed token sequence starts with a special [CLS] token and ends with a [SEP] token. Through this processing method, SentencePiece can provide an accurate structured representation for the semantic analysis of HTTP requests, providing strong support for subsequent NLP model training and adversarial attack detection.
[0075] The specific steps are as follows:
[0076] (1) Using the payload string in the HTTP request as the input text for preprocessing, and the preprocessing includes removing unnecessary spaces, special characters or other noises to ensure the purity of the text.
[0077] (2) Load the SentencePiece model and use the pre-trained SentencePiece model to tokenize the preprocessed text. The SentencePiece model breaks through the limitations of traditional word- or character-based tokenization methods by splitting text into subword units, effectively alleviating the out-of-vocabulary (OOV) problem. In addition, since SentencePiece does not rely on a predefined vocabulary, it can flexibly handle texts in multiple languages, demonstrating extremely high generality.
[0078] (3) Use the SentencePiece model to split the text into subword units. SentencePiece splits the input text into a series of subword units through statistical learning methods. These subword units can be parts of words (such as prefixes, suffixes, or roots) or complete words. For example, Figure 3 the word "1966yesterday" in
[0079] may be split into two subword units: "1966" and "yesterday". Figure 3 shows an example of payload parsing for a MySQL SELECT query.
[0080] (5) Add special tokens: The parsed token sequence starts with the [CLS] token and ends with the [SEP] token. In this way, SentencePiece can provide an accurate structured representation for the semantic analysis of HTTP requests, providing strong support for subsequent NLP model training and adversarial attack detection.
[0081] (6) Output the structured token sequence: Output the token sequence after tokenization as a structured representation , where is the token, is the token sequence and
[0082] is the last token of the token sequence
[0083] Here, the training data is augmented using adversarial examples generated by perturbing the training data in the input space, and then the original samples and the adversarial samples are used to train the classification model of the present invention. In previous work on adversarial training, especially in the field of computer vision, adversarial examples were typically generated between each mini-batch and used to train the model. However, in practice, when using NLP adversarial attacks, it is difficult to generate adversarial examples between each mini-batch update. This is because NLP adversarial attacks usually require other neural networks as their sub-components (such as sentence encoders, masked language models, BERT models, etc.). Since Transformer-based models such as BERT and RoBERTa models also require a large amount of GPU memory to store the computational graph during training, it is impossible to run the adversarial attack and train the model on the same GPU. Therefore, the present invention maximizes GPU utilization by first generating adversarial examples before each training batch and then using the generated samples to train the model.
[0084] The present invention constructs two key steps when building the attack to achieve accelerated generation of adversarial samples. First, words are sorted based on gradients, and then adversarial samples are generated based on semantic text similarity constraints.
[0085] (1) Perform gradient-based word importance ranking. The generation of adversarial samples in previous work generally generates adversarial examples by iteratively replacing a single word in the original text at a time. To determine the order of the words to be replaced, they rank the words by the change in the confidence of the target model in the ground truth label when a word is removed from the input. The present invention refers to this as deletion-based word importance ranking. One problem with this method is that an additional model forward pass must be performed for each word to calculate its importance. For long text inputs, this may mean that the model must perform up to hundreds of forward passes to generate a single adversarial example. Instead, the present invention uses the gradient of the classification loss function to determine the importance of each word.
[0086] A1) Detect the model forward propagation to calculate the gradient,
[0087] Use the BERT model to perform a forward pass on the sample, that is, the token sequence after word segmentation to calculate the cross-entropy loss function of the classification model as follows:
[0088] ,
[0089] where, represents the number of samples, is the original sample 's label, are the parameters of the classification model, The probability of correctly classifying the sample into the category, and the token sequence is the original sample ;
[0090] A2) Calculate the word importance,
[0091] Calculate the importance of each word in the token sequence with the formula: , where
[0092] ,
[0093] where represents the L1 norm, is the gradient of and is the word embedding corresponding to the word
[0094] For the BERT model, the input is tokenized into subwords, and the importance of each word is calculated by taking the average of all subwords that make up the word; in the present invention, the importance of each word is calculated by taking the average of all subwords that make up the word. This only requires one forward and backward pass and saves having to make an additional forward pass for each word. Previous studies have shown that the gradient ranking method is the fastest search method and provides a competitive attack success rate compared to deletion-based methods.
[0095] A3) Rank the word importance,
[0096] Rank the words according to the values, and the top m words in terms of importance ranking are preferentially replaced to generate adversarial samples; for the BERT model, the input text is tokenized into subwords, and if the word is composed of multiple subwords, then the average of the gradients of these subwords is calculated as the importance of the word .
[0097] (2) Apply the DistilBERT semantic text similarity constraint.
[0098] Adversarial attack research typically uses Universal Sentence Encoders (USE) to compare the sentence encodings of the original text and the perturbed text. If the cosine similarity between the two encodings is below a certain threshold, the perturbed text is ignored, thereby constraining the generated adversarial samples to still retain the attack semantics of the original text. One of the challenges of using a large encoder like USE is that it may consume a large amount of GPU memory. This invention uses the DistilBERT model as its constraint module instead of using USE. This is because DistilBERT requires 10 times less GPU memory than USE and fewer operations.
[0099] B1) Generate perturbed text.
[0100] During the process of generating adversarial samples, the token sequence after tokenization is used as the original text for perturbation, replacing the words with the top m importance rankings to generate candidate perturbed texts; for each token to be replaced, retrieve its Top-k similar words from the DistilBERT word embedding space and randomly sample to generate replacement candidates, and finally construct the perturbed text.
[0101] B2) Semantic encoding calculation. Use the DistilBERT model to encode the original text and the perturbed text to obtain semantic vectors and respectively;
[0102] B3) Cosine similarity calculation. Calculate the cosine similarity between and , and the formula is: ,
[0103] ,
[0104] where represents the L2 norm;
[0105] B4) Similarity threshold constraint.
[0106] If the cosine similarity is below the preset threshold, discard the perturbed text , considering that its semantics is too different from the original text ; otherwise, retain as the adversarial attack sample.
[0107] Step 3, adversarial training: Train the malicious Web request detector based on the original samples and adversarial samples, minimizing the loss of the original training dataset and the loss of the adversarial examples. The present invention uses the original attack samples and the adversarial attack samples generated in the previous section to train the model, and the training objective is to minimize the loss of the original training dataset and the loss of the adversarial examples.
[0108] (1) Use the original samples and adversarial attack samples to train the BERT model, and the training objective is to minimize the loss of the original samples and the adversarial attack samples, that is:
[0109] ,
[0110] where, represents the process of generating adversarial attack samples, is used to measure the adversarial loss, and set = 1 to balance the two losses, represents the parameters of the classification network, respectively represent the original sample and its label, E(.) represents the expectation function, and L(.) is the cross-entropy loss of the classification model.
[0111] (2) The specific steps of adversarial training are as follows:
[0112] C1) Forward propagation and loss calculation of the original data,
[0113] Sample a batch of original samples x and their labels y from the training set D, and calculate the original cross-entropy loss ,
[0114] ,
[0115] N represents the number of samples, are the parameters of the classification model, is the classification label, 0 represents normal, 1 represents attack, is the original sample correctly classified as the probability of the class;
[0116] C2) Forward propagation and loss calculation of the adversarial samples,
[0117] Input the adversarial attack sample Put input into the BERT model and calculate the adversarial loss :
[0118] ,
[0119] N represents the number of samples, are the parameters of the classification model, For classification labels, 0 indicates normal and 1 indicates attack. Is an adversarial attack sample Correctly classified as Probability of the class;
[0120] C3) Joint loss and parameter update, calculate the total loss :
[0121] ,
[0122] Subsequently, calculate the gradient of the total loss, use the optimizer to backpropagate the gradient, and update the model parameters .
[0123] Fourth step, real-time malicious Web request detection: Deploy the trained malicious Web request detector to the web application server to monitor Web requests in real time, collect HTTP / HTTPS traffic and perform preprocessing. The preprocessing includes application layer payload parsing and tokenization operations. The detection engine of the web application server outputs the detection classification result.
[0124] The Web attack detection method based on adversarial training of the present invention can be deployed in the enterprise's network security protection system, especially in the detection module of WAF or the Security Operation Center (SOC), and is specifically used to detect and defend against attacks on Web applications, preventing attackers from bypassing the existing WAF protection mechanism. Figure 4 A specific implementation example is presented. By introducing adversarial training technology, it can actively simulate the behavior of attackers, generate adversarial samples, and integrate them into the training process of the Web attack detection model. This process mainly relies on the security defense chain of "attack sample → adversarial sample generation → model learning → attack detection" to specifically detect and defend against various attack types such as SQL injection, cross-site scripting, and command injection commonly found in Web applications. Through the collection of attack samples, the generation of adversarial samples, and the continuous optimization of the model, the characteristics and patterns of attack behaviors are extracted and integrated into the detection model. Finally, the detection results are output with high accuracy and low false alarm rate, helping the security team to detect and respond to Web attacks in a timely manner, ensuring the security and stability of Web applications, and providing strong support for the enterprise's network security protection.
[0125] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, various changes and improvements will occur to the present invention, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. A Web defense method for evasive attacks based on adversarial training, characterized in that: The following steps are involved: 11) Web request payload parsing: Parsing HTTP requests into token sequences; 12) Generation of adversarial samples: Rapidly generate adversarial samples based on gradients; 13) Adversarial training: training malicious web request detectors based on original samples and adversarial samples; 14) Real-time malicious web request detection: The trained malicious web request detector is deployed on the web application server to monitor web requests in real time, collect HTTP / HTTPS traffic and perform preprocessing. The preprocessing includes application layer load parsing and word segmentation operations. The detection engine of the web application server outputs the detection classification results.
2. The Web defense method for evasive attacks based on adversarial training according to claim 1, characterized in that: The web request payload parsing comprises the following steps: 21) Preprocess the payload string in the HTTP request as input text; 22) Load the SentencePiece model and use the pre-trained SentencePiece model to perform word segmentation on the preprocessed text; 23) Use the SentencePiece model to split the text into subword units; 24) After the word segmentation process is completed, the subword units are converted into a sequence of tokens, each token corresponding to a subword unit or a special tag; 25) Add special tags: The parsed token sequence starts with the [CLS] token and ends with the [SEP] token; 26) Output structured token sequence: Output the token sequence after word segmentation The output is a structured representation , in, is a token, is a sequence of tokens The last token of , which is the word.
3. The Web defense method for evasive attacks based on adversarial training according to claim 1, characterized in that: The generation of the adversarial sample includes the following steps: 31) Perform gradient-based word importance ranking; 32) Perform DistilBERT semantic text similarity constraints.
4. The Web defense method for evasive attacks based on adversarial training according to claim 1, characterized in that: The adversarial training comprises the following steps: 41) Use the original samples and adversarial attack samples to train the BERT model. The training goal is to minimize the loss of the original samples and adversarial attack samples, that is: , in, represents the process of generating adversarial attack samples, Used to measure adversarial loss, set =1Weighing the two losses, represents the classification network parameters, Represent the original sample and its label respectively, E(.) represents the expected function, and L(.) is the cross entropy loss of the classification model; 42) The specific steps of adversarial training are as follows: 421) Raw data forward propagation and loss calculation, Sample a batch of original samples x and their labels y from the training set D and calculate the original cross entropy loss , , N represents the number of samples, are the parameters of the classification model, is the classification label, 0 means normal, 1 means attack, For the original sample Correctly classified as Probability of the class; 422) Adversarial sample forward propagation and loss calculation, Input adversarial attack sample Will Input BERT model and calculate adversarial loss : , N represents the number of samples, are the parameters of the classification model, is the classification label, 0 means normal, 1 means attack, For adversarial attack samples Correctly classified as Probability of the class; 423) Combine loss and parameter update to calculate total loss : , The gradient of the total loss is then calculated and the optimizer is used to backpropagate the gradient to update the model parameters. .
5. The Web defense method for evasive attacks based on adversarial training according to claim 3, characterized in that: The gradient-based word importance sorting comprises the following steps: 51) Detection model forward propagation calculates gradient, Use the BERT model for samples, that is, the token sequence after word segmentation Perform forward propagation and calculate the cross entropy loss function of the classification model as follows: , in, represents the number of samples, The original sample Tags, are the parameters of the classification model, The sample is correctly classified as The probability of the category, the token sequence is the original sample ; 52) Calculate word importance, Calculate each token sequence Chinese words Importance , the formula is: , in, represents the L1 norm, for The gradient of is the word corresponding to word embeddings; For the BERT model, the input is tokenized into subwords and the importance of each word is calculated by taking the average of all the subwords that make up the word; 53) Word importance ranking, according to The words are sorted by the value of , and the top m words in importance are replaced first to generate adversarial samples; for the BERT model, the input text is marked as subwords, if the word If it is composed of multiple subwords, the average value of the gradients of these subwords is calculated as the word The importance of.
6. The Web defense method for evasive attacks based on adversarial training according to claim 3, characterized in that: The DistilBERT semantic text similarity constraint comprises the following steps: 61) Generate perturbation text, In the process of generating adversarial samples, the token sequence after word segmentation is , as the original text for perturbation, replacing the importance The top m ranked words are used to generate candidate perturbation texts; for each token to be replaced, its Top-k similar words are retrieved from the DistilBERT word embedding space, and replacement candidates are randomly sampled to finally construct the perturbation text. ; 62) Semantic encoding calculation, using DistilBERT model to calculate the original text and perturbation text Encode and obtain semantic vectors and ; 63) Cosine similarity calculation, calculation and The cosine similarity between , the formula is: , in, represents the L2 norm; 64) Similarity threshold constraint, If the cosine similarity is lower than a preset threshold, the perturbed text is discarded , which is considered to be semantically similar to the original text The difference is too large; otherwise, keep as adversarial attack samples.
7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program. When the computer program is executed by the processor, the Web defense method for evasive attacks based on adversarial training as described in any one of claims 1 to 6 is implemented.
8. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, a Web defense method for evasive attacks based on adversarial training as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Structured query language injection detection method
CN114598526A
XSS attack load automatic generation method and system based on generative adversarial network
CN116545767A
Character-level adversarial sample generation method using affix embedding space
CN117312955A
Robust training method based on phase upset
CN118051772A
SQL injection attack detection method and system and related equipment
CN119324840A
Cited By
Data management method based on sentence semantic perception and related equipment
CN120277117A
A data management method and related equipment based on sentence semantic perception
CN120277117B