Method and system for constructing large Web attack detection model
By building a big model of web attack detection, the problem of sample imbalance and lack of interpretability in web attack detection is solved, high-precision and low-cost web attack detection is achieved, and professional analysis support is provided.
Patent Information
- Application Number
- CN202510268435.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-18
AI Technical Summary
Existing Web attack detection methods face the problems of attack sample imbalance and lack of interpretability of detection results, and the fine-tuning of large language models is high, making it difficult to implement on consumer downgraded hardware.
By constructing a Web attack detection model, including data preparation and processing, expected detection results and interpretable analysis, instruction fine-tuning data set construction, and quantitative low-rank adaptation technology, the Chinese model is fine-tuned to generate a high-precision and interpretable detection model.
It realizes high-precision web attack detection, reduces model training and deployment costs, provides professional and comprehensive analysis results, and supports users to understand attack intentions and respond to decisions.
Smart Images

Figure CN120342652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of application-layer network attack detection, and particularly to a method and system for constructing a large model for Web attack detection. Background Art
[0002] With the rapid development of the Internet, network security issues have become increasingly prominent. As a core component of the Internet, Web applications carry a large amount of data transmission and key business logics, and have also become one of the main targets of hacker attacks. From SQL injection, cross-site scripting attack (XSS) to remote command execution, the forms of Web attacks have become increasingly diverse and complex, posing a major threat to the data security of enterprises and users. In recent years, the application of deep learning technology in Web attack detection has achieved certain results and has also become a current research hotspot. From convolutional neural network (CNN) to long short-term memory network (LSTM), these methods can discover some hidden patterns that cannot be captured by traditional methods through feature extraction and learning of network traffic. However, these methods face the following main challenges:
[0003] 1. Imbalance of attack samples: In actual scenarios, the proportion of attack traffic is extremely low compared to normal traffic. The model is easily affected by the data imbalance problem, resulting in unsatisfactory detection effects.
[0004] 2. Lack of interpretability of detection results: Traditional Web attack detection models can only output classification labels, such as normal or abnormal, etc. This is not conducive to users understanding the attack intent of Web attack requests, nor is it conducive to users locating system vulnerabilities and further response decisions.
[0005] The emergence of large language models provides new ideas for solving these problems. Large language models represented by GPT, BERT, etc. have demonstrated excellent performance in natural language processing tasks. These models can capture deep semantics and complex relationships in text sequences, providing potential advantages for network attack detection. Considering that Web traffic data (such as HTTP requests and logs) can be regarded as a kind of "language" to a certain extent. The information such as request headers, parameters, and request bodies contained in these data has grammatical structures and semantic features, and large language models are good at capturing these features. At present, researchers have explored the preliminary application of large language models in Web attack detection. For example, by converting HTTP requests into "sentences" and using BERT or GPT to classify them, potential malicious requests are detected. Some experiments show that this method is significantly superior to traditional machine learning models in terms of accuracy and recall. However, high computational cost is still a challenge faced by fine-tuning large language models. This study applies quantization low-rank adaptation technology to greatly reduce the training cost of fine-tuning large language models on the premise of ensuring detection accuracy, enabling this study to be implemented on consumer-grade degraded hardware. Summary of the Invention
[0006] The object of the present invention is to provide a method and system for constructing a large model for Web attack detection in view of the deficiencies of the prior art.
[0007] The object of the present invention is achieved by the following technical solutions:
[0008] A method for constructing a large model for Web attack detection includes the following steps:
[0009] (1) Obtain Web request data, and extract the request line, request headers, and request body from the Web request data; perform missing value processing on the request headers and URL decoding on the extracted data to obtain Web request samples;
[0010] (2) Generate expected detection results and interpretive analysis results according to the Web request samples obtained in step (1);
[0011] (3) Construct an instruction fine-tuning dataset according to the Web request samples obtained in step (1), the expected detection results and interpretive analysis results obtained in step (2);
[0012] (4) Use the fine-tuning strategy and the instruction fine-tuning dataset obtained in step (3) to fine-tune and train the Chinese large model to obtain a large model for Web attack detection.
[0013] Further, in the step (1), the sources of the Web request data include open-source datasets and locally built OWASP WebGoat penetration testing platforms; use Metasploit, BeFF, and sqlmap automated penetration tools to generate attack scripts, and then use the Burpsuites tool to capture packets for extraction; among the obtained Web request data, all sample data including normal requests, XSS attacks, SQLi, malicious URLs, and remote command executions are not less than 10,000, and the number of samples for each attack is not less than 1,000.
[0014] Further, in the step (1), the missing value processing of the request headers and URL decoding of the extracted data include: filling the missing values of the request headers with "NONE", and performing URL decoding on the URL part of the request line and the form data in the request body.
[0015] Further, in the step (2), for the samples obtained from the open-source dataset, the expected detection results are given by the labels of the dataset; for the samples recorded in the locally built penetration testing platform, the expected detection results are manually labeled; if it is a normal request, it is labeled as "no Web attack", and if it is an abnormal attack request, it is labeled with the corresponding attack type; the interpretive analysis results are the basis for sample discrimination, generated by manual analysis, including Web request analysis, security analysis, and security suggestions.
[0016] Further, in the instruction fine-tuning dataset of the step (3), the fine-tuning samples include a fine-tuning instruction part, an input part, and an output part; the fine-tuning instruction part is the system prompt word of the task, the input part is the Web request sample, and the output part is the expected label and interpretive analysis information of the Web request sample; the instruction fine-tuning dataset is saved in json format.
[0017] Further, in the step (4), the Chinese large models include Tongyi Qianwen and Baichuan large models, and the samples are adapted during training according to the input-output templates of the selected large models.
[0018] Further, in the step (4), the fine-tuning strategy is the quantization low-rank adaptation technology, which is used to reduce the training cost required for fine-tuning the model; after loading the base model, the parameters of the loaded base model are quantized to 4-bit, and all the linear layers of the model are traversed and trainable low-rank adaptation modules are inserted. During the training process, the weights of the base model are frozen, and only the parameters of the inserted low-rank adaptation modules are updated; after the fine-tuning training is completed, the weights of the low-rank adaptation modules are added and merged with the corresponding linear layer weight matrices of the base model.
[0019] A system for constructing a large model for Web attack detection includes:
[0020] A data preparation and processing module, which is used to obtain Web request data, extract the request line, request headers, and request body from the Web request data; perform missing value processing on the extracted request headers and URL decoding to obtain Web request samples;
[0021] An expected detection result and interpretive analysis module, which is used to generate expected detection results and interpretive analysis results according to the Web request samples;
[0022] An instruction fine-tuning dataset construction module, which is used to construct an instruction fine-tuning dataset according to the Web request samples, expected detection results, and interpretive analysis results;
[0023] A large model quantization fine-tuning module, which is used to fine-tune and train the Chinese large model by using the fine-tuning strategy and the instruction fine-tuning dataset to obtain a large model for Web attack detection.
[0024] The beneficial effects of the present invention are as follows:
[0025] 1. High detection accuracy: Web attacks often rely on complex request sequences or specific semantic patterns. The deep representation learning ability of large language models can capture the deep semantic relationships and context dependencies in language, so it can significantly improve the discrimination ability for complex features, accurately distinguish normal traffic from malicious traffic, and thus achieve high-precision detection of Web attacks.
[0026] 2. High generalization ability: Through pre-training on a large-scale corpus, large language models have learned rich general features. After fine-tuning with a small amount of data in downstream fields, the model can quickly adapt to different types of attack scenarios and still maintain high performance even in cases of scarce or imbalanced samples.
[0027] 3. Lower training cost: Fine-tuning the parameters of a super-large model with full precision requires high hardware costs and is difficult to achieve on consumer-grade hardware. This method applies the QLoRA fine-tuning training technology to compress the model weights from 32-bit floating-point numbers to 4-bit integers. The overall reduction in video memory requirements can reach about 70%, greatly reducing the occupancy of video memory. At the same time, by freezing the original model parameters and only fine-tuning the smaller-scale LoRA layers, the performance loss caused by quantized weights can be avoided.
[0028] 4. More professional and comprehensive analysis results: Traditional detection methods can only give classification labels and lack interpretive analysis of the results, making it difficult to provide further support for users' responses and decisions. The large model for Web attack detection of the present invention fully learns the interpretive analysis aligned with experts, can intelligently analyze Web request samples, simulate the judgment process of security experts, and output comprehensive and professional analysis bases including Web request analysis, security risk analysis, and security suggestions, providing more guiding bases for users to understand attack intentions and make further response decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic structural diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0031] First, the technical terms are explained as follows:
[0032] LoRA: Low-Rank Adaptation;
[0033] QLoRA: Quantized Low-Rank Adaptation;
[0034] XSS: Cross Site Scripting;
[0035] Sqli SQL: Sql injection;
[0036] OWASP: Open Web Application Security Project.
[0037] A method for constructing a large model for Web attack detection provided by the present invention includes the following steps:
[0038] Step 1: Data preparation and processing phase. First is data acquisition. The sources of Web request data are divided into two parts: one is the open-source dataset, and the other is the locally built WebGoat penetration testing platform of OWASP (Open Web Application Security Project). Attack scripts are generated using automated penetration tools such as Metasploit, BeFF, and sqlmap, and then packet capture and extraction are performed using the Burpsuites tool (a Web application security testing tool). Combining the above two data sources, ensure that all sample data, including normal requests, XSS attacks, SQLi, malicious URLs, and remote command executions, is not less than 10,000, and the number of samples for each type of attack is not less than 1,000.
[0039] Next is the extraction of Web request data. The specific information paradigm to be extracted includes three parts:
[0040] One is the request line, which specifically includes the request method, target URL, and HTTP version. The request method is limited to five types: GET, POST, HEAD, PUT, and DELETE.
[0041] The second is the request header, which specifically includes Accept, Accept-Encoding, Accept-Charset, Accept-Language, Accept-Ranges, Authorization, Connection, Cookie, Content-Length, Content-Type, Cache-control, Host, Pragma, Referer, and User-Agent.
[0042] Thirdly, there is the request body, which is the form data actually submitted to the server in this request. The specific format of the request body is specified by the content-type parameter in the request header.
[0043] After extracting Web request data according to the above paradigm, it is necessary to further clean the extracted data, including handling missing values in the request header and URL decoding. Specifically, all missing values in the request header are filled with "NONE", and the URL part of the request line and the form data in the request body are URL decoded to obtain the original string containing special characters; the special characters may include: slash / : path separator; equal sign =: separating the name and value of the parameter; ampersand &: connecting multiple query parameters; question mark?: marking the start of the query string. Instance 1 is a sample of the Web request obtained through the above processing, and the domain name and resources have been desensitized.
[0044] Instance 1: A Web request sample obtained through the data generation and processing module.
[0045] [{
[0046] "Request line":
[0047] "POST http: / / example.com:8080 / comments HTTP / 1.1",
[0048] "Request header":
[0049] "Host:example.com
[0050] User-Agent:Mozilla / 5.0(compatible;Konqueror / 3.5;Linux)
[0051] Content-Type:application / x-www-form-urlencoded
[0052] Content-Length:60
[0053] Pragma:no-cache
[0054] Cache-control:no-cache
[0055] Accept:text / xml,application / xml,application / xhtml+xml
[0056] Accept-Encoding:x-gzip,x-deflate,gzip,deflate
[0057] Accept-Charset: utf-8, utf-8; q=0.5, *; q=0.5
[0058] Accept-Language: en
[0059] Authorization: NONE
[0060] Referer: NONE
[0061] Cookie: JSESSIONID=B490DFC47CA0FE199F1218FDA4F96367
[0062] Content-Type: application / x-www-form-urlencoded
[0063] Connection: close"
[0064] "Request body":
[0065] "username=Tom&comment=" <script>alert('Stored XSS')< / script> "
[0066] }]
[0067] The Web request sample in the above example contains three parts: the request line, the request headers, and the request body, which respectively represent:
[0068] In the request line, the HTTP method is defined as POST, indicating data submission. The target URL of the request points to the comments resource under the example.com domain name, and the port number is 8080.
[0069] In the request headers:
[0070] Host: The target host of the request is example.com.
[0071] User-Agent: The user agent string of the client, indicating that the Konqueror browser is used and running on the Linux operating system.
[0072] Content-Type: Indicates that the content type of the request body is application / x-www-form-urlencoded, usually used for submitting form data.
[0073] Content-Length: The length of the request body is 60 bytes.
[0074] Pragma and Cache-control: Both of these fields contain "no-cache", indicating that the request does not want to be cached.
[0075] Accept: Represents the content types that the client can accept, listing text / xml, application / xml, and application / xhtml+xml, etc.
[0076] Accept-Encoding: The encoding methods that the client can accept, listing gzip and deflate, etc.
[0077] Accept-Charset: Represents the character sets that the client can accept, including utf-8.
[0078] Accept-Language: Represents that the language type accepted by the client is English.
[0079] Authorization and Referer: These two fields are NONE, which may indicate that the request has no authentication information and source.
[0080] Cookie: Contains a Cookie named JSESSIONID with the value B490DFC47CA0FE199F1218FDA4F96367.
[0081] Connection: Indicates that the connection will be closed after the request is completed.
[0082] In the request body, the '&' character connects two parts of data: username=Tom and comment=. <script>alert('Stored XSS')< / script> The former defines that the username submitted by the user is Tom. comment= <script>alert('Stored XSS')< / script> Represents the comment submitted by the user, which contains a potential Stored Cross-Site Scripting (XSS) code segment that attempts to execute the alert('Stored XSS') script in the browser.
[0083] Step 2: Based on the extracted Web request samples, generate expected detection results and explanatory analyses. For samples obtained from open-source datasets, the expected detection results are given by the labels of the datasets; for samples recorded by a locally built penetration testing platform, the expected detection results are labeled by the personnel implementing the attack penetration. The labeling criteria are: for normal requests, label as "no Web attack", and for abnormal attack requests, label as the corresponding attack type, such as "XSS attack detected".
[0084] Explanatory analysis is the basis for sample discrimination. It is generated by manual analysis of security experts. On the one hand, it provides the model with the paradigms and behavioral intents of attacks for learning. On the other hand, it provides decision-making basis and security suggestions for security personnel. Explanatory analysis needs to include Web request analysis, security analysis, and security suggestions.
[0085] Step 3: According to the extracted Web request samples, expected detection results, and explanatory analysis results, construct an instruction fine-tuning dataset. In the instruction fine-tuning dataset, each fine-tuning sample contains three parts: instruction, input, and output. The instruction part is the system prompt word for the task, which is fixed in the template. The input part is the extracted Web request sample, and the output part is the expected label and explanatory analysis information of the Web request sample, which is also the overall expected output of the model. Finally, save the fine-tuning dataset in json format. Example 2 is an example of a fine-tuning data sample:
[0086] Example 2: Composition of fine-tuning data sample [{
[0087] "instruction":"
[0088] Please determine whether the following Web request message has a web attack risk: ",
[0089] "input":
[0090] "POST http: / / example.com:8080 / comments HTTP / 1.1
[0091] Host:example.com
[0092] User-Agent:Mozilla / 5.0(compatible;Konqueror / 3.5;Linux)
[0093] Content-Type:application / x-www-form-urlencoded
[0094] Content-Length:62
[0095] Pragma:no-cache
[0096] Cache-control:no-cache
[0097] Accept:text / xml,application / xml,application / xhtml+xml
[0098] Accept-Encoding: x-gzip, x-deflate, gzip, deflate
[0099] Accept-Charset: utf-8, utf-8;q=0.5, *;q=0.5
[0100] Accept-Language: en
[0101] Authorization: NONE
[0102] Referer: NONE
[0103] Cookie: JSESSIONID=B490DFC47CA0FE199F1218FDA4F96367
[0104] Content-Type: application / x-www-form-urlencoded
[0105] Connection: close
[0106] username=JohnDoe&comment= <script>alert('Stored XSS')< / script> ",
[0107] "output":
[0108] "XSS attack risk detected.\n
[0109] Web request analysis: \n HTTP method and path... Request header analysis...
[0110] Security analysis: \n 1. Input validation issue: The comment field contains <script>标签,可能出发存储型跨站脚本攻击,危害包括……\n2.Cookie使用问题:…"
[0111] 安全建议:\n 1.输入数据验证和清理,使用输入验证对username和comment的内容严格限制……\n2.存储输出编码……"
[0112] }]
[0113] 上述微调数据样本示例中各部分具体意义如下:
[0114] Instruction是任务的系统提示词,规定了大语言模型的任务是判断输入中的Web请求报文中是否存在Web攻击风险。
[0115] Input是用户提供的待检测Web请求报文,按照步骤1的定义,它包含请求行、请求头和请求体三部分,其具体表示已在步骤1中说明。
[0116] Output是预期得到的大语言模型输出,它包含步骤2中定义的预期检测结果和解释性分析。
[0117] 步骤4:量化微调训练,利用指令微调数据集对中文大模型的全连接层进行微调,得到Web攻击检测大模型。可以选择的开源中文大模型包括,通义千问、百川大模型等,根据选择的大模型的输入输出模板,在训练时对样本进行适配。本方法采用的微调策略为量化低秩适配技术(QLoRA),以降低微调模型所需的训练成本,在加载基础模型后,对加载的基础模型参数进行4-bit量化,遍历查找模型所有的线性层并插入可训练的低秩适配(LoRA)模块,训练的过程中冻结基础模型权重,仅仅更新插入的LoRA模块参数。微调训练结束后,将LoRA模块权重与其对应的基础模型的线性层权重矩阵相加合并。
[0118] 参见图1,本发明提供的一种Web攻击检测大模型的构建系统,包括:
[0119] 数据准备和处理模块,用于获取Web请求数据,从Web请求数据中提取请求行、请求头、请求体;对提取出的数据进行请求头缺失值处理和URL解码,得到Web请求样本;
[0120] 预期检测结果和解释性分析模块,用于根据Web请求样本,生成预期检测结果和解释性分析结果;
[0121] 指令微调数据集构造模块,用于根据Web请求样本、预期检测结果和解释性分析结果,构造指令微调数据集;
[0122] 大模型量化微调模块,用于利用微调策略和指令微调数据集,对中文大模型进行微调训练,得到Web攻击检测大模型。
[0123] 以上实施例仅用于说明本发明的设计思想和特点,其目的在于使本领域内的技术人员能够了解本发明的内容并据以实施,本发明的保护范围不限于上述实施例。所以,凡依据本发明所揭示的原理、设计思路所作的等同变化或修饰,均在本发明的保护范围之内。< / script>
Claims
1. A method for constructing a large model for Web attack detection, characterized in that, It includes the following steps: (1) Obtain Web request data, and extract the request line, request headers, and request body from the Web request data; perform missing value processing on the extracted headers and URL decoding to obtain a Web request sample; (2) Generate an expected detection result and an interpretive analysis result based on the Web request sample obtained in step (1); (3) Construct an instruction fine-tuning dataset based on the Web request sample obtained in step (1), the expected detection result obtained in step (2), and the interpretive analysis result; (4) Use a fine-tuning strategy and the instruction fine-tuning dataset obtained in step (3) to fine-tune and train a Chinese large model to obtain a Web attack detection large model.
2. The method for constructing a large model for Web attack detection according to claim 1, wherein In step (1), the sources of Web request data include open-source datasets and locally built OWASP WebGoat penetration testing platforms; use Metasploit, BeFF, and sqlmap automated penetration tools to generate attack scripts, and then use the Burpsuites tool to capture packets for extraction; among the obtained Web request data, the total number of sample data including normal requests, XSS attacks, SQLi, malicious URLs, and remote command executions is not less than 10,000, and the number of samples for each attack is not less than 1,000.
3. The method for constructing a large Web attack detection model according to claim 1, wherein In step (1), the missing value processing of the extracted headers and URL decoding includes: filling the missing values in the headers with "NONE", and performing URL decoding on the URL part of the request line and the form data in the request body.
4. The method for constructing a large model for Web attack detection according to claim 2, wherein In step (2), for the samples obtained from the open-source dataset, the expected detection result is given by the label of the dataset; for the samples recorded in the locally built penetration testing platform, the expected detection result is manually labeled; if it is a normal request, it is labeled as "no Web attack", and if it is an abnormal attack request, it is labeled with the corresponding attack type; the interpretive analysis result is the basis for sample discrimination, generated by manual analysis, and includes Web request analysis, security analysis, and security suggestions.
5. The method for constructing a large model for Web attack detection according to claim 1, characterized in that In the instruction fine-tuning dataset of step (3), the fine-tuning sample includes a fine-tuning instruction part, an input part, and an output part; the fine-tuning instruction part is the system prompt word of the task, the input part is the Web request sample, and the output part is the expected label and interpretive analysis information of the Web request sample; save the instruction fine-tuning dataset in json format.
6. The method for constructing a large model for Web attack detection according to claim 1, wherein In step (4), the Chinese large model includes Tongyi Qianwen and Baichuan large model, and adapts the samples during training according to the input-output templates of the selected large model.
7. The method for constructing a large model for Web attack detection according to claim 1, wherein In step (4), the fine-tuning strategy is the quantization low-rank adaptation technique, which is used to reduce the training cost required for the fine-tuning model; after loading the base model, perform 4-bit quantization on the parameters of the loaded base model, traverse and find all the linear layers of the model and insert trainable low-rank adaptation modules, freeze the weights of the base model during the training process, and only update the parameters of the inserted low-rank adaptation modules; after the fine-tuning training is completed, add and merge the weights of the low-rank adaptation modules with the corresponding linear layer weight matrices of the base model.
8. A construction system for a large model of Web attack detection, characterized in that, It includes: A data preparation and processing module for obtaining Web request data, extracting a request line, request headers, and a request body from the Web request data; Performing missing value processing on the extracted headers and URL decoding to obtain a Web request sample; An expected detection result and interpretive analysis module for generating an expected detection result and an interpretive analysis result based on the Web request sample; An instruction fine-tuning dataset construction module for constructing an instruction fine-tuning dataset based on the Web request sample, the expected detection result, and the interpretive analysis result; A large model quantization fine-tuning module for fine-tuning and training a Chinese large model using a fine-tuning strategy and the instruction fine-tuning dataset to obtain a Web attack detection large model.