Network security vulnerability analysis method and system based on large-model low-rank adaptation fine tuning training
By constructing a large network security vulnerability analysis model with low-rank adaptive training, the problem of the inability to identify new threats in existing technologies has been solved, efficient and accurate vulnerability analysis and attack pattern recognition have been achieved, and network security protection capabilities have been improved.
Patent Information
- Application Number
- CN202510984881.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-10
AI Technical Summary
Existing network security vulnerability analysis methods rely on static knowledge bases and preset vulnerability libraries, which cannot effectively identify new threats and unknown attacks. They lack the contextual understanding and reasoning capabilities of large models, resulting in insufficient security protection capabilities.
By collecting network security corpora, performing preprocessing and supervised fine-tuning, a large network security vulnerability analysis model with low-rank adaptive training is constructed. Real-time monitoring is performed using large language models and low-rank adaptation technology to identify potential security vulnerabilities and attack patterns.
It achieves efficient and accurate vulnerability analysis, can identify new attack patterns, improves network security protection capabilities, reduces computing resource requirements, and improves the robustness and generalization capabilities of the model.
Smart Images

Figure CN120768626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a network security vulnerability analysis method and system based on large-model low-rank adaptive fine-tuning training. Background Art
[0002] At present, network security vulnerability analysis technology mainly relies on traditional rules and feature engineering methods. These methods usually rely on manually designed vulnerability features and rules. These rules are often limited to the detection of known vulnerabilities and cannot effectively deal with zero-day vulnerabilities or new unknown attacks.
[0003] With the rapid development of large language models (LLMs), vulnerability analysis methods based on these models have shown great potential. Large language models process large amounts of structured and unstructured data during pre-training, enabling them to extract valuable insights from massive amounts of network logs, vulnerability reports, and threat intelligence. Adaptively training large language models with domain-specific knowledge not only significantly improves the detection accuracy of known vulnerabilities but also effectively identifies potential new attack patterns.
[0004] Existing network vulnerability analysis methods mostly rely on expert rules and manual features, and vulnerability determination is based on predefined logical chains. For example, a network security knowledge graph construction method and apparatus for dynamic threat analysis, published in CN110113314B, leverages the interplay between network attack behaviors and system vulnerabilities to construct a knowledge graph model for dynamic network threat analysis. However, this method relies on a static knowledge base, lacks dynamic adaptability, and has limited capabilities for mining associations.
[0005] Another example is a network security vulnerability detection method and system based on big data analysis in CN115412354A. This network security vulnerability detection method is based on a two-level classifier (vulnerability risk level + frequency screening), which improves detection accuracy by matching vulnerability attack samples with ordinary sample data sets; however, this method relies on a preset vulnerability library and only considers frequency and risk level, lacking the fusion of multi-dimensional features such as code semantics and attack context.
[0006] In summary, traditional rule-based or machine learning detection methods are typically only able to identify known vulnerabilities and attacks, but are unable to address new threats. As network environments become increasingly complex and attack methods continue to evolve, existing methods fail to fully leverage the contextual understanding and reasoning capabilities of large models to address dynamically changing network security threats. This results in many potential vulnerabilities going undetected, significantly limiting security protection capabilities. Therefore, improving the intelligence and data processing capabilities of vulnerability analysis systems is key to enhancing network security protection capabilities. Summary of the Invention
[0007] The present application aims to overcome the defects of the prior art and provides a network security vulnerability analysis method and system based on large model low-rank adaptation fine-tuning training.
[0008] The object of the present application can be achieved by the following technical solutions:
[0009] According to one aspect of the present application, a network security vulnerability analysis method based on large model low-rank adaptation fine-tuning training is provided, the method comprising the following steps:
[0010] S1, selecting a large language model and pre-training it;
[0011] S2, collecting a network security corpus and pre-processing it to build a supervised fine-tuning data set;
[0012] S3, based on the supervised fine-tuning data set, fine-tuning the pre-trained large language model through low-rank adaptation training to obtain a network security vulnerability analysis large model;
[0013] S4, based on the network security vulnerability analysis large model, real-time monitoring and analysis of the network system to identify potential security vulnerabilities and attack patterns.
[0014] As a preferred technical solution, the method of collecting the network security corpus in S2 includes network collection and document extraction;
[0015] The network collection process uses a web crawler to collect network security-related corpus data, including vulnerability libraries, attack records, network traffic, and log data;
[0016] The document extraction process uses Python-related libraries and OCR (Optical Character Recognition) tools to extract document content and further extract key information such as vulnerability descriptions, attack methods, and impact ranges from the document content to build structured data.
[0017] As a preferred technical solution, the pre-processing process in S2 includes: using oversampling or undersampling techniques to balance the data and using data augmentation methods to expand the data set;
[0018] After the pre-processing process is completed, the pre-processed data set is quality evaluated, the quality evaluation content including: checking whether there are missing values or invalid data in the data, if there are, performing corresponding filling or rejection; ensuring the consistency of data format and content between different data sources to avoid analysis bias caused by inconsistent formats; each piece of data is stored in a json file according to its corresponding category.
[0019] As a preferred technical solution, the supervised fine-tuning data set in S2 is constructed based on the opinions and professional knowledge of experts in the field of network security, combined with the preprocessed data set, by designing an effective question and answer pair form; the specific construction process is as follows: first, classify and arrange vulnerability related text data, including vulnerability description, analysis, solution and detection method, then use the pre-trained large language model to generate question and answer pairs; finally, the question and answer pairs are reviewed by experts and processed in batches by API (programming interface); the question and answer pairs are constructed into a supervised fine-tuning data set in Alpaca format, i.e. prompt-question-answer format, and are manually sampled and reviewed and iteratively optimized.
[0020] As a preferred technical solution, the large language model in S3 adopts the Qwen2.5-7B-Instruct model, and the training fine-tuning process is based on the LLaMA Factory framework.
[0021] As a preferred technical solution, the specific process of low-rank adaptation training of the large language model in S3 includes a training process and a quantization process.
[0022] The training process is to use LoRA technology to perform low-rank adaptation training on the large language model, and adjust the key low-rank part of the model; the quantization process is to use 4-bit quantization on the large language model to reduce the storage requirements of model weights and activation values, and quantize the model weights to 4-bit integers.
[0023] As a preferred technical solution, in the training process, LoRA technology is used to update the weights of the pre-trained large language model locally, and only the key low-rank part of the model is adjusted; in the low-rank adaptation training process, llora_rank is set to 8, i.e. the low-rank matrix dimension of the model weights is 8 in the low-rank adaptation process; lora_target is set to all, i.e. the LoRA method is applied to all model parameters.
[0024] As a preferred technical solution, the specific formula for low-rank adaptation training of the large language model using LoRA technology is as follows:
[0025] ΔW=BA,B∈R d×r ,A∈R r×k ,r<<min(d,k)
[0026] where ΔW is the update amount of the large language model weight matrix; B and A are trainable low-rank matrices; d is the input dimension of the weight matrix, k is the output dimension of the weight matrix, and r is the rank of the low-rank matrix; the parameter amount is d x r + r x k.
[0027] As a preferred technical solution, the specific formula for 4-bit quantization of the large language model is as follows:
[0028]
[0029] Where s is the scaling factor; W block is the original weight matrix block; z is the zero offset; W int4 To quantify the results.
[0030] According to another aspect of the present invention, a network security vulnerability analysis system based on large model low-rank adaptive fine-tuning training is provided, the system including a large language model pre-training module, a network security corpus processing module, a low-rank adaptive fine-tuning module and a network security vulnerability monitoring module;
[0031] The large language model pre-training module is used to select a large language model and pre-train it;
[0032] The network security corpus processing module is used to collect and preprocess the network security corpus to construct a supervised fine-tuning dataset;
[0033] The low-rank adaptive fine-tuning module performs low-rank adaptive training and fine-tuning on the pre-trained large language model based on the supervised fine-tuning dataset to obtain a large model for network security vulnerability analysis.
[0034] The network security vulnerability monitoring module monitors and analyzes the network system in real time based on the network security vulnerability analysis model, thereby identifying and outputting potential security vulnerabilities and attack patterns.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. In the present invention, a supervised fine-tuning dataset is constructed from a security corpus, and the supervised fine-tuning dataset is used to perform low-rank adaptive training and fine-tuning on a large language model to obtain a large model for network security vulnerability analysis. Based on the large model for network security vulnerability analysis, the network system is monitored and analyzed in real time, thereby identifying and outputting potential security vulnerabilities and attack patterns; through the large language model and low-rank adaptive training, efficient and accurate vulnerability analysis is achieved without increasing excessive computing resources; at the same time, the context understanding and reasoning capabilities of the large model are fully utilized, so that the present invention can effectively identify new and unseen attack patterns.
[0037] 2. In this invention, network security corpus is collected through dual channels of web crawlers and document extraction. This not only covers real-time network data to ensure freshness, but also deeply explores vulnerability knowledge in historical or professional documents through document extraction. This makes the data set cover dynamic attack records and static professional analysis, providing a more comprehensive data foundation for the subsequent training and fine-tuning of large language models.
[0038] 3. In this invention, to address sample imbalances (e.g., in network security vulnerability data, common attack patterns may account for the majority of the data, while unknown attack patterns are less common), oversampling or undersampling techniques are used to balance the data, ensuring the sensitivity of the model to different types of vulnerabilities during training. For small amounts of sample data, data augmentation methods (such as text generation and simulated data generation) are used to expand the dataset, improving the robustness and generalization ability of the model. Furthermore, the quality of the preprocessed dataset is assessed to ensure its suitability for subsequent vulnerability analysis tasks and to ensure consistency in data format and content across different data sources, thereby avoiding analysis bias caused by inconsistent formats.
[0039] 4. In this invention, the large language model adopts the Qwen2.5-7B-Instruct model, which is a high-quality base adapted to Chinese scenarios and has instruction-following capabilities. The training and fine-tuning process is based on the LLaMA Factory framework, providing an efficient training and fine-tuning environment. Through the selection of models and frameworks, the technical threshold is lowered and the training efficiency is improved.
[0040] 5. In the present invention, the specific process of low-rank adaptive training includes a training process and a quantization process. The training process uses LoRA technology to perform low-rank adaptive training on the large language model, adjusts the key low-rank parts of the model, greatly reduces the video memory occupancy, and optimizes the memory and computing overhead; in the quantization process, 4-bit quantization is used for the large language model to reduce the storage requirements of the model weights and activation values, and quantize the model weights into 4-bit integers, further improving efficiency, making the model more efficient to train and easier to deploy in network security analysis scenarios.
[0041] 6. During the low-rank adaptation training process of the present invention, 1lora_rank is set to 8, that is, during the low-rank adaptation process, the low-rank matrix dimension of the model weight is 8; this setting optimizes memory and computing overhead while ensuring the model effect; lora_target is set to all, that is, the LoRA method is applied to all model parameters, which enables the entire model to quickly adapt to new tasks through low-rank adaptation, thereby improving the accuracy and efficiency of network security vulnerability analysis tasks.
[0042] 7. The present invention proposes a network security vulnerability analysis system based on large-scale low-rank adaptive fine-tuning training, comprising a large language model pre-training module, a network security corpus processing module, a low-rank adaptive fine-tuning module, and a network security vulnerability monitoring module. Pre-training, corpus processing, fine-tuning, and monitoring each perform their respective functions, forming a complete "data-model-application" system closed loop. Through the intelligent system designed by the present invention, the model can analyze potential attack behaviors in historical data and reason based on context, promptly capturing latent or progressive attacks, thereby significantly improving network security protection capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 Schematic diagram of the steps of a network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training in the present invention;
[0044] Figure 2 Schematic diagram of the process of establishing a large network security vulnerability analysis model in the embodiment;
[0045] Figure 3 This is a schematic diagram of the JSON file format for storing preprocessed data in an embodiment;
[0046] Figure 4 This is a schematic diagram of the JSON file format for storing supervised fine-tuning data in an embodiment. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0048] Example
[0049] In this embodiment, a network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training is adopted, and the method steps are as follows: Figure 1 As shown, specifically including:
[0050] S1. Select a large language model and pre-train it;
[0051] S2. Collect and preprocess the cybersecurity corpus to construct a supervised fine-tuning dataset;
[0052] S3. Based on the supervised fine-tuning dataset, the pre-trained large language model is fine-tuned using low-rank adaptive training to obtain a large model for network security vulnerability analysis.
[0053] S4. Real-time monitoring and analysis of network systems based on a large network security vulnerability analysis model, thereby identifying and outputting potential security vulnerabilities and attack patterns.
[0054] In this embodiment, the process of establishing a large network security vulnerability analysis model is as follows: Figure 2 The specific implementation process is as follows:
[0055] 1) Collection and preprocessing of network security corpus: By extensively collecting network security-related corpus data, including vulnerability libraries, attack records, network traffic, and log data, and performing effective screening and preprocessing, a high-quality dataset suitable for vulnerability analysis is constructed.
[0056] The collection and preprocessing of a cybersecurity corpus is fundamental to building a vulnerability analysis model. To ensure the comprehensiveness and high quality of training data, this solution collects a wide range of cybersecurity-related data through various methods, including vulnerability databases, attack logs, network traffic, log files, and other raw data.
[0057] The data sources of this solution include two channels: network collection and document extraction. In order to efficiently collect vulnerability information and security incident records disclosed on the Internet, crawler technology is used to crawl data from multiple network security platforms. At the same time, many technical reports and vulnerability analysis documents in the field of network security often exist in PDF or Docx format. This solution uses python related libraries and OCR tools to extract document content, and further extract key information such as vulnerability description, attack method, impact range, etc. to construct structured data. The collected network security corpus usually contains a lot of redundant information and noise, so this solution performs comprehensive preprocessing on the collected network security corpus to ensure that the data can adapt to subsequent vulnerability analysis tasks. For data from different sources, a unified JSON standard format is used for integration, and different feature extraction schemes are designed according to the data type and purpose.
[0058] 2) Construct a supervised fine-tuning dataset: Combining the opinions and domain knowledge of cybersecurity experts, a supervised fine-tuning dataset is constructed based on the preprocessed dataset to ensure that the dataset covers various types of vulnerability information and their characteristics, providing an effective training foundation.
[0059] The construction of the supervised fine-tuning dataset is based on the opinions and expertise of experts in the field of cybersecurity, combined with the pre-processed dataset, and constructed by designing an effective question-answer pair format. This solution generates high-quality question-answer pairs by classifying and organizing vulnerability-related text data (such as vulnerability descriptions, analysis, solutions, and detection methods), and using a pre-trained large language model. Combined with expert review and API batch processing, the accuracy and professionalism of the generated data are ensured. The question-answer pairs are constructed into a supervised fine-tuning dataset in the Alpaca format (prompt-question-answer), and undergo manual sampling review and iterative optimization to ensure balanced data quality and comprehensive coverage, thereby improving the accuracy and generalization ability of the model in vulnerability analysis tasks. This process ensures that the dataset covers a variety of vulnerability information and features, providing targeted and high-quality training data for subsequent large language model training.
[0060] 3) Low-rank adaptive fine-tuning training: Low-rank adaptation techniques are used to fine-tune pre-trained large language models. Low-rank adaptation methods can effectively adapt to specific tasks (such as vulnerability analysis) without increasing excessive computing resources, thereby improving the accuracy and efficiency of the model.
[0061] In this example, low-rank adaptation technology is used to fine-tune the open-source large language base model. This allows the model to quickly adapt to specific tasks (such as vulnerability analysis) without increasing excessive computing resources, thereby improving the model's accuracy and efficiency.
[0062] In this embodiment, the open source large language model Qwen2.5-7B-Instruct is first selected. The model has been pre-trained on a large-scale data set and has strong language understanding and generation capabilities. In order to enable it to handle vulnerability analysis tasks in the field of network security more accurately, this solution is fine-tuned based on the LLaMA Factory framework. LLaMAFactory is an efficient deep learning framework that can support the efficient implementation of low-rank adaptation methods and provide the necessary tools to manage resource scheduling and parameter optimization in large-scale training processes. During the training process, by real-time monitoring of loss functions and accuracy indicators, it is ensured that the model gradually converges after each training cycle and can effectively adapt to the needs of vulnerability analysis tasks. At the same time, verification is performed regularly during the training process, and an independent verification set is used to evaluate the generalization ability and accuracy of the model to ensure that the fine-tuning process does not lead to overfitting.
[0063] 4) Training and Evaluation: By training and evaluating the fine-tuned large language model, we build an intelligent system specifically for network security vulnerability analysis. This system can accurately identify potential security vulnerabilities when analyzing massive amounts of network data and identify new, unseen attack patterns through contextual reasoning, significantly improving the vulnerability analysis capabilities of network systems.
[0064] Based on a fine-tuned large language model, this solution builds an intelligent system specifically for network security vulnerability analysis. This system can accurately identify potential security vulnerabilities and capture new and unseen attack patterns through contextual reasoning, significantly improving the vulnerability analysis capabilities of network systems.
[0065] To further validate the model's ability to identify novel attack patterns, this solution was tested using unseen attack data. This data included new vulnerabilities and attack methods that have emerged in recent years, requiring the model to identify these novel security threats through contextual reasoning. Evaluation results demonstrated that the fine-tuned model performed well in vulnerability identification and analysis tasks, effectively supporting cybersecurity personnel in quickly and accurately identifying and responding to security risks.
[0066] Step 1: First, collect high-quality cybersecurity data from the internet and documents. Online platforms include vulnerability announcement websites (such as the CVE database), security technology blogs, social media, and forums. Data collection is done using a crawler. The crawler design primarily considers the following: Designing appropriate URL (Uniform Resource Locator) filtering rules ensures crawled data is relevant to cybersecurity vulnerabilities and avoids collecting irrelevant content. During the crawling process, duplicate, invalid, or low-quality web content is filtered out in real time to ensure the validity of the collected data. To extract structured information from documents, PDF (Portable Document Format) documents are parsed using a PDF parsing library. For encrypted or specially formatted PDF files, OCR (Optical Character Recognition) is used to extract text. For Docx documents, the Python python-docx library is used to parse Docx files, extracting text content and extracting and analyzing information such as tables and images. As for the imbalance problem of samples (for example, in network security vulnerability data, common attack patterns may occupy most of the data, while the data of unknown attack patterns is relatively small), the data is balanced through oversampling or undersampling technology to ensure the sensitivity of the model to different types of vulnerabilities during training; for a small amount of sample data, data enhancement methods (such as text generation, simulation data generation, etc.) are used to expand the data set to improve the robustness and generalization ability of the model. After the data preprocessing is completed, the quality of the preprocessed data set is evaluated to ensure its suitability for subsequent vulnerability analysis tasks. The evaluation content includes checking whether there are missing values or invalid data in the data, and filling or eliminating them accordingly; ensuring the consistency of data format and content between different data sources to avoid analysis bias caused by inconsistent formats. Each piece of data is stored in the following format according to different categories. Figure 3 json file shown.
[0067] Step 2: Categorize and organize the processed raw text data according to different requirement categories, such as "Vulnerability Explanation," "Attack Pattern," "Defense Method," and "Detection Method." Each data category contains different forms of cybersecurity information. "Vulnerability Explanation" primarily includes a brief description of the vulnerability, its scope of impact, and its type; "Attack Pattern" includes a technical analysis of the vulnerability, its attack methods, and its mechanism; "Defense Method" includes specific vulnerability remediation plans, patches, and countermeasures; and "Detection Method" describes the detection methods, tools, and process for the vulnerability.
[0068] After preliminary classification, the text information from these categories is fed into a pre-trained large language model. Leveraging the DeepSeek-V3 model's powerful contextual understanding and generation capabilities, appropriate prompts are designed to construct high-quality question-answer pairs. To ensure the high quality and accuracy of the generated data, this solution uses an API (Application Programming Interface) to efficiently batch process the raw data. This solution also incorporates feedback and corrections from domain experts to ensure that each question-answer pair meets the actual needs of the cybersecurity field. During this process, experts review the questions and answers generated by the model, particularly in areas with complex technical details such as vulnerability descriptions or solutions, to ensure the authenticity and professionalism of the data.
[0069] In step 3, the generated question-answer pairs are further processed and constructed into a supervised fine-tuning dataset in the Alpaca format (i.e., prompt-question-answer format). The Alpaca format can ensure that the model can clearly understand the correspondence between questions and answers during training, so that the model can better learn task-specific knowledge. In order to improve the quality and effectiveness of the dataset, this solution strictly controls the quantity and quality of each type of questions and answers to ensure that the question-answer pairs in each category are balanced and comprehensive. At the same time, a certain proportion of question-answer pairs are randomly sampled for manual inspection and adjustment to ensure that the generated dataset is free of noise and errors. In addition, this solution also adopts an iterative review mechanism to continuously adjust and optimize the question-answer pairs in the dataset based on the results and feedback of model training, so as to improve the accuracy and generalization ability of the model in actual vulnerability analysis tasks. The constructed supervised fine-tuning data is stored in the following format: Figure 4 json file shown.
[0070] In the process of constructing the supervised fine-tuning dataset, manual review and sampling verification are also performed.
[0071] In step 4, this solution uses QLoRA technology to perform low-rank adaptation training on the Qwen2.5-7B-Instruct model. LoRA is a fine-tuning method based on low-rank matrix decomposition. By locally updating the weights of the pre-trained model, only the key low-rank parts of the model are adjusted. The lora_rank is set to 8, which means that during the low-rank adaptation process, the low-rank matrix dimension of the model weight is 8. This setting optimizes memory and computational overhead while ensuring the model effect; lora_target is all, that is, the LoRA method is applied to all model parameters, which enables the entire model to quickly adapt to new tasks through low-rank adaptation, thereby improving the accuracy and efficiency of network security vulnerability analysis tasks. The mathematical principles of LoRA technology are as follows:
[0072] ΔW=BA,B∈R d×r ,A∈R r×k ,r< <min(d,k)
[0073] Among them, d represents the input dimension of the weight matrix, k represents the output dimension of the weight matrix (usually equal to d), and r represents the rank of the low-rank matrix; B and A represent trainable low-rank matrices, and the number of parameters is only d×r+r×k. When LoRA (r=8) is applied, the original matrix parameter volume of 16,777,216 (4096×4096) with a dimension of 4096 can be compressed to 4096×8+8×4096=65,536, which is compressed by 256 times, greatly reducing the memory usage. To further improve efficiency, this solution also applies 4-bit quantization during the model fine-tuning process. By reducing the storage requirements of model weights and activation values, the model weights are quantized into 4-bit integers (INT4), and the bitsandbytes technology is applied, that is, the improved quantization scheme of LLM.int8() is adopted to support asymmetric quantization. The specific formula of the quantization technology is as follows:
[0074]
[0075] Where block_size represents the quantization block size, s represents the scale factor, z represents the zero point offset, and W_{int4} represents the 4-bit integer quantization result. For the original FP16 model, the size is ≈14GB, and after 4-bit quantization, the size is ≈5.6GB.
[0076] In addition, the training process was experimented and adjusted to set the optimal parameters such as learning rate, warm-up ratio, and gradient accumulation to ensure stable convergence of the model in multiple rounds of training and avoid overfitting. Some training parameter configurations are shown in Table 1.
[0077] Table 1 Some training parameters
[0078]
[0079] Where model_name_or_path specifies the name or path of a pre-trained large language model. In this solution, we use "Qwen2.5-7B-Instruct," a 7-billion-parameter instruction-tuned model. The Qwen series of models excels in Chinese contexts and instruction understanding tasks, making them suitable for text analysis in cybersecurity (such as vulnerability descriptions and log parsing). The 7-b parameter size strikes a balance between performance and computational resource consumption, facilitating subsequent low-rank adaptation and fine-tuning.
[0080] quantization_bit is the number of bits used to quantize the model weights. Setting it to 4 means compressing the 32-bit floating-point weights into 4-bit integers. This reduces storage costs, memory access and computational complexity, and balances accuracy and efficiency. Quantization errors can be further compensated for by combining it with LoRA fine-tuning.
[0081] quantization_method:bitsandbytes refers to the use of the bitsandbytes library to implement quantization, which supports low-bit quantization and mixed precision training. finetuning_type:lora means fine-tuning with low-rank adaptation technology, training only a small number of new parameters, and freezing most of the weights of the original model; thereby reducing training costs, preventing overfitting, and adapting to dynamic changes in the field of network security (such as new vulnerability types), facilitating subsequent incremental training. lora_rank is the rank of the low-rank matrix in LoRA technology, with a dimension of r, which determines the scale of new trainable parameters; lora_target:all means applying LoRA to all trainable parameters of the model; cutoff_len indicates the maximum length of the input sequence, and the excess will be truncated; per_device_train_batch_size is the number of samples trained on each GPU device each time; gradient_accumulation_steps:16 means updating the model parameters once after accumulating 16 small batch gradients; learning_rate is the learning rate, and a smaller learning rate is suitable for fine-tuning scenarios to ensure smooth integration of domain knowledge; num_train_epochs is the number of training rounds for the entire dataset times; lr_scheduler_type: cosine indicates that the learning rate scheduling strategy adopts cosine decay, that is, the learning rate decreases with the training rounds in a cosine curve; warmup_ratio: 0.1 indicates that the learning rate is gradually increased in a linear manner at the beginning of training, and the warm-up stage accounts for 10% of the total rounds; bf16: TRUE indicates that BF16 (16-bit brain floating point) mixed precision training is enabled; this solution uses LoRA, 4-bit quantization, and BF16 mixed precision to reduce training and inference costs and adapt to network security real-time monitoring scenarios; reasonably set the learning rate, batch strategy, and training rounds to ensure that the model is lightweight while accurately capturing vulnerability characteristics and attack patterns; based on the long text and multi-source data characteristics in the field of network security, optimize the sequence length and parameter update strategy to improve the model generalization ability.
[0082] Ultimately, the fine-tuned model can quickly identify potential vulnerabilities in massive network data and effectively capture unseen attack patterns through contextual reasoning, improving the accuracy and efficiency of network security vulnerability analysis.
[0083] In summary, this solution achieves efficient and accurate vulnerability analysis through large language models and low-rank adaptive training without increasing excessive computing resources. At the same time, it fully utilizes the contextual understanding and reasoning capabilities of the large model, enabling the present invention to effectively identify new and unseen attack patterns.
[0084] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training, characterized by: The method comprises the following steps: S1. Select a large language model and pre-train it; S2. Collect and preprocess the cybersecurity corpus to construct a supervised fine-tuning dataset; S3. Based on the supervised fine-tuning dataset, the pre-trained large language model is fine-tuned using low-rank adaptive training to obtain a large model for network security vulnerability analysis. S4. Real-time monitoring and analysis of network systems based on a large network security vulnerability analysis model, thereby identifying and outputting potential security vulnerabilities and attack patterns.
2. A network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 1, characterized in that: The method for collecting the cybersecurity corpus in S2 includes network collection and document extraction; The network collection process uses a web crawler to collect network security related corpus data, including vulnerability databases, attack records, network traffic and log data; The document extraction process uses Python-related libraries and OCR tools to extract document content, and further extracts key information such as vulnerability description, attack method, and impact range from it to construct structured data.
3. A network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 2, characterized in that: The preprocessing process in S2 includes: using oversampling or undersampling technology to balance the data and using data enhancement methods to expand the data set; After the preprocessing process is completed, the quality assessment of the preprocessed data set is performed. The quality assessment content includes: checking whether there are missing values or invalid data in the data, and if so, filling or eliminating them accordingly; ensuring the consistency of data format and content between different data sources to avoid analysis bias caused by inconsistent format; each data item is stored in a json file according to its corresponding category.
4. A network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 3, characterized in that: The supervised fine-tuning dataset in S2 is constructed based on the opinions and expertise of experts in the field of network security, combined with the pre-processed dataset, by designing an effective question-answer pair format; The specific construction process is as follows: First, classify and organize vulnerability-related text data, including vulnerability descriptions, analysis, solutions, and detection methods. Then, use the pre-trained large language model to generate question-answer pairs. Finally, the question-answer pairs are reviewed by experts and batch processed by API. The question-answer pairs are constructed in the Alpaca format, i.e., prompt-question-answer format, to construct a supervised fine-tuning dataset, and are manually sampled, reviewed, and iteratively optimized.
5. The network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 1 is characterized in that: The large language model in S3 adopts the Qwen2.5-7B-Instruct model, and the training and fine-tuning process is based on the LLaMAFactory framework.
6. The network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 1 is characterized in that: The specific process of performing low-rank adaptive training on the large language model in S3 includes a training process and a quantization process; The training process is as follows: using LoRA technology to perform low-rank adaptation training on the large language model and adjust the key low-rank parts of the model; the quantization process is as follows: using 4-bit quantization on the large language model to reduce the storage requirements of the model weights and activation values, and quantizing the model weights into 4-bit integers.
7. A network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 6, characterized in that: During the training process, LoRA technology is used to locally update the weights of the pre-trained large language model, and only the key low-rank parts of the model are adjusted; during the low-rank adaptation training process, 1lora_rank is set to 8, that is, during the low-rank adaptation process, the low-rank matrix dimension of the model weight is 8; lora_target is set to all, that is, the LoRA method is applied to all model parameters.
8. The network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 6 is characterized in that: The specific formula for using LoRA technology to perform low-rank adaptation training on a large language model is: ΔW=BA,B∈R d×r ,A∈R r×k ,r<<min(d,k) Among them, ΔW is the update amount of the large language model weight matrix; B and A are trainable low-rank matrices; d is the input dimension of the weight matrix, k is the output dimension of the weight matrix, and r is the rank of the low-rank matrix; the number of parameters is d×r+r×k.
9. The network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training according to claim 6 is characterized in that: The specific formula for using 4-bit quantization for the large language model is: Where s is the scaling factor; W block is the original weight matrix block; z is the zero offset; W int4 To quantify the results.
10. A network security vulnerability analysis system based on large model low-rank adaptive fine-tuning training, characterized by: The system applies a network security vulnerability analysis method based on large model low-rank adaptive fine-tuning training as described in any one of claims 1-9, and the system includes a large language model pre-training module, a network security corpus processing module, a low-rank adaptive fine-tuning module and a network security vulnerability monitoring module; The large language model pre-training module is used to select a large language model and pre-train it; The network security corpus processing module is used to collect and preprocess the network security corpus to construct a supervised fine-tuning dataset; The low-rank adaptive fine-tuning module performs low-rank adaptive training and fine-tuning on the pre-trained large language model based on the supervised fine-tuning dataset to obtain a large model for network security vulnerability analysis; The network security vulnerability monitoring module monitors and analyzes the network system in real time based on the network security vulnerability analysis model, thereby identifying and outputting potential security vulnerabilities and attack patterns.
Citation Information
Patent Citations
Method and apparatus for constructing cybersecurity knowledge graphs for dynamic threat analysis
CN110113314B
Network security vulnerability detection method and system based on big data analysis
CN115412354A
Cited By
Network alarm identification method and system based on data distillation and efficient parameter fine tuning
CN121509184A