Large language model fine tuning data set construction method and system for network security scene

By collecting, converting, expanding and cleaning network security data, a high-quality fine-tuning data set is built, which solves the problem of scarcity of large language models in network security scenarios, improves the performance and applicability of small models, and supports intelligent application of network security.

CN120372275APending Publication Date: 2025-07-25HUAZHONG UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510258058.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The fine-tuned data sets of large language models in the network security scenarios in the prior art are scarce and have low quality, resulting in insufficient application performance of small models in the field of network security and cannot meet actual needs.

Method used

By collecting network security raw data, converting it into question-and-answer data, performing deep expansion and breadth expansion, and combining the prompt word generation and data cleaning of large language models, a high-quality fine-tuned data set is built, including clustering, standardized processing and deduplication operations in the data cleaning stage.

Benefits of technology

A large-scale fine-tuning data set in the field of network security with wide coverage, high quality and strong adaptability has been built, which has significantly improved the fine-tuning performance and practical value of small models, and provided important support for intelligent application of network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372275A_ABST
    Figure CN120372275A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model fine tuning data set construction method and system for a network security scene. The method comprises the following steps: firstly, collecting network security original data, and converting the network security original data into question-answer pair data; performing deep expansion and breadth expansion on the question and answer pair data to obtain expanded question and answer pair data; and performing data cleaning on the expanded question and answer pair data to obtain a final data set. By integrating diversified data sources and combining cue word generation, extension and data cleaning technologies of a large language model, a high-quality large model fine tuning data set in the network security field is efficiently constructed at low cost, and the constructed data set has the characteristics of wide coverage range, high quality and high adaptability; the fine tuning performance and the practical value of the small model in the network security field can be remarkably improved, and important support is provided for development of intelligent application in a network security scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dataset construction, and in particular to a method and system for constructing a large language model fine-tuning dataset for network security scenarios. Background Art

[0002] In recent years, with the rapid development of large language models (LLMs), they have made breakthrough progress in many fields such as natural language processing and image recognition, and have gradually been applied to the field of network security. With its powerful generalization ability, ability to understand complex tasks, and ability to process multimodal data, large language models have brought new solutions to tasks such as threat detection, vulnerability mining, security log analysis, and intrusion behavior tracing in network security scenarios. However, due to the extremely large number of parameters of large language models (usually reaching billions or even hundreds of billions), they face high computing and storage costs in practical applications, which is particularly evident in resource-constrained edge devices or efficient real-time response scenarios.

[0003] In contrast, small models (models with fewer parameters) have become an important application choice in actual network security scenarios due to their low computing requirements and suitability for deployment on edge devices. However, the generalization ability and knowledge coverage of small models are limited, and they show obvious deficiencies in the field of network security. The knowledge of network security scenarios is highly professional, timely, and complex, such as detection rules for specific malicious behaviors, vulnerability exploitation methods for specific protocols, etc. This knowledge is usually not fully covered by ordinary training corpora, resulting in poor performance of small models in specific tasks in the security field. Therefore, fine-tuning the small model in the field and giving it stronger network security capabilities has become an important means to improve its practicality.

[0004] However, there is a relative lack of high-quality network security fine-tuning datasets. First, network security data is highly private and sensitive, and many high-value security data cannot be made public, resulting in limited access to training data. Secondly, existing public network security datasets are mostly used for specific tasks (such as malware classification, intrusion detection, etc.), rather than specifically designed for fine-tuning large language models. The diversity and coverage of the data are difficult to meet actual needs. In addition, there is currently no systematic and reusable method for building a network security large language model fine-tuning dataset, which means that developers need to invest a lot of manpower and time costs when building a fine-tuning dataset, and the data quality is difficult to guarantee, which directly affects the effect of model fine-tuning and actual application performance.

[0005] Therefore, there is an urgent need for a method for constructing a fine-tuning dataset of large language models for network security scenarios to systematically and efficiently construct a high-quality network security fine-tuning dataset, so as to support the fine-tuning and actual deployment requirements of network security small models and promote the development of intelligent technologies in the field of network security. Summary of the Invention

[0006] The present invention provides a method and system for constructing a fine-tuning dataset of large language models for network security scenarios, which solves the technical problems of data scarcity and low data quality in the prior art.

[0007] The present invention provides a method for constructing a fine-tuning dataset of large language models for network security scenarios, including:

[0008] Collecting network security raw data;

[0009] Converting the network security raw data into question-and-answer pair data;

[0010] Performing in-depth expansion and breadth expansion on the question-and-answer pair data to obtain expanded question-and-answer pair data;

[0011] Cleaning the expanded question-and-answer pair data to obtain a final dataset.

[0012] Specifically, the converting the network security raw data into question-and-answer pair data includes:

[0013] Extracting question-and-answer pair data from the network security raw data through regular expressions and / or keyword matching;

[0014] And / or,

[0015] Inputting the network security raw data into a large language model, and the large language model outputs question-and-answer pair data based on pre-designed prompt words.

[0016] Specifically, the performing in-depth expansion and breadth expansion on the question-and-answer pair data to obtain expanded question-and-answer pair data includes:

[0017] Inputting the question-and-answer pair data into a large language model, and the large language model outputs in-depth expanded and / or breadth expanded question-and-answer pair data based on pre-designed prompt words.

[0018] Specifically, the cleaning the expanded question-and-answer pair data to obtain a final dataset includes:

[0019] Clustering the expanded question-and-answer pair data to obtain question-and-answer pair data of different clusters; inputting the question-and-answer pair data of different clusters into a question-and-answer pair scoring model, and outputting question-and-answer pair data with higher scores;

[0020] And / or,

[0021] Normalize the expanded Q&A pair data to obtain normalized Q&A pair data;

[0022] And / or,

[0023] Count the number of words in the expanded Q&A pair data. If the most numerous words contain numbers and / or preset symbols, delete the corresponding Q&A pairs;

[0024] And / or,

[0025] Segment and divide the expanded Q&A pair data into sentences and paragraphs, and retain the Q&A pairs where the number of sentences and paragraphs is greater than or equal to a preset first threshold;

[0026] And / or,

[0027] Count the number of stop words in the expanded Q&A pair data, and retain the Q&A pairs where the number of stop words is less than or equal to a preset second threshold.

[0028] Specifically, cluster the expanded Q&A pair data to obtain Q&A pair data of different clusters; input the Q&A pair data of different clusters into a Q&A pair scoring model, and output Q&A pair data with higher scores, including:

[0029] Concatenate the expanded Q&A pair data to obtain concatenated Q&A pair data;

[0030] Perform vectorization processing on the concatenated Q&A pair data to obtain vectorized Q&A pairs;

[0031] Initialize the vectorized Q&A pairs to obtain spatial vector points, forming a complete vector space;

[0032] Select any Q&A pair space vector point in the vector space as point P, and determine whether the number of Q&A pair space vector points contained in the neighborhood of point P exceeds a preset third threshold;

[0033] If it does not exceed the preset third threshold, mark point P as a noise point and skip the clustering process of this point;

[0034] If it exceeds the preset third threshold, mark point P as a clustering center, and cluster all the Q&A pair space vector points in the neighborhood of point P; repeat the above process until all the Q&A pair space vector points in the vector space have been processed;

[0035] Calculate the semantic similarity between Q&A pair data in various clusters, and mark the Q&A pair data with the semantic similarity greater than or equal to a preset fourth threshold as duplicate Q&A pairs;

[0036] Input the duplicate Q&A pairs into the Q&A pair scoring model respectively, and retain the Q&A pair data with a higher score.

[0037] The present invention provides a large language model fine-tuning dataset construction system for network security scenarios, including:

[0038] An original data acquisition module for acquiring network security original data;

[0039] A data conversion module for converting the network security original data into Q&A pair data;

[0040] A data expansion module for performing in-depth and breadth expansion on the Q&A pair data to obtain expanded Q&A pair data;

[0041] A data cleaning module for cleaning the expanded Q&A pair data to obtain a final dataset.

[0042] Specifically, the data conversion module includes:

[0043] A first data conversion unit for extracting Q&A pair data from the network security original data through regular expressions and / or keyword matching;

[0044] And / or,

[0045] A second data conversion unit for inputting the network security original data into a large language model, and the large language model outputs Q&A pair data based on pre-designed prompt words.

[0046] Specifically, the data expansion module is specifically configured to input the Q&A pair data into a large language model, and the large language model outputs in-depth and / or breadth-expanded Q&A pair data based on pre-designed prompt words.

[0047] Specifically, the data cleaning module includes:

[0048] A first data cleaning unit for clustering the expanded Q&A pair data to obtain Q&A pair data in different clusters; inputting the Q&A pair data in different clusters into the Q&A pair scoring model, and outputting the Q&A pair data with a higher score;

[0049] And / or,

[0050] The second data cleaning unit is used to count the number of words in each of the expanded Q&A pair data. If the most numerous words contain numbers and / or preset symbols, delete the Q&A pairs to which they belong;

[0051] and / or,

[0052] The third data cleaning unit is used to split sentences and paragraphs for each of the expanded Q&A pair data, and retain the Q&A pairs whose number of sentences and paragraphs is greater than or equal to a preset first threshold;

[0053] and / or,

[0054] The fourth data cleaning unit is used to count the number of stop words in each of the expanded Q&A pair data, and retain the Q&A pairs whose number of stop words is less than or equal to a preset second threshold;

[0055] and / or,

[0056] The data standardization unit is used to perform standardization processing on each of the expanded Q&A pair data to obtain standardized Q&A pair data.

[0057] Specifically, the first data cleaning unit includes:

[0058] The data splicing subunit is used to splice each of the expanded Q&A pair data to obtain spliced Q&A pair data;

[0059] The data vectorization subunit is used to perform vectorization processing on the spliced Q&A pair data to obtain vectorized Q&A pairs;

[0060] The data initialization subunit is used to initialize the vectorized Q&A pairs to obtain spatial vector points, forming a complete vector space;

[0061] The data comparison subunit is used to select any Q&A pair space vector point in the vector space as point P, and determine whether the number of Q&A pair space vector points contained in the neighborhood of point P exceeds a preset third threshold;

[0062] The noise point marking subunit is used to, if it does not exceed the preset third threshold, mark point P as a noise point and skip the clustering process of this point;

[0063] The data clustering subunit is used to, if it exceeds the preset third threshold, mark point P as a clustering center, and cluster all the Q&A pair space vector points in the neighborhood of point P; repeat the above process until all the Q&A pair space vector points in the vector space have been processed;

[0064] A repeated Q&A pair marking subunit, which is used to calculate the semantic similarity between Q&A pair data in various clusters, and mark the Q&A pair data whose semantic similarity is greater than or equal to a preset fourth threshold as repeated Q&A pairs;

[0065] A data cleaning subunit, which is used to input the repeated Q&A pairs into a Q&A pair scoring model respectively, and retain the Q&A pair data with higher scores.

[0066] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0067] First, collect raw network security data and convert the raw network security data into Q&A pair data; then perform in-depth expansion and breadth expansion on the Q&A pair data to obtain the expanded Q&A pair data; then perform data cleaning on the expanded Q&A pair data to obtain the final data set. By integrating diverse data sources and combining the prompting word generation, expansion, and data cleaning techniques of large language models, it is possible to efficiently and low-costly construct a high-quality fine-tuning data set for large models in the field of network security. The constructed data set has the characteristics of wide coverage, high quality, and strong adaptability, can significantly improve the performance and practical value of fine-tuning small models in the field of network security, and provides important support for the development of intelligent applications in network security scenarios. Description of the Drawings

[0068] Figure 1 It is a flowchart of a method for constructing a fine-tuning data set for a large language model in a network security scenario provided by an embodiment of the present invention;

[0069] Figure 2 It is a module diagram of a system for constructing a fine-tuning data set for a large language model in a network security scenario provided by an embodiment of the present invention. Detailed Embodiments

[0070] By providing a method and system for constructing a fine-tuning data set for a large language model in a network security scenario, the embodiments of the present invention solve the technical problems of data scarcity and low data quality in the prior art.

[0071] The technical solutions in the embodiments of the present invention for solving the above technical problems generally have the following ideas:

[0072] The embodiments of the present invention propose a systematic method for constructing a fine-tuning data set for large models in the field of network security, which includes four core stages: data collection, data generation, data enhancement, and data cleaning. It systematically solves the key pain points in the prior art and significantly improves the quality and applicability of the fine-tuning data set for network security large models. The specific content is as follows:

[0073] Data collection phase: Obtain comprehensive and diverse raw cybersecurity data from various sources such as security books and papers, vulnerability databases, security forums and communities, security reports, open-source data, and source code data, laying the foundation for building a high-quality dataset.

[0074] Data generation phase: Leverage large language models (LLMs) to deeply understand and process plain text data. By designing precise prompts, guide the model to generate high-quality question-and-answer pairs, generating rich and highly domain-relevant fine-tuning data from different dimensions to meet the diverse needs of fine-tuning cybersecurity large models. In this process, the large language model does not simply convert plain text into questions and answers, but needs to generate in-depth and practical question-and-answer pairs by combining context logic, content relevance, and domain expertise. For some highly specialized content, such as vulnerability descriptions and attack principles, the model needs to show pertinence and specificity in question design.

[0075] Data augmentation phase: Further expand and enrich the question-and-answer dataset through deep and broad expansion, enabling it to cover more diverse cybersecurity scenarios and tasks.

[0076] Data cleaning phase: Dedup and standardize the generated and augmented data, and combine a high-quality screening mechanism to finally obtain a high-quality cybersecurity fine-tuning dataset. Among them, the main purpose of the deduplication operation is to identify and delete redundant question-and-answer pairs in the data to improve the quality and effectiveness of the dataset. Specifically, the embodiments of the present invention can automatically adjust the clustering process based on the characteristics of the question-and-answer pairs through a density-based clustering enrichment, and perform efficient classification according to data structures with different densities, thereby effectively screening out duplicate question-and-answer pairs and ensuring the diversity and quality of the dataset.

[0077] The embodiments of the present invention ensure the comprehensiveness, adaptability, and high quality of the dataset through systematic design, significantly enhancing the practical value and performance of the fine-tuning model in the field of cybersecurity.

[0078] To better understand the above technical solutions, the above technical solutions will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0079] As Figure 1 shown, the method for constructing a large language model fine-tuning dataset for cybersecurity scenarios provided by the embodiments of the present invention includes:

[0080] Step S110: Collect raw cybersecurity data;

[0081] Specifically describe this step. Obtain comprehensive and diverse raw data from multi-channel and multi-type data sources, specifically including the following aspects:

[0082] Security Books and Papers: Extract theoretical knowledge and research results related to network security from professional books and academic papers, covering basic concepts, methodologies, and cutting-edge technologies, etc.

[0083] Vulnerability Databases: Collect publicly disclosed vulnerability information, such as CVE databases, vulnerability reports, etc., and extract vulnerability descriptions, exploitation methods, and their repair solutions.

[0084] Security Forums and Communities: Extract discussion content from security technology forums and community exchanges, covering actual security incidents, solutions to technical problems, and experts' experience sharing.

[0085] Security Reports: Utilize security reports released by network security companies and research institutions to obtain high-value information such as security threat trends, attack analysis, and countermeasure strategies.

[0086] Open-Source Data: Extract code samples and cases related to network security from open-source code libraries, open-source tools, and public datasets.

[0087] Source Code Data: Collect source code data related to network security, including implementation codes of security tools and sample codes for vulnerability exploitation.

[0088] Step S120: Convert the original network security data into Q&A pair data;

[0089] Specifically explain this step. Converting the original network security data into Q&A pair data includes:

[0090] Extract Q&A pair data from the original network security data through regular expressions and / or keyword matching;

[0091] Specifically, for data with a high degree of formatting (such as Q&A posts or community discussions in forums), use regular expressions to efficiently extract Q&A pairs. For example:

[0092] Design corresponding regular expressions to quickly match Q&A paragraphs for common Q&A markers (such as Q:xxx and A:xxx).

[0093] Write exclusive extraction rules for the fixed formats of some specific platforms (such as the "Question" tag and "Accepted Answer" tag in StackOverflow). Example regular matching pattern:

[0094] pattern=r"(Q[::]?\s*(.*?)\n.*?A[::]?\s*(.*?)(?=\nQ|\Z))"

[0095] For discussion data with a lower degree of structure, a keyword matching algorithm is used to identify question sentences and their corresponding answers. For example, the question is located by identifying common question keywords (such as "what", "how", "why", etc.), and the corresponding answer is extracted based on its context.

[0096] Question identification: Scan the text for question keywords (such as "What is X vulnerability").

[0097] Answer extraction: Extract consecutive sentences from the text paragraph after the question until the next question appears or the text ends. In the extraction process, in order to ensure the accuracy and practicality of the question-answering data, denoising and data cleaning mechanisms are also combined:

[0098] Remove questions and answers that are not related to network security (such as purely daily life questions).

[0099] Filter and retain question-and-answer data with high domain relevance (for example, check whether the answer contains words in the network security field such as "vulnerability", "attack", and "repair" through keywords).

[0100] Remove redundant information, such as irrelevant advertising content, user signatures, etc.

[0101] and / or,

[0102] The original network security data is input into the large language model, and the large language model outputs question-answer pair data based on pre-designed prompt words.

[0103] Specifically, the collected plain text data is first input into the model in the form of paragraphs or articles. For example, a text describing a security vulnerability may be: "CVE-2002-1371 is a vulnerability in Apple Common Unix Printing System (CUPS). An attacker can exploit this vulnerability by sending a specially crafted print request, which may lead to a denial of service attack or even allow the attacker to execute arbitrary instructions." To ensure that the model generates high-quality questions and answers, prompt words suitable for the field of network security are designed, such as: "I will send you a text about network security. Please deeply understand the text content and try to convert it into multiple sets of dialogue-style texts, which will be returned in JSON format. When generating these question-answer pairs, you need to maintain a certain contextual logic and chronological order, while avoiding ambiguous questions. The generated question-answer pairs include explanations of the content, supplementary background information, and solutions to potential problems."

[0104] Guided by prompts, the model deeply analyzes the text content and generates multiple question-and-answer pairs. For example, for text describing a certain vulnerability, the model can generate the following question-and-answer pairs: The questions include "What is the CVE-2002-1371 vulnerability?", "What threats can the CVE-2002-1371 vulnerability pose to the system?", and "How to mitigate the risks of the CVE-2002-1371 vulnerability?", and corresponding detailed answers are generated.

[0105] To ensure the quality of the generated questions and answers, a series of quality control measures are also adopted. First, a logical check is performed on the generated question-and-answer pairs to ensure that the question-and-answer logic is clear and the answer content is accurate, avoiding content misdirection caused by incorrect generation. Second, through the optimization of prompts, the generation of vague or meaningless questions is avoided, such as the vague question "Do you know this vulnerability?". Finally, to ensure that the generation of questions and answers meets the specific requirements of the network security field, domain-specific content is added in the design of prompts, making the generated questions and answers more in line with the actual application scenario.

[0106] It should be noted here that the design of prompts is crucial. The prompts need to cover specific task descriptions and add necessary restrictive conditions to ensure that the output results meet expectations. For example, the task descriptions can include generating question-and-answer pairs for specific network security scenarios, deeply explaining security concepts, and proposing targeted solutions. The restrictive conditions can include output format requirements (such as JSON format), topic scope (such as vulnerability description, attack defense, intrusion detection, etc.), as well as the logical order and content clarity to be followed during the generation process. Well-designed prompts can maximize the knowledge reserve of the teacher model and ensure that the generated data has a high level of domain relevance and structurality. The prompt design can be based on the following multiple generation templates to meet the needs of different tasks:

[0107] 0-shot template: Directly guide the model to generate the required question-and-answer data through a concise and clear task description. For example: "Please generate detailed question-and-answer pairs for network security vulnerabilities. The questions need to be clear, and the answers need to include vulnerability descriptions, scope of influence, and repair suggestions."

[0108] 1-shot template: Provide an example question-and-answer pair to give the model a specific reference framework. For example: "The following is an example question and answer: Question: What is the CVE-2021-1234 vulnerability? Answer: CVE-2021-1234 is a remote code execution vulnerability in a specific application. Attackers can gain full control of the system through this vulnerability. Please generate similar question-and-answer pairs based on the following text: …"

[0109] Few-shot Template: By providing multiple example question-answer pairs, further restrict the output format and style of the model. For example, provide 3-5 sets of question-answer pairs to enable the model to capture more complex patterns.

[0110] Chain-of-Thought (CoT) Template: Guide the model to generate question-answer data with a reasoning process. For example: "Please analyze the following description of a cybersecurity vulnerability step by step and generate corresponding questions and detailed answers, ensuring that the exploitation method, scope of impact, and solution of the vulnerability are reflected in the answers."

[0111] Role-Playing Template: Set a specific identity background to guide the model to generate question-answer data from a specific perspective. For example: "Suppose you are a cybersecurity expert. Please generate questions and answers based on the following description. Your answers need to have clear professional logic and rich details."

[0112] Expert Knowledge Injection Template: Supplement domain background information by embedding highly specialized or knowledge points that may be unknown to the large language model in the prompt, and guide the model to generate more accurate question-answer pairs. For example: "The following are the security repair details and industry best practices for a specific vulnerability. Please generate question-answer pairs based on this information, ensuring that specific repair steps and relevant technical details are cited in the answers."

[0113] By calling the APIs of large language models (such as OpenAI API, Anthropic API, etc.) and combining the above prompt templates, gradually generate diverse structured question-answer data. In actual operation, not only can the prompts be dynamically adjusted according to task requirements, but also the generation results can be optimized through multiple rounds of interaction. For example, for inaccuracies or incompleteness that may exist in the generated data, the prompts can be adjusted again through the model and regenerated to ensure that the data quality meets the expectations.

[0114] In this embodiment, the large language models are ChatGPT, GPT-4, Claude, Gemini, etc.

[0115] Step S130: Deeply expand and broadly expand the question-answer pair data to obtain the expanded question-answer pair data;

[0116] Specifically explain this step. Deeply expand and broadly expand the question-answer pair data to obtain the expanded question-answer pair data, including:

[0117] Input the question-answer pair data into the large language model, and the large language model outputs the deeply expanded and / or broadly expanded question-answer pair data based on the pre-designed prompts.

[0118] Specifically, depth expansion aims to enhance the complexity and reasoning depth of Q&A data, making the generated Q&A more valuable for practical applications and having a logical hierarchy. In this stage, the original data is complicated through strategies such as adding constraints, deepening the problem background, and introducing reasoning chains to generate higher-quality Q&A data.

[0119] Adding constraints makes the problem closer to practical applications by adding specific restrictive conditions or scenario requirements to existing Q&A pairs. For example:

[0120] Original question: How to detect SQL injection attacks existing in the network?

[0121] Expanded question: On resource-constrained embedded devices, how to efficiently detect SQL injection attacks in the network?

[0122] Deepening the problem background makes the problem more complicated by adding background information to the problem, requiring more context factors to be considered when answering. For example:

[0123] Original question: How to fix the CVE-2021-44228 vulnerability?

[0124] Expanded question: The CVE-2021-44228 vulnerability (Log4j vulnerability) was discovered in an enterprise distributed microservices architecture that is implemented in multiple languages, including Java and Python. How to efficiently fix this vulnerability in such an architecture and ensure that the fix does not have a significant impact on system performance?

[0125] Designing the reasoning chain guides the question to design more reasoning steps, making the answer require multiple layers of logical derivation, thereby enhancing the depth of the question. For example:

[0126] Original question: What are the main protection measures against DDoS attacks?

[0127] Expanded question: When facing a DDoS attack, starting from the analysis of the source of the attack traffic, explain how to identify abnormal traffic through traffic analysis techniques, restrict traffic using network devices, and combine cloud security services to prevent the attack.

[0128] During the depth expansion process, the large language model is used to complete the complication of the question through carefully designed prompts. For example, the prompts can be: "Please add a constraint to the following question to make it closer to a specific scenario."; "According to the background of the following question, add more technical details to make the question more complicated."; "Design a reasoning chain for the following question, requiring the answer to explain the key steps in sequence." Through depth expansion, the quality and complexity of Q&A data can be significantly improved, making the model perform better when dealing with actual network security scenarios.

[0129] The focus of breadth expansion is to cover a wider range of task categories and application areas by adding new topics and scenarios, making up for the lack of topics in existing datasets and generating diversified data. Unlike deep expansion, breadth expansion does not focus on mining the details of a single topic, but guides the generation of new instructions across domains and topics to ensure that the dataset has wide applicability.

[0130] Domain coverage expansion analyzes the topic distribution from existing datasets, prioritizes fields or tasks that have not been fully covered (such as cryptography, blockchain, privacy protection, etc.), and uses prompt words to guide the model to generate question-answer pairs on new topics to supplement the areas not covered in the dataset.

[0131] The diversification of task types has expanded from a single task type to multiple task types, including generative tasks, classification tasks, reasoning tasks, etc. For example:

[0132] Original Task Type: "Describe how to fix a SQL injection vulnerability."

[0133] Breadth-expanded generation: Reasoning task: "Is there a SQL injection vulnerability in the following code snippet? If so, please explain why." Classification task: "Classify the following network attack types (DDoS, XSS, SQL injection) and explain their main characteristics." Generation task: "Generate a technical document on the best practices of SQL injection protection."

[0134] The long-tail topic expansion guides the model to generate rare but valuable long-tail topic questions, supplementing the data of unpopular or cutting-edge fields. For example:

[0135] Original instruction: "Describe common types of cyber attacks."

[0136] Breadth expansion generation: "How does quantum computing affect existing cryptography techniques?"; "What are the unique security challenges in edge computing devices?"

[0137] In the process of breadth expansion, the large language model is used to generalize the question through carefully designed prompt words. For example, the prompt word can be:

[0138] “Generate semantically related but topically different new questions from the following seed questions, ensuring that the questions belong to new domains or industries.”

[0139] “Based on existing instructions, expand and generate instructions covering new task types (such as generation, classification, and reasoning), and ensure that the logic is reasonable.”

[0140] “Generate interdisciplinary (e.g., law, sociology) instruction based on cybersecurity, covering new issues in cybersecurity.”

[0141] Step S140: Clean the expanded Q&A pair data to obtain the final dataset.

[0142] A specific description of this step is as follows. Cleaning the expanded Q&A pair data to obtain the final dataset includes:

[0143] Cluster the expanded Q&A pair data to obtain Q&A pair data of different clusters; input the Q&A pair data of different clusters into the Q&A pair scoring model, and output the Q&A pair data with higher scores.

[0144] Specifically, clustering the expanded Q&A pair data to obtain Q&A pair data of different clusters; inputting the Q&A pair data of different clusters into the Q&A pair scoring model, and outputting the Q&A pair data with higher scores, including:

[0145] Concatenate the expanded Q&A pair data to obtain the concatenated Q&A pair data. For example, for the question "How to prevent SQL injection attacks?" and its answer "Attacks can be prevented through parameterized queries and input validation", the concatenated input text is: "Question: How to prevent SQL injection attacks? Answer: Attacks can be prevented through parameterized queries and input validation."

[0146] Perform vectorization processing on the concatenated Q&A pair data to obtain vectorized Q&A pairs. Specifically, input the concatenated Q&A pair data into the Sentence-BERT model for vectorization processing to obtain vectorized Q&A pairs. Through this process, each Q&A pair can be converted into a 768-dimensional semantic vector, so that each pair of Q&A has a clear semantic position in the high-dimensional space.

[0147] Initialize the vectorized question-answer pairs to obtain spatial vector points, forming a complete vector space. Specifically, the 768-dimensional vector corresponding to each question-answer pair can be regarded as a point of the question-answer pair in the high-dimensional space. This space is usually a Euclidean space. In this space, the position of each question-answer pair through its vector representation reflects their semantic information. The closer the vector distances of two question-answer pairs are, the higher their semantic similarity is, and the farther the distances are, the greater their semantic differences are. In this way, each vector becomes a spatial point in the high-dimensional space, and different spatial points can be clustered into the same cluster according to their semantic similarity. Suppose there is a pair of question and answer: Question: "How to prevent SQL injection attacks?" Answer: "Attacks can be prevented through parameterized queries and input validation." After being processed by the Sentence-BERT model, this pair of question-answer pairs is converted into a 768-dimensional vector, v1 = [0.23, 0.15, 0.94,..., 0.32], which is the position of this question-answer pair in the 768-dimensional space. Each dimension of this vector represents a certain semantic feature of the question-answer. For example, some dimensions may reflect the security theme of "SQL injection attacks", while other dimensions may describe the relevance of protection measures such as "parameterized queries" or "input validation".

[0148] Select any question-answer pair spatial vector point in the vector space as point P, and determine whether the number of question-answer pair spatial vector points contained in the neighborhood of point P exceeds a preset third threshold;

[0149] If it does not exceed the preset third threshold, mark point P as a noise point and skip the clustering process of this point;

[0150] If it exceeds the preset third threshold, mark the point P as a clustering center, cluster all the question-answer pair spatial vector points in the neighborhood of point P, and add them to the current clustering cluster. This process will recursively check the neighborhoods of these spatial vector points until no new points can be added; repeat the above process until all the question-answer pair spatial vector points in the vector space have been processed;

[0151] Calculate the semantic similarity between the question-answer pair data in each cluster, and mark the question-answer pair data with a semantic similarity greater than or equal to a preset fourth threshold as duplicate question-answer pairs;

[0152] Specifically, through the formula Calculate the semantic similarity between question-answer pairs; where A and B are the vector representations of question-answer pairs, · represents the dot product, and ∣∣A∣∣ and ∣∣B∣∣ are the norms of the vectors respectively.

[0153] To screen for duplicate question-and-answer pairs, a similarity threshold is set, such as 0.9. When the similarity between two question-and-answer pairs is higher than this threshold, they are marked as duplicate data.

[0154] The duplicate question-and-answer pairs are respectively input into the question-and-answer pair scoring model, and the question-and-answer pair data with a higher score is retained.

[0155] Specifically, the reward-model-deberta-v3-large-v2 model developed by OpenAssistant is used to score the question-and-answer pairs. For each pair of questions and answers, the model scores them based on their semantics and logic. The question-and-answer pair with a higher score is considered better because its content is more comprehensive and the expression is clearer. The question-and-answer pair with a higher score will be retained as part of the final dataset.

[0156] Suppose there are the following two question-and-answer pairs. They are of similar length, but one answer is relatively concise while the other provides a more detailed and clearly structured answer:

[0157] Question-and-answer pair A:

[0158] Question: How to prevent SQL injection attacks?

[0159] Answer: SQL injection attacks can be prevented by using parameterized queries. This method ensures that the data entered by users is not directly embedded in the SQL statement, thus avoiding the execution of malicious SQL code. In this way, the program can safely process user input and prevent malicious attacks.

[0160] Question-and-answer pair B:

[0161] Question: How to prevent SQL injection attacks?

[0162] Answer: SQL injection attacks can be prevented in various ways. The most important thing is to use parameterized queries. Parameterized queries can ensure that the input data is not directly spliced into the SQL statement, thus preventing the execution of malicious code. In addition, techniques such as input validation and Web Application Firewall (WAF) can be combined to further strengthen the protection.

[0163] Although the answer in question-and-answer pair A is concise, it only mentions one protection method - parameterized queries. Although this is an effective method, it does not further expand or discuss other possible protection measures and seems rather brief. Question-and-answer pair B provides multi-angle protection suggestions. In addition to parameterized queries, it also mentions additional protection measures such as input validation and Web Application Firewall (WAF). This not only covers more protection methods but also provides a more comprehensive solution for users.

[0164] Through scoring, Q&A pair A scored 0.8067 and Q&A pair B scored 1.4977. Since Q&A pair B has a higher score, it indicates that it is superior to Q&A pair A in terms of content integrity and information breadth. Therefore, Q&A pair B will be retained, while Q&A pair A may be excluded due to lack of sufficient details and richness.

[0165] and / or,

[0166] Normalize the data of each extended Q&A pair to obtain normalized Q&A pair data;

[0167] Specifically, integrate all Q&A data into a unified JSON format, adopting a three-part structure of instruction, input, and output. In this process, convert each Q&A pair into a normalized data entry, where instruction represents the specific description of the task, input represents the context information related to the task, and output represents the answer or execution result of the task. Specifically, for Q&A data, instruction corresponds to the question description, input is empty or provides necessary background information, and output is the answer content. For example:

[0168] {

[0169] "instruction":"How to prevent SQL injection attacks?",

[0170] "input":"",

[0171] "output":"SQL injection attacks can be prevented by using parameterized queries and input validation."

[0172] }

[0173] For Q&A with context dependence, such as scenarios involving multi-round conversations or requiring background information, the context can be filled in the input part. For example:

[0174] {

[0175] "instruction":"How to prevent SQL injection attacks?",

[0176] "input":"Scenario: E-commerce website",

[0177] "output":"Prevent SQL injection by enabling a web application firewall and using parameterized queries."}

[0178] and / or,

[0179] Count the number of words in each Q&A pair data after expansion. If the most frequent words contain numbers and / or preset symbols, it indicates that the document format does not meet the requirements, and the corresponding Q&A pair should be deleted.

[0180] For example:

[0181] Document A: "The CVE-2021-44228 vulnerability is a remote code execution vulnerability in Apache Log4j, which can be triggered by an attacker through carefully crafted input." It contains "CVE-2021-44228" (a CVE vulnerability number), which belongs to a common format in the field of network security and will not be deleted.

[0182] Document B: "12345!!@#%&*()#CVE". It contains "12345!!@#%&*()", which belongs to symbols and numbers, and the most frequent words are non-alphabetic, so it will be deleted.

[0183] and / or,

[0184] Perform sentence splitting and paragraph segmentation on each Q&A pair data after expansion, delete documents with too little content, and retain Q&A pairs whose sentence count and paragraph count are greater than or equal to a preset first threshold.

[0185] For example:

[0186] Document C: "SQL injection attack refers to an attacker injecting malicious SQL statements into input fields, causing database operation anomalies." Although this document mentions SQL injection, it only contains one sentence and cannot provide sufficient background and details, so it needs to be deleted.

[0187] Document D: "First paragraph: Definition of SQL injection; Second paragraph: SQL injection defense methods". Document D contains two paragraphs, and each paragraph is relatively short, which does not meet the quality standard and should be deleted.

[0188] and / or,

[0189] Count the number of stop words in each Q&A pair data after expansion, and retain Q&A pairs whose stop word count is less than or equal to a preset second threshold.

[0190] For example:

[0191] Document E: "In the field of network security, SQL injection attack is a common attack method. It can attack the database through malicious SQL code input by users, thereby destroying the integrity of the database. By using protection measures such as parameterized queries and input validation, such attacks can be effectively avoided."

[0192] This document involves the definition and prevention methods of SQL injection attacks, but contains a large number of stop words, such as "in", "of", "in", "of", etc. By tokenizing the words in the document and counting the number of stop words. The total number of words N total = 40, the number of stop words N stopwords = 15, the proportion of stop words is 37.5%. This indicates that in this document, the proportion of stop words is 37.5%, exceeding the common 30% threshold. Therefore, the proportion of stop words in the document is too high, which may lead to insufficient information in the document and thus be considered a redundant and low-information document.

[0193] such as Figure 2 As shown, the large language model fine-tuning dataset construction system for network security scenarios provided by the embodiments of the present invention includes:

[0194] The original data acquisition module 100 is used to acquire network security original data;

[0195] Specifically, the original data acquisition module 100 is specifically used to obtain comprehensive and diverse original data from multi-channel and multi-type data sources.

[0196] The data conversion module 200 is used to convert network security original data into question-and-answer pair data;

[0197] Specifically, the data conversion module 200 includes:

[0198] The first data conversion unit is used to extract question-and-answer pair data from network security original data through regular expressions and / or keyword matching;

[0199] and / or,

[0200] The second data conversion unit is used to input network security original data into a large language model, and the large language model outputs question-and-answer pair data based on pre-designed prompt words.

[0201] The data expansion module 300 is used to perform in-depth and breadth expansion on the question-and-answer pair data to obtain expanded question-and-answer pair data;

[0202] Specifically, the data expansion module 300 is specifically used to input the question-and-answer pair data into a large language model, and the large language model outputs in-depth and / or breadth-expanded question-and-answer pair data based on pre-designed prompt words.

[0203] The data cleaning module 400 is used to clean the expanded question-and-answer pair data to obtain the final dataset.

[0204] Specifically, the data cleaning module 400 includes:

[0205] The first data cleaning unit is used to cluster the expanded question-and-answer pair data to obtain question-and-answer pair data of different clusters; input the question-and-answer pair data of different clusters into the question-and-answer pair scoring model, and output the question-and-answer pair data with higher scores;

[0206] In this embodiment, the first data cleaning unit includes:

[0207] The data splicing subunit is used to splice the expanded question-and-answer pair data to obtain the spliced question-and-answer pair data;

[0208] The data vectorization subunit is used to perform vectorization processing on the spliced question-and-answer pair data to obtain vectorized question-and-answer pairs;

[0209] The data initialization subunit is used to initialize the vectorized question-and-answer pairs to obtain spatial vector points and form a complete vector space;

[0210] The data comparison subunit is used to select any question-and-answer pair spatial vector point in the vector space as the P point, and judge whether the number of question-and-answer pair spatial vector points contained in the neighborhood of the P point exceeds a preset third threshold;

[0211] The noise point marking subunit is used to, if it does not exceed the preset third threshold, mark the P point as a noise point and skip the clustering process of this point;

[0212] The data clustering subunit is used to, if it exceeds the preset third threshold, mark the P point as a clustering center, cluster all the question-and-answer pair spatial vector points in the neighborhood of the P point, and add them to the current cluster. This process will recursively check the neighborhoods of these spatial vector points until no new points can be added; repeat the above process until all the question-and-answer pair spatial vector points in the vector space have been processed;

[0213] The repeated question-and-answer pair marking subunit is used to calculate the semantic similarity between the question-and-answer pair data in each cluster, and mark the question-and-answer pair data with semantic similarity greater than or equal to a preset fourth threshold as repeated question-and-answer pairs;

[0214] The data cleaning subunit is used to input the repeated question-and-answer pairs into the question-and-answer pair scoring model respectively, and retain the question-and-answer pair data with higher scores.

[0215] And / or

[0216] The second data cleaning unit is used to count the number of words in the expanded question-and-answer pair data. If the most numerous words contain numbers and / or preset symbols, it indicates that the document format does not meet the requirements, and the belonging question-and-answer pairs are deleted;

[0217] And / or

[0218] The third data cleaning unit is used to perform sentence splitting and paragraph splitting on each pair of question and answer data after expansion, delete documents with too little content, and retain pairs of questions and answers whose number of sentences and paragraphs is greater than or equal to a preset first threshold.

[0219] And / or,

[0220] The fourth data cleaning unit is used to count the number of stop words in each pair of question and answer data after expansion, and retain pairs of questions and answers whose number of stop words is less than or equal to a preset second threshold.

[0221] And / or,

[0222] The data standardization unit is used to perform standardization processing on each pair of question and answer data after expansion to obtain standardized pair of question and answer data.

[0223] The embodiments of the present invention propose a complete and systematic method for constructing a fine-tuning data set for large models in the field of network security, covering four core stages: data collection, data generation, data augmentation, and data cleaning. Among them, in the data generation stage, relying on the powerful generation ability of the large language model, high-quality pairs of questions and answers are automatically generated through carefully designed prompt words. In the data augmentation stage, deep expansion and breadth expansion are adopted to improve the complexity, coverage, and diversity of the data set. In the data cleaning stage, a deduplication algorithm based on Sentence-BERT, a unified standard format, and a high-quality screening mechanism based on semantic scoring are introduced. This multi-level and multi-stage cleaning method ensures the uniqueness, consistency, and high quality of the data set, providing more accurate and reliable training data for model fine-tuning. The embodiments of the present invention significantly reduce the cost of constructing a network security fine-tuning data set, improve the practicality and performance of the model in the field of network security, and provide strong support for the development of intelligent network security technology.

[0224] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0225] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.

[0226] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.

[0227] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks. Figure 1 in one or more flows and / or blocks Figure 1 in one or more blocks.

[0228] Details not described in the embodiments of the present invention are well-known techniques to those skilled in the art. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for constructing a fine-tuning dataset of a large language model for network security scenarios, characterized in that, Including: Collecting raw network security data; Converting the raw network security data into Q&A pair data; Performing in-depth expansion and breadth expansion on the Q&A pair data to obtain expanded Q&A pair data; Cleaning the expanded Q&A pair data to obtain the final dataset.

2. The method for constructing a fine-tuning dataset of a large language model for a network security scenario according to claim 1, wherein, The converting the raw network security data into Q&A pair data includes: Extracting Q&A pair data from the raw network security data through regular expressions and / or keyword matching; And / or, Inputting the raw network security data into a large language model, and the large language model outputs Q&A pair data based on pre-designed prompt words.

3. The method for constructing a large language model fine-tuning dataset for a network security scenario according to claim 1, wherein, The performing in-depth expansion and breadth expansion on the Q&A pair data to obtain expanded Q&A pair data includes: Inputting the Q&A pair data into a large language model, and the large language model outputs in-depth expanded and / or breadth expanded Q&A pair data based on pre-designed prompt words.

4. The method for constructing a large language model fine-tuning dataset for a network security scenario according to claim 1, wherein, The cleaning the expanded Q&A pair data to obtain the final dataset includes: Clustering the expanded Q&A pair data to obtain Q&A pair data of different clusters; inputting the Q&A pair data of different clusters into a Q&A pair scoring model and outputting Q&A pair data with higher scores; And / or, Performing normalization processing on the expanded Q&A pair data to obtain normalized Q&A pair data; And / or, Counting the number of words in each expanded Q&A pair data, and if the most numerous words contain numbers and / or preset symbols, deleting the belonging Q&A pairs; And / or, Performing sentence splitting and paragraph splitting on the expanded Q&A pair data, and retaining the Q&A pairs whose number of sentences and paragraphs is greater than or equal to a preset first threshold; And / or, Counting the number of stop words in each expanded Q&A pair data, and retaining the Q&A pairs whose number of stop words is less than or equal to a preset second threshold.

5. The method for constructing a large language model fine-tuning dataset for a network security scenario according to claim 4, wherein, The clustering the expanded Q&A pair data to obtain Q&A pair data of different clusters; Inputting the Q&A pair data of different clusters into a Q&A pair scoring model and outputting Q&A pair data with higher scores includes: Concatenating the expanded Q&A pair data to obtain concatenated Q&A pair data; Performing vectorization processing on the concatenated Q&A pair data to obtain vectorized Q&A pairs; Initializing the vectorized Q&A pairs to obtain spatial vector points and forming a complete vector space; Selecting any Q&A pair spatial vector point in the vector space as point P, and judging whether the number of Q&A pair spatial vector points included in the neighborhood of point P exceeds a preset third threshold; If it does not exceed the preset third threshold, marking point P as a noise point and skipping the clustering process of this point; If it exceeds the preset third threshold, marking point P as a clustering center and clustering all Q&A pair spatial vector points in the neighborhood of point P; repeating the above process until all Q&A pair spatial vector points in the vector space have been processed; Calculating the semantic similarity between Q&A pair data in each cluster, and marking the Q&A pair data whose semantic similarity is greater than or equal to a preset fourth threshold as duplicate Q&A pairs; Input the repeated Q&A pairs into the Q&A pair scoring model respectively, and retain the Q&A pair data with higher scores.

6. A large language model fine-tuning dataset construction system for network security scenarios, characterized in that, It includes: An original data acquisition module for acquiring original network security data; A data conversion module for converting the original network security data into Q&A pair data; A data expansion module for performing in-depth and breadth expansion on the Q&A pair data to obtain expanded Q&A pair data; A data cleaning module for cleaning the expanded Q&A pair data to obtain the final data set.

7. The large language model fine-tuning dataset construction system for network security scenarios according to claim 6, wherein The data conversion module includes: A first data conversion unit for extracting Q&A pair data from the original network security data through regular expressions and / or keyword matching; And / or A second data conversion unit for inputting the original network security data into a large language model, and the large language model outputs Q&A pair data based on pre-designed prompt words.

8. The large language model fine-tuning dataset construction system for network security scenarios according to claim 6, characterized in that, The data expansion module is specifically used to input the Q&A pair data into a large language model, and the large language model outputs in-depth expanded and / or breadth expanded Q&A pair data based on pre-designed prompt words.

9. The large language model fine-tuning dataset construction system for the network security scenario according to claim 6, wherein, The data cleaning module includes: A first data cleaning unit for clustering the expanded Q&A pair data of each item to obtain Q&A pair data of different clusters; inputting the Q&A pair data of different clusters into the Q&A pair scoring model, and outputting Q&A pair data with higher scores; And / or A second data cleaning unit for counting the number of words in the expanded Q&A pair data of each item, and if the most numerous words contain numbers and / or preset symbols, deleting the affiliated Q&A pairs; And / or A third data cleaning unit for splitting sentences and paragraphs of the expanded Q&A pair data of each item, and retaining the Q&A pairs with the number of sentences and paragraphs greater than or equal to a preset first threshold; And / or A fourth data cleaning unit for counting the number of stop words in the expanded Q&A pair data of each item, and retaining the Q&A pairs with the number of stop words less than or equal to a preset second threshold; And / or A data standardization unit for performing standardization processing on the expanded Q&A pair data of each item to obtain standardized Q&A pair data.

10. The large language model fine-tuning dataset construction system for network security scenarios according to claim 9, characterized in that, The first data cleaning unit includes: A data splicing subunit for splicing the expanded Q&A pair data of each item to obtain spliced Q&A pair data; A data vectorization subunit for performing vectorization processing on the spliced Q&A pair data to obtain vectorized Q&A pairs; A data initialization subunit for initializing the vectorized Q&A pairs to obtain spatial vector points and form a complete vector space; A data comparison subunit for selecting any Q&A pair space vector point in the vector space as point P, and judging whether the number of Q&A pair space vector points contained in the neighborhood of point P exceeds a preset third threshold; A noise point marking subunit for, if it does not exceed the preset third threshold, marking point P as a noise point and skipping the clustering process of this point; A data clustering subunit, which is used to mark the P point as a clustering center if it exceeds the preset third threshold, and cluster all the question-and-answer pair space vector points within the neighborhood of the P point; repeat the above process until all the question-and-answer pair space vector points in the vector space have been processed; A repeated question-and-answer pair marking subunit, which is used to calculate the semantic similarity between the question-and-answer pair data in each cluster, and mark the question-and-answer pair data whose semantic similarity is greater than or equal to the preset fourth threshold as repeated question-and-answer pairs; A data cleaning subunit, which is used to input the repeated question-and-answer pairs into the question-and-answer pair scoring model respectively, and retain the question-and-answer pair data with higher scores.