Cybersecurity large model corpus construction method and device, equipment and storage medium
Patent Information
- Application Number
- CN202611045742.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本申请的主要目的在于提供一种网络安全大模型语料构建方法、装置、设备及存储介质,旨在解决针对现有网安大模型语料构建过程中存在的自动化程度低、严重依赖人工主观评估,且缺乏对攻防技战术知识的实战有效性进行客观验证的不足的技术问题
本申请基于预设大模型对网络安全领域的文本数据进行多维度评分,并根据评分结果从文本数据中筛选目标文本,对目标文本执行战术结构化抽取,获得结构化后的靶场实战对抗参数;将结构化后的靶场实战对抗参数映射至网络靶场执行攻防演练任务,并计算各目标文本执行攻防演练任务对应的语料录用评分,将语料录用评分满足预设录用评分阈值的目标文本录入训练语料库。本申请通过筛选网络安全领域的文本,并从文本中提取出关键的攻防技战术,并将其转化为靶场环境支持的攻防演练任务,完成语料量化评分与筛选,将满足靶场实战有效性检验的语料,纳入网络安全垂域大模型语料库。相较于现有网安大模型语料构建过程中存在的自动化程度低、严重依赖人工主观评估,且缺乏对攻防技战术知识的实战有效性进行客观验证的不足,本申请通过将初步清洗分类后的网安技战术语料自动化转化为靶场攻防演练任务,基于对抗环境中的动态执行效果进行效能评估与筛选实现了高质量网安实战语料的自动化、高合格率筛选,从数据源头减少大模型的幻觉,有效提升网安垂域大模型实战能力的上限。
Smart Images

Figure CN122840190A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to methods, apparatus, devices and storage media for constructing large-scale network security model corpora. Background Technology
[0002] The higher the authenticity and effectiveness of practical attack and defense knowledge in cybersecurity corpora, the greater the performance ceiling of model training and its practical application value. Therefore, exploring an efficient, objective, and accurate method for constructing a large-scale cybersecurity model corpus is of great strategic significance for improving the intelligence level of the overall network defense system.
[0003] Currently, the construction of large-scale model corpora in vertical domains mainly relies on a paradigm that combines static rule processing with manual specification annotation. However, this approach has significant shortcomings when applied to the field of cybersecurity: on the one hand, rule-based engines and similarity-based processing remain at the surface-level semantic matching of text, failing to verify the practical effectiveness of code or tactics at the underlying level. A piece of semantically coherent attack code may completely fail in a real environment due to the lack of specific parameters. On the other hand, manual verification that heavily relies on multi-terminal collaboration among experts is not only costly, but the subjective perception of experts is also difficult to cover the ever-changing and complex network attack and defense scenarios. Summary of the Invention
[0004] The main purpose of this application is to provide a method, apparatus, device and storage medium for constructing a large network security model corpus, which aims to solve the technical problems of low automation, heavy reliance on subjective human evaluation and lack of objective verification of the practical effectiveness of attack and defense tactics in the construction of existing large network security model corpora.
[0005] To achieve the above objectives, this application proposes a method for constructing a large-scale network security model corpus, the method comprising: Based on a pre-set large model, text data in the field of cybersecurity is scored in multiple dimensions, and target text is selected from the text data according to the scoring results; Tactical structured extraction is performed on the target text to obtain structured target range combat parameters; The structured target range combat parameters are mapped to the network target range to perform attack and defense exercises, and the corpus admission score corresponding to each target text performing the attack and defense exercises is calculated. The target texts whose corpus selection scores meet the preset selection score threshold are entered into the training corpus.
[0006] In one implementation, before performing multi-dimensional scoring on text data in the cybersecurity field based on a preset large model, and filtering target text from the text data according to the scoring results, the method further includes: Text cleaning is performed on the original multimodal security data to obtain cleaned data; A paragraph-level deduplication mechanism based on the minimum hash algorithm is used to remove duplicate paragraphs from the cleaned data to obtain a candidate corpus set. Based on a pre-defined large model, the texts in the candidate corpus are classified by topic, and text data belonging to the field of cybersecurity are selected.
[0007] In one implementation, the step of performing multi-dimensional scoring on text data in the cybersecurity field based on a preset large model, and filtering target text from the text data according to the scoring results, includes: Based on a pre-defined large model, domain analysis is performed on text data in the field of cybersecurity to obtain a set of cybersecurity subdomains. Based on the weight information corresponding to the set of network security subdomains, the text data is scored using a multi-dimensional weighted score to obtain a comprehensive quality score result. Based on the comprehensive quality score, it is determined whether the text data meets the quality requirements, and the text data that meets the quality requirements is taken as the target text.
[0008] In one implementation, the step of performing tactical structured extraction on the target text to obtain structured target range combat parameters includes: Based on the information extraction model and preset prompt word template, target range adversarial elements are extracted from the target text. The target range adversarial elements include the target target range environment, attack tactics number, payload type, executable code, and initial target range combat adversarial parameters of expected execution results. The initial target range combat parameters are converted into a preset structured format to obtain the structured target range combat parameters.
[0009] In one implementation, the step of mapping the structured target range combat parameters to the network target range to execute attack and defense exercises, and calculating the corpus acceptance score corresponding to each target text executing the attack and defense exercises, includes: The structured target range combat parameters are mapped to ternary artifacts that can be deployed in the network target range. The attack and defense exercise is performed based on the triplet artifact, and the corpus admission score corresponding to each target text performing the attack and defense exercise is calculated.
[0010] In one implementation, the triple artifact includes an environment template, an attack execution script, and a state detection probe. The step of performing an attack and defense exercise based on the triple artifact and calculating the corpus acceptance score corresponding to each target text performing the attack and defense exercise includes: The environment template is deployed in the network test range, and attack and defense drills are performed based on the attack execution script. The target machine state vector before the attack and the target machine state vector after the attack are collected based on the state detection probe. The state transition vector is determined based on the target machine state vector before the attack and the target machine state vector after the attack. Based on the execution result of the attack execution script and the state consistency between the state transition vector and the expected execution result, the corpus selection score for each target text executing the attack and defense exercise task is determined.
[0011] In one implementation, determining the corpus acceptance score for each target text performing the attack and defense exercise task based on the execution result corresponding to the attack execution script, the state transition vector, and the state consistency between the expected execution result and the state execution result includes: Determine whether the execution result corresponding to the attack execution script was successfully executed within the limited time, and obtain the execution success indication function; The state fit is determined based on the similarity between the actual adversarial feature set corresponding to the state transition vector and the expected feature set corresponding to the expected execution result. The weighted summation of the execution success indication function and the state consistency is used to determine the corpus selection score corresponding to each target text executing the attack and defense exercise task.
[0012] Furthermore, to achieve the above objectives, this application also proposes a network security large-scale model corpus construction device, which includes: The text filtering module is used to perform multi-dimensional scoring on text data in the field of cybersecurity based on a preset large model, and to filter target text from the text data according to the scoring results. The tactical extraction module is used to perform tactical structured extraction on the target text to obtain structured target range combat parameters. The task training module is used to map the structured target range combat parameters to the network target range to execute attack and defense training tasks, and to calculate the corpus admission score corresponding to each target text executing the attack and defense training task. The corpus entry module is used to enter target texts whose corpus entry scores meet the preset entry score thresholds into the training corpus.
[0013] Furthermore, to achieve the above objectives, this application also proposes a corpus construction device for a large network security model. The corpus construction device for the large network security model includes: a memory, a processor, and a corpus construction program for the large network security model stored in the memory and executable on the processor. The corpus construction program for the large network security model is configured to implement the corpus construction method for the large network security model.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium storing a corpus construction program for a large network security model, wherein the corpus construction program for the large network security model is executed by a processor to implement the corpus construction method for the large network security model.
[0015] One or more technical solutions proposed in this application have at least the following technical effects: This application uses a pre-defined large-scale model to perform multi-dimensional scoring on text data in the cybersecurity field. Based on the scoring results, target texts are selected from the text data, and tactical structure extraction is performed on the target texts to obtain structured target range combat parameters. These structured target range combat parameters are then mapped to cyber range attack and defense exercises, and a corpus adoption score is calculated for each target text performing these exercises. Target texts whose corpus adoption scores meet a pre-defined adoption score threshold are added to the training corpus. This application completes the corpus quantification, scoring, and selection by screening texts in the cybersecurity field, extracting key attack and defense techniques and tactics from the texts, and transforming them into attack and defense exercises supported by the target range environment. Corpus data that meets the effectiveness verification requirements in the target range is then included in the cybersecurity vertical domain large-scale model corpus. Compared to existing cybersecurity large-scale model corpus construction processes, which suffer from low automation, heavy reliance on subjective human evaluation, and a lack of objective verification of the practical effectiveness of offensive and defensive tactical knowledge, this application automates the transformation of the initially cleaned and classified cybersecurity tactical terminology corpus into range offensive and defensive exercise tasks. Based on the dynamic execution effect in the adversarial environment, it conducts performance evaluation and screening, achieving automated and high-pass-rate screening of high-quality cybersecurity practical corpus. This reduces the illusion of large models from the data source and effectively improves the upper limit of the practical capabilities of cybersecurity vertical domain large-scale models. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating an embodiment of the method for constructing a large-scale network security model corpus in this application. Figure 2 This is a text deduplication diagram provided in Embodiment 1 of the method for constructing a large network security model corpus in this application; Figure 3This is a flowchart illustrating Embodiment 2 of the method for constructing a large network security model corpus in this application. Figure 4 This is a flowchart illustrating Embodiment 3 of the method for constructing a large network security model corpus in this application. Figure 5 This is a schematic diagram of the overall architecture provided in Embodiment 3 of the method for constructing a large network security model corpus in this application; Figure 6 This is a schematic diagram of the module structure provided in Embodiment 1 of the network security large model corpus construction device of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the network security large model corpus construction method in the embodiments of this application.
[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0022] The main solution of this application embodiment is as follows: based on a preset large model, text data in the field of network security is scored in multiple dimensions, and target texts are selected from the text data according to the scoring results. Tactical structured extraction is performed on the target texts to obtain the target range combat parameters. The target range combat parameters are mapped to the network target range to perform attack and defense exercises, and the corpus admission score corresponding to each target text performing the attack and defense exercises is calculated. Target texts whose corpus admission scores meet the preset admission score threshold are entered into the training corpus.
[0023] The existing cybersecurity large model corpus has shortcomings in its construction process, including low automation, heavy reliance on subjective human evaluation, and lack of objective verification of the practical effectiveness of attack and defense tactics.
[0024] This application provides a solution that filters text in the cybersecurity field, extracts key offensive and defensive techniques and tactics from the text, and transforms them into offensive and defensive exercise tasks supported by a target range environment. It completes corpus quantification, scoring, and screening, and incorporates corpora that meet the requirements for effectiveness verification in real-world target range scenarios into a large-scale cybersecurity vertical model corpus. Compared to existing large-scale cybersecurity model corpus construction processes, which suffer from low automation, heavy reliance on subjective human evaluation, and a lack of objective verification of the practical effectiveness of offensive and defensive techniques and tactics, this application automates the transformation of pre-cleaned and categorized cybersecurity technical and tactical terminology into target range offensive and defensive exercise tasks. Based on dynamic execution effects in an adversarial environment, it performs performance evaluation and screening, achieving automated and high-pass-rate screening of high-quality cybersecurity practical corpora. This reduces the illusion of large-scale models from the data source and effectively improves the upper limit of the practical capabilities of large-scale cybersecurity vertical models.
[0025] It should be noted that the execution subject of this embodiment can be a network security large model corpus construction device. The network security large model corpus construction device can be a device with an intelligent agent installed, such as a mobile terminal, personal computer, server or other electronic device used by the user. It can also be other devices that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment and the following embodiments, the network security large model corpus construction method is described using a network security large model corpus construction device.
[0026] The network security large model corpus construction device can also be an electronic device that can control and detect devices equipped with intelligent agents; this embodiment does not impose any restrictions on this.
[0027] Based on this, embodiments of this application provide a method for constructing a large-scale network security model corpus, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the network security large model corpus construction method of this application.
[0028] In this embodiment, the network security large model corpus construction method is applied to the sending end, including steps S10~S40: Step S10: Based on a preset large model, perform multi-dimensional scoring on text data in the field of cybersecurity, and filter target text from the text data according to the scoring results.
[0029] It should be noted that the expert large language model is called to score the selected cybersecurity text data from multiple preset dimensions, and the weight of each dimension is dynamically assigned according to the cybersecurity subdomain to which the text belongs, so as to calculate the comprehensive quality score of each text and select the target text that meets the quality requirements.
[0030] Furthermore, before step S10, the method further includes: performing text cleaning on the original multimodal security data to obtain cleaned data; removing duplicate paragraphs from the cleaned data using a paragraph-level deduplication mechanism based on the minimum hash algorithm to obtain a candidate corpus set; and classifying the text in the candidate corpus set by topic based on a preset large model to filter out text data belonging to the cybersecurity field.
[0031] It should be noted that text preprocessing and filtering are performed on the original multimodal security data to obtain text data in the field of cybersecurity. Data reduction is achieved through document deduplication. For the massive and redundant internet-collected data, a paragraph-level deduplication mechanism based on the MinHash algorithm is adopted to quickly calculate and compare the hash signatures of text fragments, effectively eliminating highly similar or repetitive redundant paragraphs, improving the purity of the basic corpus and the efficiency of subsequent processing. After obtaining the original multimodal security data and cleaning it, a paragraph-level deduplication mechanism based on the minimum hash algorithm is used to remove duplicate paragraphs, resulting in a candidate corpus set. An expert large language model with cybersecurity knowledge is then invoked, and few-shot prompt word engineering is used to classify the text in the candidate corpus set by topic, filtering out text data belonging to the cybersecurity field.
[0032] Understandably, when classifying topics based on expert models, considering that text crawled from wide area networks inevitably contains irrelevant data from non-cybersecurity domains, this application introduces a large-scale security expert model and combines it with Few-Shot (few-shot) prompt word engineering to quickly and accurately determine the domain of the deduplicated text, remove irrelevant noise, and ensure the topic purity of the corpus. The pre-set large-scale model can be a pre-trained and constructed large-scale security expert language model. The large-scale security expert language model refers to a large-scale language model with the ability to understand cybersecurity knowledge, analyze vulnerabilities, identify attack chains, analyze threat intelligence, and map attack and defense tactics. It can be implemented in two ways: the first is prompt word enhancement, which uses an existing large-scale cybersecurity domain model or a general large-scale model, and uses system prompt words, few-shot examples, and structured output constraints to enable it to complete topic classification, quality rating, and TTPs extraction tasks; the second is fine-tuning enhancement, which uses cybersecurity corpus to fine-tune the basic large-scale model, making the model more suitable for cybersecurity corpus discrimination and structure extraction tasks. If fine-tuning enhancement is adopted, the fine-tuning dataset may include vulnerability analysis reports, CVE announcements, PoC / Exp descriptions, threat intelligence reports, APT analysis reports, security operation alert explanations, range mission descriptions, and manually labeled cybersecurity / non-cybersecurity classification samples. Training can employ supervised fine-tuning, constructing instruction samples from the input text and expected JSON output, enabling the model to learn whether the output belongs to the cybersecurity domain, its subdomain, confidence level, quality score, and technical / tactical elements. In the preferred embodiment of this application, the expert large model is primarily implemented using a "cybersecurity vertical domain large model, Few-Shot cue word engineering, and JSON output constraints" approach, without requiring retraining or fine-tuning of the model. In other alternative embodiments, an expert large model fine-tuned with cybersecurity instruction data can also be used to further improve the consistency of topic classification, scoring, and extraction.
[0033] In one implementation, for the raw multimodal security data scraped from the massive amounts of the internet, basic text cleaning (such as removing HTML tags and filtering special characters) is first performed. Then, a paragraph-level deduplication mechanism based on the MinHash algorithm is used to reduce computational complexity and improve deduplication recall. The specific implementation method is as follows: Text segmentation and N-gram extraction: The cleaned original document is divided into several paragraphs, and an N-gram (e.g., 3-gram) feature set is extracted for each paragraph. Let the N-gram feature sets of paragraph A and paragraph B be S_A and S_B, respectively.
[0034] Hash signature generation: Introduce k independent hash functions h_1, h_2, ..., h_k. For the feature set S_A of paragraph A, calculate the minimum value under each hash function to generate a signature vector of length k: V_A=[min(h_1(S_A)),min(h_2(S_A)),...,min(h_k(S_A))] Similarity calculation and deduplication determination: According to the mathematical properties of MinHash, the probability that corresponding elements in two paragraph signature vectors are equal is equivalent to their Jaccard similarity J(A,B). This application estimates the similarity by comparing signature vectors:
[0035] Where k represents the total number of independent hash functions; The function is an indicator function that takes the value 1 when the condition in parentheses is true, and 0 otherwise. θ represents the system's preset deduplication similarity threshold (e.g., 0.85). When J(A,B) is greater than or equal to θ, it is determined to be a duplicate paragraph and is removed. The deduplicated candidate corpus is then output.
[0036] It should be noted that the MinHash algorithm in this application is only used for similarity estimation and duplicate content removal of candidate corpora, and does not directly undertake the task of cybersecurity topic classification. The system first uses MinHash to perform paragraph-level deduplication on the original webpage text, outputting a low-redundancy candidate corpus set; then, this candidate corpus set is input into a large-scale security expert model, which performs cybersecurity topic identification, subdomain classification, and subsequent quality scoring. Therefore, the "minimum hash processing" and "expert model classification" in this application are two independent processing steps that are interconnected: the former solves the problem of duplicate data filtering, and the latter solves the problem of corpus topic identification and professional judgment.
[0037] In the specific implementation, refer to Figure 2 The text deduplication diagram shown can be further broken down into the following execution process using MinHash: S01, perform basic cleaning of the original webpage text to remove webpage tags, navigation bars, advertising slogans, copyright notices, script code and garbled characters; S02, the cleaned document is divided into multiple text segments according to natural paragraphs, heading segments, code blocks or fixed-length windows; S03, extract character-level or word-level N-gram features from each text segment to form a segment feature set; S04, multiple independent hash functions are used to map each N-gram feature set, and the minimum hash value under each hash function is taken to form the MinHash signature vector of the segment; S05, compare the signature vectors of different segments to estimate the Jaccard similarity between segments; S06, when the similarity exceeds the preset threshold, the later-appearing segment is identified as a duplicate or near-duplicate segment and removed; when the similarity is below the threshold, the segment is retained as a candidate corpus. S07 outputs the deduplicated candidate corpus set and enters the expert large model topic classification stage.
[0038] Step S20: Perform tactical structured extraction on the target text to obtain structured target range combat parameters.
[0039] It should be noted that for high-quality corpora that pass the static quality rating, this application integrates them into a high-fidelity simulated network range (such as a network range) for effect verification. The system first uses automated information extraction technology to extract key offensive and defensive tactics (TTPs) from the target text, obtaining structured network range combat parameters.
[0040] Furthermore, step S30 also includes: extracting target range adversarial elements from the target text based on the information extraction expert big model and the preset prompt word template. The target range adversarial elements include the target target range environment, attack tactics number, payload type, executable code, and initial target range combat adversarial parameters of expected execution results; converting the initial target range combat adversarial parameters into a preset structured format to obtain structured target range combat adversarial parameters.
[0041] It should be noted that the system invokes an expert-level large model focused on Information Extraction (IE), utilizing structured prompt word templates to precisely extract executable adversarial elements from the candidate corpus selected in step three (quality rating). The prompt word templates are as follows: System Commands You are an advanced cybersecurity intelligence analysis engine. Extract key practical adversarial parameters from the input unstructured cybersecurity analysis text and map them into a standardized JSON structure.
[0042] Extraction Rules and Output Specifications Please extract the following fields. If any item is missing from the text, please enter null: target_env: Target environment components and versions (e.g., Apache 2.4.49, Redis 5.0).
[0043] attack_tactic_id: The mapped MITRE ATT&CK tactic / technical number (e.g., T1190).
[0044] payload_type: Payload type (e.g., HTTP Request, Bash Command, Python Script).
[0045] executable_code: Exploit code (PoC / Exp), malicious payload, or defense configuration instructions.
[0046] expected_outcome: Expected execution result characteristics (such as "returning root privileges" or "generating a specific file").
[0047] In one implementation, a text fragment is input: "Regarding the unauthorized access vulnerability in a certain OA system, an attacker can send a GET request containing host=127.0.0.1;id to the / api / v1 / system / ping interface. If the system has the vulnerability, the HTTP response will directly return the current server's running user group information (such as uid=0(root)), thereby achieving command injection." Model outputs JSON: { "target_env": "OA System / Web Server", "attack_tactic_id": "T1190", "payload_type": "HTTP Request", "executable_code": "GET / api / v1 / system / ping?host=127.0.0.1;id HTTP / 1.1", "expected_outcome": "uid=0(root)" } Step S30: Map the structured target range combat parameters to the network target range to execute attack and defense exercises, and calculate the corpus admission score corresponding to each target text executing the attack and defense exercises.
[0048] It should be noted that for high-quality corpora that pass the static quality rating, this application integrates them into a high-fidelity simulated network range (such as a network range) for effectiveness verification. The system first utilizes automated information extraction technology to extract key attack and defense tactics (TTPs), vulnerability exploitation code (PoC), or defense configurations from the text, and transforms them into attack and defense exercise tasks supported by the range environment. Subsequently, adversarial exercises are executed in the range environment. The system performs quantitative scoring and final selection based on the physical effects of the actual adversarial exercise (such as whether vulnerabilities are successfully triggered or whether defense scripts successfully block attacks). Only corpora that successfully pass the effectiveness verification in the network security vertical domain will be formally included in the large-scale model corpus of cybersecurity.
[0049] Step S40: Input the target texts whose corpus admission scores meet the preset admission score threshold into the training corpus.
[0050] Understandably, target texts whose corpus selection scores meet the preset selection score thresholds are entered into the training corpus.
[0051] This embodiment provides a method for constructing a large-scale cybersecurity model corpus. This embodiment performs multi-dimensional scoring on text data in the cybersecurity field based on a preset large-scale model, and filters target texts from the text data according to the scoring results. Tactical structure extraction is performed on the target texts to obtain structured target range combat parameters. These structured target range combat parameters are mapped to cyber range attack and defense exercises, and a corpus adoption score is calculated for each target text performing the attack and defense exercises. Target texts whose corpus adoption scores meet a preset adoption score threshold are entered into the training corpus. This application completes the corpus quantification, scoring, and filtering by filtering texts in the cybersecurity field, extracting key attack and defense techniques and tactics from the texts, and transforming them into attack and defense exercises supported by the target range environment. Corpus data that meets the effectiveness verification requirements in the target range is then included in the large-scale cybersecurity model corpus. Compared to existing cybersecurity large-scale model corpus construction processes, which suffer from low automation, heavy reliance on subjective human evaluation, and a lack of objective verification of the practical effectiveness of offensive and defensive tactical knowledge, this application automates the transformation of the initially cleaned and classified cybersecurity tactical terminology corpus into range offensive and defensive exercise tasks. Based on the dynamic execution effect in the adversarial environment, it conducts performance evaluation and screening, achieving automated and high-pass-rate screening of high-quality cybersecurity practical corpus. This reduces the illusion of large models from the data source and effectively improves the upper limit of the practical capabilities of cybersecurity vertical domain large-scale models.
[0052] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, the embodiments of this application provide a method for constructing a large-scale network security model corpus. In this embodiment, refer to... Figure 3 , Figure 3This is a flowchart illustrating the second embodiment of the network security large model corpus construction method of this application. Step S10 further includes: steps S101 to S103; Step S101: Perform domain analysis on text data in the cybersecurity field based on a pre-set large model to obtain a set of cybersecurity subdomains; It should be noted that the cybersecurity subdomains are determined by classifying topics using expert large-scale model ew-shot prompts and filtering topics using Few-Shot prompt engineering. This application obtains a set of cybersecurity subdomains by dividing text data in the cybersecurity field based on a pre-set large-scale model.
[0053] Understandably, since data crawled from wide area networks inevitably contains irrelevant information, this step introduces expert-level large language models with cybersecurity knowledge (such as Trend Micro's Primus and Cisco's Foundation-Sec model) to filter topics using a Few-Shot tooltip project. The tooltip template needs to include a task description, classification criteria, and 3-5 positive and negative sample (Few-Shot) tooltips. The template is as follows: System Prompt You are a domain expert model with top-notch security knowledge. Your task is to analyze the input web scraping text and determine whether it substantially falls under the domain of "cybersecurity, attack and defense, vulnerability analysis, or threat intelligence".
[0054] Output Specifications Please output strictly in JSON format, without including any unnecessary explanatory text. Output fields include: is_cybersec (Boolean): Whether the data pertains to cybersecurity.
[0055] confidence (floating-point number): The confidence level of the judgment, ranging from [0.0, 1.0].
[0056] Few-Shot Examples Example 1 (Positive Sample) Input text: "In the newly disclosed CVE-2024-XXXX, attackers can bypass WAF restrictions by using a carefully crafted serialized object and write a webshell into the / tmp directory of the target server, thereby achieving remote code execution (RCE)." Expected output: {"is_cybersec": true, "confidence": 0.98} Example 2 (Negative Sample) Enter text: "The technology sector saw a surge this week, with XX company releasing its latest generation smartphone, featuring the newest 5nm chip and significantly improved battery life, which has garnered widespread attention from consumers." Expected output: {"is_cybersec": false, "confidence": 0.05} The system parses the JSON result returned by the large model and directly extracts the value of the confidence field. If $\ge0.5$, it proceeds to step S102.
[0057] Step S102: Perform multi-dimensional weighted scoring on the text data based on the weight information corresponding to the set of network security sub-domains to obtain a comprehensive quality score result; Understandably, multi-dimensional weighted scoring refers to weighted scoring based on preset dimensions and corresponding weights for sub-domains. The preset dimensions include four dimensions: readability, logical fluency, structural integrity, and technical richness. Different weight vectors are assigned to these four dimensions in different domains. The weights for each dimension are dynamically allocated, including: when the text is determined to belong to the vulnerability discovery and attack / defense practice category, the highest weight is assigned to the technical richness dimension; when the text is determined to belong to the security defense and automation script category, a relatively high weight is assigned to the structural integrity dimension; when the text is determined to belong to the reverse engineering and low-level analysis category, a relatively high weight is assigned to the logical fluency and structural integrity dimensions; and when the text is determined to belong to the threat intelligence and situational awareness category, a relatively high weight is assigned to the structural integrity and logical fluency dimensions.
[0058] In one implementation, to further refine the corpus quality, the Few-Shot cue capabilities of the expert large model are utilized to perform fine-grained automated scoring of the corpus from multiple dimensions, including text readability, logical coherence, technical richness (such as whether it contains detailed code snippets or configuration instructions), and structural integrity, combined with domain-specific weights. A dynamic admission control mechanism is set up, and the corpus is only eligible to enter the next round of practical verification when the comprehensive score exceeds a preset threshold.
[0059] Understandably, to avoid low-quality, semantically fragmented cybersecurity text entering the time-consuming target range evaluation stage, this step proposes a multi-dimensional weighted quality assessment algorithm, continuing to utilize an expert large model for fine-grained scoring of the corpus. The large model is required to complete simultaneous evaluation of all four dimensions in a single inference iteration and output a structured score. The template is as follows: System Prompt You are a rigorous cybersecurity data quality assessment and dynamic scoring engine. Please perform in-depth analysis and quantitative scoring of the input cybersecurity text according to the following quality rating and dynamic scoring system, and finally output a comprehensive quality score Q(x).
[0060] [Quality Rating and Dynamic Scoring System] Please first independently score the text across its four core dimensions (each with a maximum score of 100 points): Readability (S_r): Focuses on evaluating text formatting and the absence of garbled characters; Logical fluency (S_f): This focuses on evaluating the coherence of the contextual semantics; Structural integrity (S_c): Focus on assessing whether the analysis elements such as vulnerability background and attack preconditions are complete; Technical richness (S_t): As a core evaluation indicator, it focuses on assessing whether the text contains substantial hard-core security knowledge (such as specific vulnerability exploit code PoC / Exp, malicious instruction payload, CVE number, registry modification path or defense rule script).
[0061] After completing the independent scoring of each dimension, please analyze the text content in depth, categorize it into the most relevant cybersecurity subdomain, and assign corresponding weight vectors. Make a comprehensive judgment based on the core technology objects, keywords, code / configuration content, IOC / TTPs elements, and attack and defense scenarios in the input text: If the focus is on vulnerability number, PoC / Exp, attack payload, vulnerability reproduction, etc., then it should be classified as Domain A; If the focus is on detection rules, defense scripts, hardening configurations, alarm handling, or automated responses, then it should be classified as Domain B. If the focus is on malicious code, assembly instructions, function calls, control flow, or low-level execution logic, then it should be classified as domain C. If the focus involves IOCs, TTPs, APT groups, attack chains, technical and tactical designations, and situational assessment, then it should be classified under Domain D.
[0062] If a text involves multiple subfields, the most relevant subfield is determined based on its core purpose and main technical content.
[0063] Area A (Vulnerability Discovery and Attack / Defense Practice): This area highly emphasizes the practical value of vulnerability verification code and attack payloads.
[0064] The weight vector [w_r, w_f, w_c, w_t] = [0.05, 0.10, 0.15, 0.70] Domain B (Security Defense and Automation Scripts): Emphasizes the accuracy of defense rule instructions and the completeness of the configuration environment. Weight vector [w_r, w_f, w_c, w_t] = [0.10, 0.15, 0.25, 0.50] Domain C (Reverse Engineering and Low-Level Analysis): Emphasizes the rigor of assembly instruction execution flow analysis and the coherence of contextual logic.
[0065] The weight vector [w_r, w_f, w_c, w_t] = [0.10, 0.20, 0.30, 0.40] Domain D (Threat Intelligence and Situation Awareness): Focuses on the analysis of APT attack chains, emphasizing the closed loop and completeness of intelligence elements (IOC / TTP mapping).
[0066] The weight vector [w_r, w_f, w_c, w_t] = [0.15, 0.25, 0.40, 0.20] Output Specifications Output must be a standard JSON object, and no other explanatory text is allowed. The JSON must explicitly contain the subdomain you determined, the array of weights applied, the raw scores for each dimension, and the final calculated composite score. An example output format is as follows: { "Sub_Category": "Vulnerability Discovery and Attack / Defense Practice", "Weights_Applied": [0.05, 0.10, 0.15, 0.70], "Raw_Scores": { "S_r": 90, "S_f": 85, "S_c": 80, "S_t": 95 }, "Final_Score_Q": 91.5 }
[0067] Step S103: Based on the comprehensive quality score result, determine whether the text data meets the quality requirements, and use the text data that meets the quality requirements as the target text.
[0068] It should be understood that this application abandons the traditional large-scale model data annotation method of fixed weights and constructs a dynamic weight adaptive allocation mechanism based on cybersecurity combat scenarios. This mechanism can identify the evaluation focus of different cybersecurity knowledge systems (such as code verification for attack and defense and logical closed loop for intelligence), thereby maximizing the preservation of the high-value characteristics of heterogeneous cybersecurity data and providing more accurate data input for subsequent automated range exercises.
[0069] To further illustrate, consider the following scoring example: Assume the input corpus is a vulnerability analysis blog post, including the affected scope of a certain web service version, the vulnerability triggering interface, exploitation conditions, HTTP request payload, return result characteristics, and remediation suggestions. The expert model first determines that the corpus belongs to "Domain A: Vulnerability Discovery and Attack / Defense Practice," therefore using a weight vector [w_r, w_f, w_c, w_t] = [0.05, 0.10, 0.15, 0.70]. Here, S_r represents readability, S_f represents logical fluency, S_c represents structural integrity, and S_t represents technical richness.
[0070] Based on expert evaluation using a large-scale model, the scores for each dimension of this corpus are as follows: S_r = 88 indicates that the text layout is relatively clear, with only a small amount of formatting noise; S_f = 84 indicates that the logic between the vulnerability background, triggering method, and impact is relatively coherent; S_c = 80 indicates that the vulnerability background, target environment, triggering conditions, and remediation suggestions are included, but a more complete description of environment dependencies is lacking. S_t = 92 indicates that it includes specific interfaces, parameters, request payloads, and expected response characteristics, and has a high level of practical technical content.
[0071] The calculation process for the overall quality score Q(x) is as follows: Q(x)=0.05×88+0.10×84+0.15×80+0.70×92=4.4+8.4+12.0+64.4=89.2.
[0072] If the system presets a static quality threshold of 80 points, the corpus will pass the quality rating and enter the subsequent TTPs extraction and target range verification stages; if Q(x) < 80, the corpus will be judged as low-quality corpus and discarded or enter the manual review queue.
[0073] The corresponding structured scoring output is as follows: { "Sub_Category": "Vulnerability Discovery and Attack / Defense Practice", "Weights_Applied": [0.05, 0.10, 0.15, 0.70], "Raw_Scores": {"S_r": 88, "S_f": 84, "S_c": 80, "S_t": 92}, Calculation: 0.05 88 + 0.10 84 + 0.15 80 + 0.70 92", "Final_Score_Q": 89.2, "Decision":"Pass" } Scoring Consistency Test: To verify the stability of the expert large-scale model scoring results, this application may set up a scoring consistency test mechanism. Specifically, multiple rounds of repeated scoring are performed on the same batch of candidate corpora, or two independent expert large-scale models are used for cross-scoring, and the mean absolute difference, variance, or rank consistency rate between the multiple rounds of scoring for the same corpus is calculated. For example, in one embodiment, 100 candidate cybersecurity corpora are selected, and three rounds of independent quality scoring are performed on each. If the difference in the comprehensive score of the three rounds does not exceed 5 points, the corpus is considered to have a consistent score; if the difference exceeds 5 points, it enters the stage of manual review or secondary model evaluation. The test results show that 91 out of 100 corpora have a three-round score difference of no more than 5 points, with a scoring consistency rate of 91%; the mean absolute difference in the comprehensive score of the three rounds is 2.8 points, indicating that this scoring mechanism has good stability and repeatability.
[0074] This embodiment provides a method for constructing a large-scale cybersecurity model corpus. This embodiment performs multi-dimensional scoring on text data in the cybersecurity field based on a preset large-scale model, and filters target texts from the text data according to the scoring results. Tactical structure extraction is performed on the target texts to obtain structured target range combat parameters. These structured target range combat parameters are mapped to cyber range attack and defense exercises, and a corpus adoption score is calculated for each target text performing the attack and defense exercises. Target texts whose corpus adoption scores meet a preset adoption score threshold are entered into the training corpus. This application completes the corpus quantification, scoring, and filtering by filtering texts in the cybersecurity field, extracting key attack and defense techniques and tactics from the texts, and transforming them into attack and defense exercises supported by the target range environment. Corpus data that meets the effectiveness verification requirements in the target range is then included in the large-scale cybersecurity model corpus. Compared to existing cybersecurity large-scale model corpus construction processes, which suffer from low automation, heavy reliance on subjective human evaluation, and a lack of objective verification of the practical effectiveness of offensive and defensive tactical knowledge, this application automates the transformation of the initially cleaned and classified cybersecurity tactical terminology corpus into range offensive and defensive exercise tasks. Based on the dynamic execution effect in the adversarial environment, it conducts performance evaluation and screening, achieving automated and high-pass-rate screening of high-quality cybersecurity practical corpus. This reduces the illusion of large models from the data source and effectively improves the upper limit of the practical capabilities of cybersecurity vertical domain large-scale models.
[0075] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, the embodiments of this application provide a method for constructing a large network security model corpus, referring to... Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the network security large model corpus construction method of this application. In this embodiment, step S30 further includes S301~S302: Step S301: Map the structured target range combat parameters to a network target range deployable triplet artifact.
[0076] It should be noted that the triplet artifact includes an environment template, an attack execution script, and a state detection probe. To achieve the automatic conversion of unstructured corpus into executable adversarial tasks, this step designs an orchestration function Φ to map the TTPs structured object J extracted in step 1 to a target-deployable triplet artifact Φ(J)=(E,A,P).
[0077] Where E is the environment template, A is the attack execution script, and P is the state detection probe. The environment template E is generated using the `target_env` field as the index key, matching the corresponding version of the container image or virtual machine snapshot from a pre-configured vulnerability environment image library; if the library is missing, the template is generated on the spot based on the environment. The attack execution script A is generated based on a parameterized template engine. The system maintains a script template distribution table based on the `payload_type` field, and pre-configures script skeleton templates for different payload types such as HTTP requests, Shell commands, and Python scripts. Placeholders mark variable fields (such as target address, exploit code, expected result characteristics, etc.) in the templates. At runtime, the system completes script rendering according to the following process: A=Render(T_payload_type,J∪C_env) Render is a standard term in template processing, referring to the process of combining a template skeleton with variable data to generate a final executable text output. T_payload_type represents the template categorized by payload type, and C_env represents the dynamically allocated environment context (such as IP address, port, and credentials) in the test range. The rendering process injects the field values of J and C_env into the template placeholders and automatically adds common logic such as timeout control, exception handling, and exit code normalization, outputting a script file that can be directly executed in the test range. The state detection probe P automatically generates a set of monitoring rules based on the expected_outcome field, including regular expression pattern matching, file hash comparison, process fingerprint comparison, and port monitoring changes. Finally, the system packages the triple (E, A, P) into a standardized Playbook and submits it to the test range execution engine for scheduling.
[0078] Understandably, the advantages of using a parameterized template engine instead of hard-coded string concatenation are: templates are decoupled from data, adding new payload types only requires expanding the template library without modifying the core scheduling logic; the template rendering process automatically escapes variables, which can effectively prevent special characters from breaking the script syntax; and it supports template inheritance and fragment reuse, which significantly reduces maintenance costs in large-scale corpus processing scenarios.
[0079] Step S302: Execute an attack and defense exercise based on the triplet artifact, and calculate the corpus admission score corresponding to each target text executing the attack and defense exercise.
[0080] It should be noted that the environment template is deployed through the triplet artifact, and attack and defense exercises are performed to calculate the corpus acceptance score.
[0081] Furthermore, step S302 further includes: deploying the environment template in the network range and performing an attack and defense exercise based on the attack execution script; collecting the target machine state vector before the attack and the target machine state vector after the attack based on the state detection probe; determining the state transition vector based on the target machine state vector before the attack and the target machine state vector after the attack; and determining the corpus acceptance score based on the execution result corresponding to the attack execution script, the state transition vector and the state consistency between the expected execution result.
[0082] It should be noted that after deploying environment template E, probe P collects the initial state vector S0 of the target machine before the attack is launched, launches attack script A, and waits for a limited time T before collecting the final state S1 and calculating the state transition ΔS. The target machine state vector is defined as a quadruple: S=(S_proc,S_file,S_net,S_log) S_proc is the process view (adding / exiting processes and PID set), S_file is the file system view (changing file paths and hashes), S_net is the network view (port listening and connection 5-tuple), and S_log is the security log set.
[0083] State difference: ΔS=S1 S0 in The difference operator is computed in dimension and outputs the set of adversarial features actually observed, F_observed.
[0084] Furthermore, determining the corpus acceptance score for each target text performing the attack and defense exercise task based on the execution result corresponding to the attack execution script, the state transition vector, and the state consistency between the expected execution result and the state transition vector includes: determining whether the execution result corresponding to the attack execution script was successfully executed within a limited time and obtaining an execution success indicator function; determining the state consistency based on the similarity between the actual adversarial feature set corresponding to the state transition vector and the expected feature set corresponding to the expected execution result; and performing a weighted summation of the execution success indicator function and the state consistency to determine the corpus acceptance score for each target text performing the attack and defense exercise task.
[0085] It should be noted that this step uses the actual execution results of the target range offensive and defensive confrontation as the sole criterion for corpus selection, quantifies "real combat effectiveness" into a calculable indicator, and automatically removes low-value corpus that is invalid, inaccurate in description, or unreproducible.
[0086] Define a success indicator function E(x):
[0087] Define the state fit M(x) as the Jaccard similarity between the observed feature set F_observed and the expected feature set F_expected:
[0088] The final score V(x) for corpus acceptance is weighted by two indicators: execution success E(x) and state matching M(x).
[0089] Where β+γ=1, it is recommended to take β=0.4 and γ=0.6 to highlight the weights of state feature matching. When V(x)≥τ_V (e.g. 0.80), it is included in the final training corpus; otherwise, it is discarded to avoid unverifiable "paper knowledge" from entering the large model training data.
[0090] In one implementation, taking the OA system command injection corpus as an example, after execution in the target range: The attack script ran successfully with an exit code of 0, and E(x) = 1. The probe captures a response containing the field "uid=0(root)", which perfectly matches the expected feature, M(x)=1.0; V(x) = 0.4 × 1 + 0.6 × 1.0 = 1.0 The corpus has been validated in real-world testing at the target range and has been officially added to the training database. Conversely, if the target machine response has no UID feature, then E(x) = 0 or M(x) = 0, V(x) ≤ 0.8, and it will be automatically discarded if it does not meet the threshold.
[0091] In one embodiment, for further illustration, the application references Figure 5 The overall architecture diagram shown below, using the OA system command injection corpus as an example, illustrates the execution in the target environment: The attack script ran successfully with an exit code of 0, and E(x) = 1. The probe captures a response containing the field "uid=0(root)", which perfectly matches the expected feature, M(x)=1.0; V(x) = 0.4 × 1 + 0.6 × 1.0 = 1.0 The corpus has been validated in real-world testing at the target range and has been officially added to the training database. Conversely, if the target machine response has no UID feature, then E(x) = 0 or M(x) = 0, V(x) ≤ 0.8, and it will be automatically discarded if it does not meet the threshold.
[0092] Complete processing path example: Using a publicly available vulnerability analysis blog as the raw corpus input. The blog describes a command injection vulnerability in a web application component, where an attacker can trigger system command execution by constructing specific HTTP request parameters and return system user information in the response content.
[0093] Raw text input Example input text: "The / api / v1 / system / ping interface of a certain web application has a command injection risk. When an attacker appends a system command to the host parameter, for example, host=127.0.0.1;id, the server will directly pass the parameter to the system ping command for execution. If the target is vulnerable, the HTTP response will return system user information such as uid and gid. It is recommended to upgrade to the patched version and perform whitelist verification on the host parameter." The output is a raw webpage text object, including metadata such as title, body, publication time, source URL, body length, and crawl time.
[0094] Deduplication The input is the original webpage text object. The system cleans and segments the text, extracts N-gram features and generates a MinHash signature, which is then compared with signatures in an existing corpus for similarity.
[0095] Output example: {"duplicate":false,"similarity_max":0.42,"decision":"Keep","reason":"Maximum similarity with existing corpus is below the deduplication threshold of 0.85"}.
[0096] Theme Classification The input is the deduplicated candidate text. The expert model determines whether it belongs to the fields of cybersecurity, vulnerability analysis, or attack and defense based on the Few-Shot prompts.
[0097] Output example: {"is_cybersec":true,"category":"Vulnerability Analysis and Attack / Defense Practice","confidence":0.96,"decision":"Entering Quality Rating"}.
[0098] Quality rating The input consists of candidate texts identified as belonging to the cybersecurity field. An expert model scores the texts based on four dimensions: readability, logical fluency, structural completeness, and technical richness, and calculates a comprehensive score according to the weights of vulnerability discovery and attack / defense practice.
[0099] Output example: {"Sub_Category":"Vulnerability Discovery and Attack / Defense Practice","Weights_Applied":[0.05,0.10,0.15,0.70],"Raw_Scores":{"S_r":88,"S_f":84,"S_c":80,"S_t":92},"Final_Score_Q":89.2,"decision":"Passed static quality rating, proceed to TTPs extraction"}.
[0100] TTPs extraction The input is a vulnerability analysis text that has passed the quality rating. The information extraction expert's large model extracts the target environment, attack technique ID, payload type, executable request, and expected result.
[0101] Output example: {"target_env":"Web Application / vulnerable ping API", "attack_tactic_id": "T1190", "payload_type": "HTTP Request", "executable_code": "GET / api / v1 / system / ping?host=127.0.0.1;id HTTP / 1.1", "expected_outcome": "The HTTP response contains a uid or gid field", "defense_suggestion": "Upgrade the version and whitelist the host parameter"}.
[0102] Target range conversion The input is a TTPs structured object. The system automatically matches the target environment image, network topology, and task template based on target_env, and encapsulates the HTTP request payload into a target environment executable task.
[0103] Output example: {"task_name":"Web API Command Injection Validation", "target_image": "vulnerable-web-api:demo", "attacker_node": "kali-client", "victim_node": "web-server", "network_topology": "attacker-victim single-hop topology", "execution_command": "curl 'http: / / victim / api / v1 / system / ping?host=127.0.0.1;id'", "success_condition": "Response content contains uid or gid"}.
[0104] Target range execution The input is the target range task object. The target range system automatically starts attack nodes and victim nodes, executes tasks in an isolated environment, and collects HTTP responses, system logs, network traffic, and exit status.
[0105] Output example: {"execution_status":"completed","http_status":200, "response_snippet": "uid=1000(www-data) gid=1000(www-data)", "log_evidence": "web-server access log contains injected parameter", "traffic_evidence": "attacker sent crafted HTTP request to victim"}.
[0106] Target Range Rating The inputs are the execution result and the expected result. The system scores based on dimensions such as environment matching, task executability, attack effectiveness hit rate, evidence completeness, and security isolation. The target range performance score R(x) can be calculated as follows: R(x)=0.20×E_env+0.25×E_exec+0.30×E_effect+0.15×E_evidence+0.10×E_safety Where E_env represents the environment matching degree, E_exec represents the task executability, E_effect represents the attack or defense effect hit rate, E_evidence represents the integrity of log, traffic, and response evidence, and E_safety represents the target range isolation and security execution status. Taking the above execution results as an example... R(x)=0.20×90+0.25×100+0.30×95+0.15×90+0.10×100=18+25+28.5+13.5+10=95.
[0107] If the target range scoring threshold is set to 80 points, the corpus will pass the practical effectiveness verification.
[0108] Inbound / Discard The inputs are a static quality score Q(x) = 89.2 and a range performance score R(x) = 95. The system makes the final determination based on a dual-threshold admission strategy.
[0109] Output example: {"quality_score":89.2,"range_score":95,"final_decision":"Input","corpus_type":"Exploitation and Attack / Defense Practice Corpus","stored_fields":["Original Text","Cleaned Text","Topic Classification Result","Quality Score Result","TTPs Structured Result","Target Range Task Object","Execution Evidence","Final Score"]}.
[0110] If the corpus fails to meet the threshold requirement at any stage, the system records the reason for discarding it.
[0111] For example, if the topic classification confidence score is below 0.5, it is marked as "non-cybersecurity corpus"; if the quality score is below 80, it is marked as "low-quality corpus"; if the target range execution fails or fails to achieve the expected results, it is marked as "invalid combat corpus". Discarded corpus is not included in the training corpus, but can be included in the abnormal sample pool for subsequent manual review or prompt word optimization.
[0112] The final high-quality cybersecurity large-scale model corpus includes not only the raw text but also the structured metadata generated during processing. Its data structure includes: 1) Raw corpus fields: title, source, URL, publication time, crawling time, and text content; 2) Deduplication fields: MinHash signature, maximum similarity, and duplicate detection result; 3) Topic classification fields: whether it is a cybersecurity corpus, cybersecurity sub-domain, classification confidence; 4) Quality scoring fields: scores in four dimensions, weight vector, overall quality score, and admission result; 5) TTPs fields: Target environment, attack technique ID, payload type, executable code, and expected result; 6) Range mission fields: mirror environment, network topology, attacking node, victim node, executed command, success criteria; 7) Execution evidence fields: return results, log evidence, traffic evidence, screenshots or file evidence; 8) Final judgment fields: target range score, entry / discard result, reason for discard or entry category.
[0113] Through the aforementioned structured fields, this application can ensure that each training corpus has a traceable source, an interpretable scoring process, and verifiable evidence of practical execution.
[0114] This application presents a complete closed loop from raw cybersecurity text to high-quality large-scale model training corpus: First, it utilizes MinHash for low-cost deduplication, avoiding redundant corpus pollution; second, it employs expert large-scale models to perform cybersecurity topic classification and multi-dimensional quality rating, preventing non-cybersecurity text and low-quality text from entering subsequent stages; third, it uses structured extraction to convert vulnerability environments, attack payloads, TTPs numbers, and expected results from the text into target range tasks; finally, it objectively verifies the practical effectiveness of the corpus through network target range execution results. Compared to traditional corpus construction methods that rely solely on static semantic similarity or manual scoring, this application not only improves corpus selection efficiency but also provides complete scoring criteria, execution evidence, and traceable processing paths for each piece of corpus, significantly enhancing the quality of large-scale model corpus and supporting high-quality training of large-scale models in vertical domains.
[0115] This embodiment provides a method for constructing a large-scale cybersecurity model corpus. This embodiment performs multi-dimensional scoring on text data in the cybersecurity field based on a preset large-scale model, and filters target texts from the text data according to the scoring results. Tactical structure extraction is performed on the target texts to obtain structured target range combat parameters. These structured target range combat parameters are mapped to cyber range attack and defense exercises, and a corpus adoption score is calculated for each target text performing the attack and defense exercises. Target texts whose corpus adoption scores meet a preset adoption score threshold are entered into the training corpus. This application completes the corpus quantification, scoring, and filtering by filtering texts in the cybersecurity field, extracting key attack and defense techniques and tactics from the texts, and transforming them into attack and defense exercises supported by the target range environment. Corpus data that meets the effectiveness verification requirements in the target range is then included in the large-scale cybersecurity model corpus. Compared to existing cybersecurity large-scale model corpus construction processes, which suffer from low automation, heavy reliance on subjective human evaluation, and a lack of objective verification of the practical effectiveness of offensive and defensive tactical knowledge, this application automates the transformation of the initially cleaned and classified cybersecurity tactical terminology corpus into range offensive and defensive exercise tasks. Based on the dynamic execution effect in the adversarial environment, it conducts performance evaluation and screening, achieving automated and high-pass-rate screening of high-quality cybersecurity practical corpus. This reduces the illusion of large models from the data source and effectively improves the upper limit of the practical capabilities of cybersecurity vertical domain large-scale models.
[0116] This application also provides a device for constructing a large-scale network security model corpus; please refer to... Figure 6 , Figure 6 A schematic diagram of the module structure provided in Embodiment 1 of the network security large model corpus construction device, wherein the network security large model corpus construction device includes: The text filtering module is used to perform multi-dimensional scoring on text data in the field of cybersecurity based on a preset large model, and to filter target text from the text data according to the scoring results. The tactical extraction module is used to perform tactical structured extraction on the target text to obtain structured target range combat parameters. The task training module is used to map the structured target range combat parameters to the network target range to execute attack and defense training tasks, and to calculate the corpus admission score corresponding to each target text executing the attack and defense training task. The corpus entry module is used to enter target texts whose corpus entry scores meet the preset entry score thresholds into the training corpus.
[0117] The network security large-scale model corpus construction device provided in this application, employing the network security large-scale model corpus construction method in the above embodiments, can solve the technical problems of low automation, heavy reliance on subjective human evaluation, and lack of objective verification of the practical effectiveness of attack and defense tactics in the existing network security large-scale model corpus construction process. Compared with the prior art, the beneficial effects of the network security large-scale model corpus construction device provided in this application are the same as those of the network security large-scale model corpus construction method provided in the above embodiments, and other technical features in the network security large-scale model corpus construction device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0118] This application provides a cybersecurity large model corpus construction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the cybersecurity large model corpus construction method in the above embodiment 1.
[0119] The following is for reference. Figure 7 The diagram illustrates a structural schematic of a network security large model corpus construction device suitable for implementing embodiments of this application. The network security large model corpus construction device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The cybersecurity large model corpus construction device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0120] like Figure 7As shown, the network security large model corpus construction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the network security large model corpus construction device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the network security large model corpus construction device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows network security large model corpus construction devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0121] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0122] The network security large-scale model corpus construction device provided in this application, employing the network security large-scale model corpus construction method described in the above embodiments, can solve the technical problems of low automation, heavy reliance on subjective human evaluation, and lack of objective verification of the practical effectiveness of attack and defense tactics in existing network security large-scale model corpus construction processes. Compared with the prior art, the beneficial effects of the network security large-scale model corpus construction device provided in this application are the same as those of the network security large-scale model corpus construction method provided in the above embodiments, and other technical features in this network security large-scale model corpus construction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0123] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0124] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0125] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the network security large model corpus construction method in the above embodiments.
[0126] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0127] The aforementioned computer-readable storage medium may be included in the network security large model corpus construction device; or it may exist independently and not be assembled into the network security large model corpus construction device.
[0128] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the network security large-scale model corpus construction device, the network security large-scale model corpus construction device performs multi-dimensional scoring on text data in the network security field based on a preset large-scale model, and selects target texts from the text data according to the scoring results. It then performs tactical structured extraction on the target texts to obtain structured target range combat parameters. Finally, it maps the structured target range combat parameters to the network target range to perform attack and defense exercises, calculates the corpus adoption score corresponding to each target text performing the attack and defense exercises, and inputs the target texts whose corpus adoption scores meet the preset adoption score threshold into the training corpus.
[0129] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Python, Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0131] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0132] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described method for constructing a large-scale cybersecurity model corpus. This addresses the technical problems of low automation, heavy reliance on subjective human evaluation, and lack of objective verification of the practical effectiveness of attack and defense tactics in existing large-scale cybersecurity model corpus construction processes. Compared with existing technologies, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the large-scale cybersecurity model corpus construction method provided in the above embodiments, and will not be elaborated upon here.
[0133] All user-related data involved in this application (such as user privacy data, user behavior data, etc.) were obtained with the user's permission or consent; that is to say, when this application is used in a specific product or technology, user permission is required to obtain and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.
[0134] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
Claims
1. A method for constructing a large-scale network security model corpus, characterized in that, The method for constructing the large-scale cybersecurity model corpus includes: Based on a pre-set large model, text data in the field of cybersecurity is scored from multiple dimensions, and target text is selected from the text data according to the scoring results; Tactical structured extraction is performed on the target text to obtain structured target range combat parameters; The structured target range combat parameters are mapped to the network target range to perform attack and defense exercises, and the corpus admission score corresponding to each target text performing the attack and defense exercises is calculated. The target texts whose corpus selection scores meet the preset selection score threshold are entered into the training corpus.
2. The method as described in claim 1, characterized in that, Before performing multi-dimensional scoring on text data in the cybersecurity field based on a preset large model, and filtering target text from the text data according to the scoring results, the process also includes: Text cleaning is performed on the original multimodal security data to obtain cleaned data; A paragraph-level deduplication mechanism based on the minimum hash algorithm is used to remove duplicate paragraphs from the cleaned data to obtain a candidate corpus set. Based on a pre-defined large model, the texts in the candidate corpus are classified by topic, and text data belonging to the field of cybersecurity are selected.
3. The method as described in claim 1, characterized in that, The process of performing multi-dimensional scoring on text data in the cybersecurity field based on a pre-defined large model, and then filtering target text from the text data according to the scoring results, includes: Based on a pre-defined large model, domain analysis is performed on text data in the field of cybersecurity to obtain a set of cybersecurity subdomains. Based on the weight information corresponding to the set of network security subdomains, the text data is scored using a multi-dimensional weighted score to obtain a comprehensive quality score result. Based on the comprehensive quality score, it is determined whether the text data meets the quality requirements, and the text data that meets the quality requirements is taken as the target text.
4. The method as described in claim 3, characterized in that, The step of performing tactical structured extraction on the target text to obtain structured target range combat parameters includes: Based on the information extraction model and preset prompt word template, target range adversarial elements are extracted from the target text. The target range adversarial elements include the target target range environment, attack tactics number, payload type, executable code, and initial target range combat adversarial parameters of expected execution results. The initial target range combat parameters are converted into a preset structured format to obtain the structured target range combat parameters.
5. The method according to any one of claims 1-4, characterized in that, The process of mapping the structured target range combat parameters to the network target range to execute attack and defense exercises, and calculating the corpus acceptance score for each target text executing the attack and defense exercises, includes: The structured target range combat parameters are mapped to ternary artifacts that can be deployed in the network target range. The attack and defense exercise is performed based on the triplet artifact, and the corpus admission score corresponding to each target text performing the attack and defense exercise is calculated.
6. The method as described in claim 5, characterized in that, The triplet artifact includes an environment template, an attack execution script, and a state detection probe. The execution of the attack and defense exercise task based on the triplet artifact, and the calculation of the corpus acceptance score corresponding to each target text executing the attack and defense exercise task, include: The environment template is deployed in the network test range, and attack and defense drills are performed based on the attack execution script. The target machine state vector before the attack and the target machine state vector after the attack are collected based on the state detection probe. The state transition vector is determined based on the target machine state vector before the attack and the target machine state vector after the attack. Based on the execution result of the attack execution script and the state consistency between the state transition vector and the expected execution result, the corpus selection score for each target text executing the attack and defense exercise task is determined.
7. The method as described in claim 6, characterized in that, The determination of the corpus acceptance score for each target text performing the attack and defense exercise task based on the execution result corresponding to the attack execution script, the state transition vector and the expected execution result, includes: Determine whether the execution result corresponding to the attack execution script was successfully executed within the limited time, and obtain the execution success indication function; The state fit is determined based on the similarity between the actual adversarial feature set corresponding to the state transition vector and the expected feature set corresponding to the expected execution result. The weighted summation of the execution success indication function and the state consistency is used to determine the corpus selection score corresponding to each target text executing the attack and defense exercise task.
8. A device for constructing a large-scale network security model corpus, characterized in that, The network security large model corpus construction device includes: The text filtering module is used to perform multi-dimensional scoring on text data in the field of cybersecurity based on a preset large model, and to filter target text from the text data according to the scoring results. The tactical extraction module is used to perform tactical structured extraction on the target text to obtain structured target range combat parameters. The task training module is used to map the structured target range combat parameters to the network target range to execute attack and defense training tasks, and to calculate the corpus admission score corresponding to each target text executing the attack and defense training task. The corpus entry module is used to enter target texts whose corpus entry scores meet the preset entry score thresholds into the training corpus.
9. A device for constructing a large-scale network security model corpus, characterized in that, The cybersecurity large model corpus construction device includes: a memory, a processor, and a cybersecurity large model corpus construction program stored in the memory and capable of running on the processor, wherein the cybersecurity large model corpus construction program is configured to implement the cybersecurity large model corpus construction method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a network security large model corpus construction program, which, when executed by a processor, implements the network security large model corpus construction method as described in any one of claims 1 to 7.