Network threat knowledge automatic extraction method, electronic equipment and storage medium
By using large models and semantic alignment technology, the problem of insufficient semantic alignment and contextual linkage in network threat intelligence knowledge extraction is solved, achieving high-precision APT attack chain reconstruction and graph construction, and improving the automated analysis capability of network threat intelligence.
Patent Information
- Application Number
- CN202511279087.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-11
AI Technical Summary
Existing network threat intelligence knowledge extraction technologies suffer from weak semantic alignment capabilities, insufficient linkage between context and domain knowledge, difficulty in effectively merging the same attack entity with multiple aliases and abbreviations, and difficulty in fully reconstructing APT attack chains using traditional methods.
We employ an automatic network threat knowledge extraction method based on large models and semantic alignment. By combining pre-trained sentence vector models and large language models, we perform multi-example prompt-driven semantic embedding and retrieval to construct a knowledge graph of APT organization network threat intelligence, enabling the structured extraction and fusion of multiple types of threat intelligence entities and semantic relationships.
It achieves high-precision automatic extraction of APT organizations, vulnerabilities, tools, and complex attack chain relationships from multi-source heterogeneous CTI data, improving the integrity of the graph structure and the closure of the relationship chain, enhancing cross-document reasoning capabilities, and reducing computational resource consumption.
Smart Images

Figure CN120930756A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of security layer technology for the Industrial Internet, particularly to the fields of industrial Internet threats, data security, and threat prediction in computer security. Specifically, it relates to a method for automatically extracting network threat knowledge based on large models and semantic alignment, as well as electronic devices and storage media. Background Technology
[0002] As cybersecurity threats continue to evolve, cyberattacks are becoming increasingly diverse, with a surge in APT analysis reports, vulnerability announcements, and threat intelligence, resulting in a massive amount of Cyber Threat Intelligence (CTI) data. This intelligence comes from diverse sources, is presented in varying ways, and the same attacking organization or tool often appears under aliases, abbreviations, and other forms. Furthermore, attack event descriptions exhibit significant contextual dependencies and cross-document correlations, posing a significant challenge to the automatic extraction, summarization, and knowledge association analysis of this intelligence.
[0003] Existing technologies for constructing and extracting knowledge of network threat intelligence have the following shortcomings: the same entity lacks effective mapping across different sources, resulting in fragmented information; rule-based or keyword-based methods can only identify Indicators of Computation (IOCs) and are unable to extract higher-level attack chains and organizational relationships; traditional sequence models have limited understanding of context and are unable to capture cross-segment dependencies; graph neural networks are mostly based on co-occurrence or shallow connections and cannot effectively integrate knowledge bases in fields such as ATT&CK and CVE for deep reasoning, thus limiting the completeness of APT attack profiles.
[0004] Current mainstream technologies for network threat intelligence knowledge extraction and graph construction can be mainly divided into the following three categories:
[0005] (1) Rule-based and keyword matching methods can quickly locate explicit features such as IP, domain name, hash, and CVE number through IOC patterns or regular expressions. For example, YARA rules can achieve IOC detection, but it is difficult to extract high-level semantic relationships such as APT organizations and attack chain context;
[0006] (2) Sequence labeling and machine learning methods, using models such as CRF and BiLSTM to perform sequence modeling on intelligence texts to achieve multi-type entity recognition. However, this type of method relies on a large number of manually labeled samples, has limited generalization ability, and cannot capture long-distance dependencies across paragraphs and reports;
[0007] (3) Application of graph neural networks in security intelligence: IOCs, vulnerabilities, organization names, etc. are modeled as graph nodes, and node relationships are learned using methods such as GCN and GAT. For example, graph models based on co-occurrence frequency can capture shallow connections, but lack deep linkage with ATT&CK and CVE databases, making it difficult to fully reconstruct the APT attack chain.
[0008] While the above methods have promoted the structuring of cyber threat intelligence to some extent, they still have the following fundamental shortcomings: First, the semantic alignment capability is weak. Existing methods mostly remain at the string or regular expression level for comparison, and cannot effectively merge the same attack entity with multiple aliases or abbreviations, resulting in fragmented graphs. Second, the linkage between context and domain knowledge is insufficient. Sequence models and graph neural networks often have difficulty combining with security knowledge bases such as ATT&CK for cross-document reasoning, which limits the in-depth analysis of APT organization behavior patterns. Summary of the Invention
[0009] The purpose of this invention is to provide a method, electronic device, and storage medium for automatically extracting network threat knowledge based on large models and semantic alignment, so as to overcome the shortcomings of the prior art.
[0010] Based on the first main aspect of the present invention, an automatic extraction method for network threat knowledge based on large models and semantic alignment includes the following steps performed by a computer hardware system:
[0011] Threat intelligence data related to APT organizations is collected from multi-source cybersecurity texts. The collected data is then segmented into sentences, denoised, and redundant text is removed to generate a standardized threat intelligence corpus.
[0012] A pre-trained sentence vector model is used to generate semantic embeddings for the input text and example library text from the standardized threat intelligence corpus. By calculating similarity, K examples most similar to the input are retrieved from the existing manually annotated example library. An ICL prompt template containing task instructions, examples and the current input is constructed.
[0013] The ICL prompt template is input into a large language model that has been lightly tuned by LoRA to extract structured triples of multiple types of threat intelligence entities and multiple semantic relationships;
[0014] For the output threat intelligence entities and semantic relationships, a semantic aggregation method combining prompt-driven large language model and vectorization method is used to generate standardized entity nodes and updated relationship information.
[0015] Based on the obtained standardized entity nodes and updated relationship information, an APT organization network threat intelligence knowledge graph containing specific nodes and edges is constructed, and a structured graph file is output.
[0016] The above five steps respectively achieve data acquisition and preprocessing, multi-example prompt-driven retrieval, large language model extraction, coarse and fine-grained semantic fusion, and knowledge graph construction.
[0017] During the data collection and preprocessing phase, textual data on network threat intelligence was collected from various data sources, including APT analysis reports, vulnerability announcements, threat intelligence bulletins, and social media posts. This included APT attack tracing reports, CVE vulnerability analyses, intelligence aggregation blogs, and posts from security communities such as Twitter and Telegram. The collected raw text data underwent sentence-level segmentation, and regular expression matching was used to remove redundant special characters, footnotes, and advertising content. Text encoding was standardized, and HTML tags and YAML / JSON nesting redundancy were cleaned to ensure a consistent structure for the multi-source heterogeneous data during the input stage, providing stable input for subsequent semantic feature extraction, similarity retrieval, and knowledge reasoning.
[0018] In the multi-example prompt-driven retrieval stage, a pre-trained sentence vector model (such as SecureBERT, Sentence-BERT, or TFIDF+PCA) is used to perform high-dimensional vectorization encoding on the input text and the pre-built example library. By calculating the cosine similarity between the input and example library vectors, the K examples with the closest semantic structure are retrieved, and an ICL (In-Context Learning) prompt template containing task instructions, the example set, and the current input text is dynamically constructed. The prompt template can dynamically adjust the example and instruction content according to the detected context type (APT organization affiliation, exploit vulnerability, tool targeting, etc.), achieving adaptive extraction for multiple tasks and significantly improving the coverage of multi-granularity entities and complex semantic relationships.
[0019] In the large language model extraction stage, the constructed ICL prompt template is input into the large language model, which has been lightweighted and fine-tuned using LoRA. Through contextual guidance and multi-example few-shot learning, it automatically identifies multiple types of entities, including APT organizations, vulnerability IDs, attack tools, malware families, attack stages, victim industries, and geographical regions. Simultaneously, it extracts semantic relationships such as exploitation, attack, location, target orientation, affiliation, and association, outputting standardized triple structures. This stage overcomes the limitations of traditional regular expression matching or CRF / LSTM sequence models, which can only capture shallow features such as IOCs. It can extract multiple entities and relationships in parallel within a single prompt template, forming a preliminary semantic skeleton of the APT attack chain. The ICL prompt template, by dynamically adjusting task instructions and example sets, can flexibly adapt to multi-task extraction scenarios involving APT organizations, vulnerability IDs, tools, and attack chain relationships. This effectively improves the extraction generalization ability under different report language styles and paragraph contexts, reducing the risk of relying on a specific template.
[0020] In the coarse-grained and fine-grained semantic fusion stage, for the multi-source heterogeneous entity and relationship information output in the previous stage, a prompt-driven large language model is first used for coarse-grained semantic grouping, mapping entities with different expressions in different documents to a predefined category pool (APT organizations, vulnerabilities, tools, industries, regions, etc.). Subsequently, within the same category pool, high-dimensional vectors generated by SecureBERT or TFIDF are used to perform multiple rounds of semantic similarity comparison and hierarchical clustering. Nodes with close semantic distances are gradually fused through progressive thresholding, thereby solving the information fragmentation problem caused by aliases, abbreviations, and translations of APT organizations, vulnerability numbers, and tools in different sources. This generates standardized entity nodes that can be uniquely identified, and the associated upstream and downstream relationship edges are updated synchronously to maintain the integrity of the knowledge chain.
[0021] Finally, in the knowledge graph construction phase, a knowledge graph of APT organization network threat intelligence is constructed based on standardized entity nodes and their relationship information after multiple rounds of coarse-grained and fine-grained semantic fusion. The graph uses nodes to represent APT organizations, vulnerabilities, attack tools, malware, attack stages, victim industries and regions, and edges to represent diverse semantic relationships such as exploitation, attack, location, target orientation, association, and affiliation. The output is a standard graph file for persistent storage and supports invocation by APT attack chain tracing and analysis platforms or security posture visualization engines, enabling visualized replay of attack chains, aggregation of cross-report attack patterns, and automated security policy linkage.
[0022] In this invention, the standardized threat intelligence corpus generated in the first step serves as the foundational data source for subsequent processing, providing input text for the sentence vector model to generate semantic embeddings. Simultaneously, the diverse threat intelligence content it encompasses forms the core material for constructing the example library, ensuring that the example library covers various entity and relationship scenarios. The processed standardized corpus is also used for LoRA lightweight fine-tuning of the large language model, enabling the model to adapt to network threat intelligence extraction tasks. This allows for accurate extraction of entities and relationships when inputting ICL prompt templates generated based on the corpus. These three elements form a progressive dependency relationship supported by the standardized corpus, jointly ensuring the accuracy and adaptability of threat knowledge extraction.
[0023] The example library is a pre-built resource independent of the standardized threat intelligence corpus. Its core materials come from high-quality network threat intelligence samples annotated by domain experts, covering multiple types of entities such as APT organizations, vulnerability numbers, and attack tools, as well as structured examples of semantic relationships such as "exploitation" and "attack". These examples exist in the form of triples of "input sentence + structured output", and each example is labeled with its source category and entity structure. This is used to compare semantic similarity with the input text in subsequent steps, providing example support for the dynamic construction of ICL prompt templates, thereby assisting the large language model in achieving multi-task adaptive extraction.
[0024] The standardized threat intelligence corpus and the example corpus are in the same format. The goal of this invention is to use a large model to obtain more standardized threat intelligence corpora, based on the already manually annotated example corpus as a source of information. This allows for the acquisition of more data using a large model and a small number of samples.
[0025] As a further preferred embodiment, in the aforementioned method, the multi-source cybersecurity text includes at least one of or a combination of APT analysis reports, vulnerability announcements, threat intelligence notifications, and social media posts;
[0026] The threat intelligence data related to the APT group includes APT group attack attribution reports, CVE vulnerability analysis, intelligence aggregation blogs, and security community posts, including Twitter and Telegram, or a combination thereof.
[0027] As a further preferred option, in the aforementioned method, the pre-trained sentence vector model includes one or a combination of SecureBERT, Sentence-BERT, or TFIDF+PCA.
[0028] The step of generating semantic embeddings and calculating similarity using a pre-trained sentence vector model involves using a pre-trained sentence vector model to perform high-dimensional vectorization encoding on the input text and a pre-built example library, and then calculating the cosine similarity between the input and example library vectors.
[0029] As a further preferred option, the structured triplet is expressed by the following formula:
[0030]
[0031] Among them, This represents the set of triples in the entire threat knowledge graph; This represents a headentity node, such as an APT organization or vulnerability number. This represents a tail entity node, such as an attack tool, attack target, or region location. It represents the semantic relationships between entities, including "exploit", "belong to", "attack", "located at", etc. It is the set of all entities in the knowledge graph; It is the set of all relation types.
[0032] As a further preferred embodiment, in the aforementioned method, the semantic aggregation method that combines a prompt-driven large language model with a vectorization method to generate standardized entity nodes and updated relational information, i.e., the coarse-grained and fine-grained semantic fusion stage, includes the following:
[0033] A prompt-driven large language model is used to perform coarse-grained semantic grouping, unifying different expressions into a predefined semantic category pool. Within the same category pool, SecureBERT or TFIDF vector representations are used in combination with multi-round cosine similarity comparison and hierarchical clustering. Nodes with close semantic distances are gradually fused through progressive thresholding to achieve fine-grained semantic alignment. A knowledge base maintained by domain security experts and a predefined synonym mapping table are introduced to perform secondary consistency verification and boundary adjustment on the multi-round fusion results, resolving ambiguity issues caused by aliases, translations, and abbreviations, and maintaining the dependency structure of upstream and downstream attack chains.
[0034] As a further preferred embodiment, in the aforementioned method, in the APT organization network threat intelligence knowledge graph containing specific nodes and edges, the nodes are used to represent APT organizations, vulnerabilities, attack tools, malware, attack stages, victim industries and regions, and the edges are used to describe various semantic relationships including exploitation, attack, affiliation, location, target orientation, and association.
[0035] Before outputting the APT organization network threat intelligence knowledge graph containing specific nodes and edges, cross-mapping verification and enhancement based on an external knowledge base are performed. The constructed graph nodes and relationships are matched and verified with the ATT&CK technology library, CVE vulnerability library and known IOC blacklist, missing attribute information is filled in and conflicting nodes are removed, thereby improving the accuracy and comprehensiveness of the graph.
[0036] The knowledge graph of this invention supports outputting standard graph structure files (such as Neo4j or GraphML format) and can be integrated into a security situation visualization platform to automatically generate a multi-dimensional interactive view of APT attack chains, which facilitates security analysts to conduct source tracing, situation assessment and policy linkage based on dependency chains based on the graph.
[0037] As a further preferred option, in the aforementioned method, the large language model adopts the LoRA lightweight fine-tuning method, which freezes the original pre-trained weights and only updates the low-rank matrix parameters in the attention sublayer and feedforward network, thereby significantly reducing the overhead of GPU memory and computing resources while maintaining inference accuracy, and adapting to rapid deployment of computing nodes of different sizes.
[0038] The LoRA lightweight fine-tuning formula is as follows:
[0039]
[0040] in, This is the original frozen weight matrix; It is a low-rank trainable matrix.
[0041] As a further preferred embodiment, in the aforementioned method, the stepwise fusion of semantically close nodes through progressive thresholding includes:
[0042] A lower similarity threshold is used in the initial aggregation stage to capture a broad range of synonymous candidate entities;
[0043] Furthermore, the similarity threshold is gradually increased in subsequent rounds of comparison, high-confidence nodes are aggregated, and a layered fusion method is adopted to prevent excessive merging from causing the loss of details in the knowledge structure.
[0044] Based on a second key aspect of the present invention, an electronic device is provided, comprising: at least one processor; a memory communicatively connected to the at least one processor; the memory storing a computer program that, when executed by the at least one processor, enables the at least one processor to implement the aforementioned method for automatically extracting network threat knowledge based on large models and semantic alignment.
[0045] Based on a third key aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed, implements the aforementioned method for automatically extracting network threat knowledge based on large models and semantic alignment.
[0046] By employing the technical solution of this invention, compared with the prior art, this invention achieves high-precision automatic extraction of APT organizations, vulnerabilities, tools and complex attack chain relationships in multi-source heterogeneous CTI data by combining a prompt-driven large language model and ICL multi-instance context retrieval, breaking through the limitation of traditional IOC matching that can only identify explicit static features.
[0047] Furthermore, this invention employs a semantic alignment strategy of multi-round coarse-grained semantic fusion to effectively unify synonymous entities with multiple aliases, abbreviations, and translations, ensuring the integrity of the graph structure and the closure of the relationship chain.
[0048] This invention further enhances the contextual richness and cross-document reasoning capabilities of the knowledge graph by automatically completing node attribute tags through cross-mapping with the ATT&CK and CVE knowledge bases.
[0049] Experimental results show that the accuracy of attack chain reconstruction on actual APT intelligence datasets is improved by 21% compared with the co-occurrence-based aggregation model, and the node consistency index is improved by 18%. At the same time, the LoRA fine-tuning and progressive aggregation strategy reduces inference resource consumption by 30%, providing key technical support for APT attack tracing, threat situation awareness and automated security policy generation. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, obtaining other drawings based on these drawings without creative effort still falls within the scope of the present invention.
[0051] Figure 1 The flowchart of an automatic network threat knowledge extraction method based on large model and semantic alignment in one embodiment of the present invention is shown.
[0052] Figure 2 An embodiment of the present invention illustrates a threat intelligence extraction framework based on a large model.
[0053] Figure 3 The diagram illustrates the process of splitting the original corpus and preprocessing samples in one embodiment of the present invention.
[0054] Figure 4 The structure of the manual annotation and training set construction module is shown in one embodiment of the present invention.
[0055] Figure 5 This invention illustrates an embodiment of the use of LoRA for lightweight fine-tuning of a large model.
[0056] Figure 6 A flowchart illustrating similar sentence retrieval using KNN in one embodiment of the present invention is shown.
[0057] Figure 7 This paper illustrates a method for extracting threat intelligence information using a large model in one embodiment of the present invention.
[0058] Figure 8 This illustrates a mechanism for aligning extraction results with a standard terminology set in one embodiment of the invention.
[0059] Figure 9 The present invention illustrates an embodiment of the entity result clustering fusion and relation alignment process.
[0060] Figure 10 The structure for constructing a network threat intelligence knowledge graph is shown in one embodiment of the present invention. Detailed Implementation
[0061] The preferred embodiments of the present invention will be described in detail below to provide a clearer understanding of the purpose, features, and advantages of the present invention. It should be understood that the following embodiments are not intended to limit the scope of the present invention, but are merely illustrative of the essential spirit of the technical solution of the present invention.
[0062] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known techniques associated with the invention may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.
[0063] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.
[0064] Please see Figure 1 The present invention provides an embodiment of an automatic network threat knowledge extraction method based on large models and semantic alignment, comprising steps S101-S105 executed by a computer hardware system:
[0065] Step S101: Collect threat intelligence data related to APT organizations from multi-source cybersecurity texts, perform sentence-level segmentation, text denoising, and remove redundant text on the collected data to generate a standardized threat intelligence corpus.
[0066] Step S102: Use a pre-trained sentence vector model to generate semantic embeddings for the input text and example library text from the standardized threat intelligence corpus. Calculate the similarity to retrieve K most similar examples from the existing manually annotated example library and construct an ICL prompt template containing task instructions, examples, and the current input.
[0067] Step S103: Input the ICL prompt template into the large language model that has been lightly tuned by LoRA, and extract structured triples of multiple types of threat intelligence entities and multiple semantic relationships.
[0068] Step S104: For the output threat intelligence entities and semantic relationships, a semantic aggregation method combining prompt-driven large language model and vectorization method is used to generate standardized entity nodes and updated relationship information.
[0069] Step S105: Based on the obtained standardized entity nodes and updated relationship information, construct an APT organization network threat intelligence knowledge graph containing specific nodes and edges, and output a structured graph file.
[0070] The following terms cover the key technical components of this invention, including large model semantic understanding, prompting learning, context-aware retrieval, cross-document semantic fusion, and knowledge graph construction.
[0071] In-Context Learning Prompt Template (ICL) refers to a contextual prompt structure built before input to a large language model, typically including three parts: task instructions, example samples, and the current input. In this invention, the ICL template is automatically generated in conjunction with semantic similarity retrieval, supporting multi-task switching such as APT organization affiliation, vulnerability exploitation, and tool targeting, significantly improving semantic consistency and contextual guidance capabilities.
[0072] LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning technique for large language models. By inserting low-rank matrices into the attention layer and feedforward network, it achieves model adaptation tasks by updating only a small number of parameters. In this invention, LoRA is used to perform lightweight fine-tuning of large models, effectively enhancing their knowledge alignment and extraction capabilities in safe contexts and significantly reducing training resource consumption.
[0073] The SecureBERT embedding model is a pre-trained language model specifically trained on cybersecurity corpora. It generates sentence-level or short text-level semantic embedding vectors to capture contextual similarity between entities. In this invention, it is used to calculate the similarity between input and examples, supporting semantic-driven example retrieval and classification.
[0074] Semantic alignment aggregation strategy: Addressing issues such as naming ambiguity, abbreviations, and translation differences in the extraction results, this strategy combines the suggestive capabilities of the language model with the similarity measure of the embedding vectors to achieve unified merging of synonymous entities. This invention employs a three-stage strategy of coarse-grained grouping + fine-grained clustering + knowledge base mapping to ensure the uniqueness of entity nodes and the continuity of upstream and downstream relationships.
[0075] Ternary structured extraction refers to extracting unstructured input text into a structured format of "entity-relationship-entity" using a large model, such as "APT29—Exploitation—CVE-2017-11882". This structuring process relies on multi-instance ICL templates, enabling the large model to simultaneously identify multiple entity types and their relationships, thus constructing a complete attack chain.
[0076] APT attack chain reconstruction mechanism: This refers to reconstructing the attack paths and behavioral chains of APT organizations by performing graph-level fusion of entities and relationships extracted from multiple reports. This invention utilizes multi-dimensional features such as entity timestamps, context paragraph numbers, and similarity tags to construct attack event sequences, supporting time-series visualization and replay of the attack chain.
[0077] Graph-structured knowledge representation: Threat intelligence entities such as APT organizations, vulnerability IDs, attack tools, and malware families are modeled as nodes, and the abstract relationships between them, such as exploitation, attack, and target orientation, are modeled as edges to form a knowledge graph. This invention's graph supports exporting to multiple standard graph formats such as GraphML and Neo4j, and can be integrated into visualization platforms for situational analysis.
[0078] Relationship edge synchronous update mechanism: In the process of entity fusion, in order to prevent the break of upstream or downstream relationships, this invention synchronously updates all edge information related to the fusion node, retains its original connectivity, and ensures the integrity and traceability of the knowledge graph structure.
[0079] Instruction-Example Dynamic Switching Mechanism: The ICL template design of this invention supports automatic switching of task instructions and example content based on the input paragraph topic (e.g., vulnerability number priority, APT tool priority). For example, in detecting the text "APT28 exploit CVE-2022-30190", examples of the "vulnerability exploitation" category will be prioritized to improve the model's ability to focus on relationship identification.
[0080] Multi-round semantic fusion clustering: This invention adopts a progressive similarity threshold setting. The first round captures generalized near-sense candidate entities (similarity > 0.7), the second round raises the threshold to accurately merge core nodes (similarity > 0.85), and finally combines the domain expert database for manual mapping verification to balance aggregation granularity and semantic accuracy.
[0081] APT tag auto-completion: During the knowledge graph construction phase, this invention performs cross-validation with the ATT&CK tactical library and the CVE vulnerability library. For example, after identifying "APT-C-23 attack T1059 script execution", it automatically completes the attack phase "Execution" and the known exploit tool "PowerShell", improving the completeness and interpretability of the knowledge graph.
[0082] Cross-document entity fusion algorithm: This invention supports semantic merging of the same entities in different reports. For example, “APT28”, “FancyBear”, and “Sofacy Group” describe the same attack organization in different documents. The system automatically completes entity unification through text context, semantic vectors and external alias dictionaries.
[0083] Threat Intelligence Multi-Label Classifier: A weakly supervised multi-label classification mechanism is introduced during the fine-tuning of the large language model. The extracted entities are dynamically bound with labels (such as malware / tools / organizations / tactics, etc.), which supports the model to automatically determine the optimal category in semantically overlapping scenarios, effectively reducing misclassification and redundant entities.
[0084] The GraphML format export module exports the final constructed knowledge graph structure as a standardized GraphML file, which includes metadata such as entity node attributes, semantic relationship edges, and source document reference information. It can be directly connected to the Neo4j graph database or other graph engines to achieve efficient graph query and visualization interaction.
[0085] Attack chain visualization rendering engine: As one of the application components of this invention, this module automatically generates an APT attack chain path view after accepting the graph structure input. It supports interactive functions such as node highlighting, path tracing, timeline evolution, and multi-dimensional switching (by organization / industry / vulnerability / stage) to assist analysts in tracing and predicting the source.
[0086] Dependency Chain-Driven Policy Linkage Interface: This invention supports pushing discovered attack dependency chains to the defense system policy interface based on graph construction, and automatically generating defense rules by combining YARA rules, EDR engine, etc., to achieve an integrated closed loop from intelligence extraction to policy linkage.
[0087] Example 1: A Network Threat Intelligence Extraction Method Based on Large Model and Semantic Alignment
[0088] Step 1: Multi-source intelligence text collection and preprocessing
[0089] 1. Sample collection:
[0090] a) Collect network threat-related text data from APT attack reports, CVE vulnerability announcements, CNVD databases, and security analysis tweets from the X platform (formerly Twitter). b) The samples cover 10 typical APT groups (such as APT28, Lazarus, TA505), and the collected reports are in formats including PDF, HTML, Markdown, and JSON to ensure text diversity and realistic context. 2. Sample Preprocessing:
[0091] Step 2: ICL prompt template construction and semantic retrieval
[0092] 1. Sample library preparation:
[0093] a) Construct triplet example pairs in the form of "input sentence + structured output" to cover categories such as attack organization identification, vulnerability number extraction, tool attribution and attack target; b) The total number of examples N=500, and each example is labeled with the source category and entity structure to facilitate downstream classification and matching.
[0094] 2. Vector encoding and similarity retrieval:
[0095] a) Encode the input text into sentence vectors using a pre-trained sentence vector model (such as SecureBERT), represented as:
[0096]
[0097] b) Similarly, each sample in the example library is encoded as follows: The cosine similarity between the input and the example is calculated using the following formula:
[0098]
[0099] Parameter description:
[0100] : The vector representation of the text to be extracted;
[0101] : Vector representation of the example text;
[0102] L2 norm;
[0103] : The semantic similarity between the i-th example and the current input.
[0104] c) Select the top K examples based on similarity (default K=5) and construct the ICL suggestion template:
[0105] Extract instruction + Examples 1~5 + Current input.
[0106] Step 3: Large Language Model Extraction and Triple Generation
[0107] 1. Model Configuration:
[0108] a) Use large language models like LLaMA or ChatGLM, perform lightweight fine-tuning via LoRA, freeze the backbone Transformer parameters, and only update the weights of the low-rank matrix.
[0109] b) Supports multi-task extraction capabilities, including named entity recognition, relationship recognition, and context completion.
[0110] 2. Extraction and execution process:
[0111] a) Input the ICL prompt template into the large model;
[0112] b) The model outputs structured intelligence triples, including APT organization, vulnerability number, attack target, tool name, relationship type, etc.
[0113] Step 4: Entity Standardization and Semantic Alignment
[0114] 1. Vectorized representation:
[0115] a) Use SecureBERT to embed each extracted entity into an entity vector. ;
[0116] b) Calculate the cluster center vector for entities of the same category. Then calculate the L2 distance between the two:
[0117]
[0118] Parameter description:
[0119] : The i-th entity vector;
[0120] : The center vector of category k;
[0121] 1. Entity Clustering and Fusion: a) Hierarchical clustering algorithm is used to merge entities in multiple rounds; b) The threshold is set at 0.7 in the first round to form preliminary clusters, and increased to 0.85 in the second round for precise unification; c) MITRE ATT&CK organization alias dictionary (e.g., APT28 ≈ Fancy Bear ≈ Sofacy) is introduced for manual boundary verification.
[0122] Step 5: Map Construction and Structure Optimization
[0123] 1. Definition of nodes and edges: a) Nodes include categories such as APT organizations, vulnerabilities, attack tools, target industries, and attack stages; b) Edges represent semantic relationships, including "exploitation", "attack", "belong to", "target", "located in", etc.
[0124] 2. Exporting and Visualizing the Graph Structure: a) Save the graph structure as GraphML format and import it into the Neo4j database; b) Each node carries attribute information such as entity name, source document ID, and extraction confidence, supporting graph view query and path backtracking.
[0125] 3. Connectivity assessment metrics:
[0126] The formula for calculating the connectivity of a graph is as follows:
[0127]
[0128] in:
[0129] The maximum number of nodes in the connected subgraph in the graph;
[0130] : The total number of all entity nodes extracted;
[0131] C: Connectivity of the graph structure; a higher value indicates stronger semantic chain integrity.
[0132] Experiments show that the connectivity rate of the graph constructed using the method of this invention reaches an average of 92.3%, which is 19.4% higher than the version without aggregation, significantly improving the structural closed loop and path integrity.
[0133] Example 2: Validation of Semantic Robustness of Large Model
[0134] To verify the stability of this invention in threat intelligence extraction under adversarial perturbations such as semantic ambiguity, diverse aliases, and unstructured expressions, the following experiment was designed:
[0135] 1. Adversarial Example Generation: Insert 30% of the adversarial perturbation samples into the test set, including: a) Entity multireference: such as APT28, Fancy Bear, Sofacy, etc. referring to the same organization; b) Expression style perturbation: insert semantic noise and abbreviations into sentences (such as using the abbreviation "APT" to replace the full organization name); c) Context variation: by randomly changing the order of triples or inserting interfering commentary statements.
[0136] 2. Experimental Results: a) Entity Coherence Assessment: Using the multi-round semantic alignment mechanism of this invention, entities with inconsistent naming can still be merged into unified nodes, with an average consistency score (Entity Coherence Score) ≥ 0.91; b) Extraction Robustness: Under multi-round prompting, the LoRA-tuned large model can still accurately reconstruct the relationship between APT organizations, attack tools, and tactics, with an F1 score of 89.4%, which is about 13.7% higher than the unhinted large model.
[0137] Example 3: Comparative Experiment with Mainstream Threat Intelligence Extraction Technologies
[0138] On a dataset of 6,620 manually labeled APT intelligence entries, the overall performance of this invention in entity recognition and triplet relation extraction is compared with that of existing mainstream methods, as follows:
[0139] Table 1: Comparison of the overall performance of the present invention and existing mainstream methods in entity recognition and triplet relation extraction.
[0140]
[0141] Key conclusions:
[0142] Multi-round semantic alignment improves entity merging accuracy by 26.4% (compared to rule-based methods);
[0143] Compared to the traditional BERT sequence labeling model, the prompt-driven + LoRA fine-tuning mechanism improves the F1 score by 7.7% and significantly enhances cross-domain transfer capabilities.
[0144] This invention combines a multi-stage processing flow of "example-driven + entity standardization + knowledge graph construction" to effectively solve the problem of extracting secure text from multiple heterogeneous sources.
[0145] The technical terms, principles, or means related to the technical solutions of the present invention mentioned in the above embodiments, which are not described in detail above, are all well-known technologies or common practices that are known to those skilled in the art.
[0146] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for automatically extracting network threat knowledge based on large models and semantic alignment, characterized in that, This includes the following steps performed by the computer hardware system: Threat intelligence data related to APT organizations is collected from multi-source cybersecurity texts. The collected data is then segmented into sentences, denoised, and redundant text is removed to generate a standardized threat intelligence corpus. A pre-trained sentence vector model is used to generate semantic embeddings for the input text and example library text from the standardized threat intelligence corpus. By calculating similarity, K examples most similar to the input are retrieved from the existing manually annotated example library. An ICL prompt template containing task instructions, examples and the current input is constructed. The ICL prompt template is input into a large language model that has been lightly tuned by LoRA to extract structured triples of multiple types of threat intelligence entities and multiple semantic relationships; For the output threat intelligence entities and semantic relationships, a semantic aggregation method combining prompt-driven large language model and vectorization method is used to generate standardized entity nodes and updated relationship information. Based on the obtained standardized entity nodes and updated relationship information, an APT organization network threat intelligence knowledge graph containing specific nodes and edges is constructed, and a structured graph file is output.
2. The method for automatic extraction of network threat knowledge based on large models and semantic alignment according to claim 1, characterized in that, The multi-source cybersecurity text includes at least one or a combination of APT analysis reports, vulnerability announcements, threat intelligence notifications, and social media posts. The threat intelligence data related to the APT group includes APT group attack attribution reports, CVE vulnerability analysis, intelligence aggregation blogs, and security community posts, including Twitter and Telegram, or a combination thereof.
3. The method for automatic extraction of network threat knowledge based on large models and semantic alignment according to claim 1, characterized in that, The pre-trained sentence vector model includes one or a combination of SecureBERT, Sentence-BERT, or TFIDF+PCA. The step of generating semantic embeddings and calculating similarity using a pre-trained sentence vector model involves using a pre-trained sentence vector model to perform high-dimensional vectorization encoding on the input text and a pre-built example library, and then calculating the cosine similarity between the input and example library vectors.
4. The method for automatic extraction of network threat knowledge based on large models and semantic alignment according to claim 1, characterized in that, The structured triple is expressed by the following formula: ; in, This represents the set of triples in the entire threat knowledge graph; Represents the head entity node; Represents the tail entity node; Indicates the semantic relationships between entities; It is the set of all entities in the knowledge graph; It is the set of all relation types.
5. The method for automatically extracting network threat knowledge based on large models and semantic alignment according to claim 1, characterized in that, The semantic aggregation method, which combines a prompt-driven large language model with a vectorization approach, generates standardized entity nodes and updated relational information, including the following: A prompt-driven large language model is used to perform coarse-grained semantic grouping, and different expressions are uniformly aggregated into a predefined semantic category pool; Within the same category pool, SecureBERT or TFIDF vector representations are used in combination with multiple rounds of cosine similarity comparison and hierarchical clustering. Nodes with close semantic distances are gradually merged through progressive thresholding to achieve fine-grained semantic alignment. A knowledge base maintained by domain security experts and a predefined synonym mapping table are introduced to perform secondary consistency verification and boundary adjustment on the multi-round fusion results.
6. The method for automatic extraction of network threat knowledge based on large models and semantic alignment according to claim 1, characterized in that, In the APT organization network threat intelligence knowledge graph containing specific nodes and edges, the nodes are used to represent APT organizations, vulnerabilities, attack tools, malware, attack stages, victim industries and regions, and the edges are used to describe various semantic relationships including exploitation, attack, affiliation, location, target orientation, and association. Before outputting the APT organization network threat intelligence knowledge graph containing specific nodes and edges, cross-validation and enhancement based on an external knowledge base are performed. The constructed graph nodes and relationships are matched and verified with the ATT&CK technology library, CVE vulnerability library and known IOC blacklist, missing attribute information is filled in and conflicting nodes are removed, thereby improving the accuracy and comprehensiveness of the graph.
7. The method for automatic extraction of network threat knowledge based on large models and semantic alignment according to claim 1 or 6, characterized in that, The fine-tuning method for the large language model after LoRA lightweight fine-tuning includes: While freezing the original pre-trained weights, only the low-rank matrix parameters in the attention sublayer and the feedforward network are updated; The LoRA lightweight fine-tuning formula is as follows: ; in, This is the original frozen weight matrix; It is a low-rank trainable matrix.
8. The method for automatic extraction of network threat knowledge based on large models and semantic alignment according to claim 5, characterized in that, The stepwise fusion of semantically close nodes using a progressive threshold method includes: A lower similarity threshold is used in the initial aggregation stage to capture a broad range of synonymous candidate entities; Furthermore, in subsequent rounds of comparison, the similarity threshold is gradually increased, high-confidence nodes are aggregated, and a layered fusion method is adopted to prevent excessive merging from causing the loss of details in the knowledge structure.
9. An electronic device, characterized in that, include: At least one processor; A memory that is communicatively connected to the at least one processor; The memory stores a computer program that, when executed by the at least one processor, enables the at least one processor to implement the automatic network threat knowledge extraction method based on large models and semantic alignment as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed, it implements the automatic extraction method for network threat knowledge based on large models and semantic alignment as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Knowledge base alignment method and device, computer equipment and storage medium
CN109783582A
Traffic jam detection method and device, electronic equipment and storage medium
CN116311961A
Method for constructing computer education knowledge graph based on knowledge graph
CN117875412A
Equipment quality state grading method and system based on clustering algorithm
CN118965038A
Knowledge association learning method and system based on knowledge graph and virtual reality
CN119166830A
Cited By
Redundant fragment calculation verification and edge loss recovery error correction method for attack traceability graph construction
CN121509099A
Network threat behavior reasoning method and system based on large model retrieval enhancement
CN121525813A
False threat intelligence detection method and system based on dynamic graph comparative learning
CN121711149A
A method and system for detecting false threat intelligence based on dynamic graph contrastive learning
CN121711149B