Network security domain knowledge graph entity alignment method and system based on large language model
Through a multi-round entity alignment method based on a large language model, the problems of effective representation of multi-source threat intelligence data and insufficient entity alignment accuracy in the field of network security are solved, and efficient and explainable entity alignment is achieved to adapt to the complex and changing scenarios in the field of network security.
Patent Information
- Application Number
- CN202510642253.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional entity alignment methods have difficulty constructing effective representations of multi-source threat intelligence data in the field of network security. The entity alignment accuracy is insufficient and the results are poorly interpretable.
A multi-round entity alignment method based on a large language model is adopted, combined with multi-angle similarity analysis and external knowledge description. Through standardization processing, pre-alignment and multi-round entity alignment, a large language model is used for comprehensive scoring and result confirmation to construct a mapping relationship between knowledge graphs.
It significantly improves the accuracy and interpretability of entity alignment, adapts to complex and changing scenarios in the field of network security, reduces false matches and missed matches, supports system expansion and updates, and adapts to new threat intelligence needs.
Smart Images

Figure CN120675737A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for aligning entities in a knowledge graph in the field of network security, and belongs to the interdisciplinary technical field of intelligent network security analysis and knowledge graph construction. Background Art
[0002] A knowledge graph is a semantic network used to represent entities, attributes, and their relationships, and is widely used in fields such as cybersecurity, healthcare, and finance. Entity alignment is a key task in knowledge graph fusion, aiming to identify nodes representing the same entity across different knowledge graphs. Traditional entity alignment methods rely primarily on techniques such as knowledge representation learning, graph neural networks (GNNs), and pre-trained language models to extract and align entity representations from structured or unstructured data. However, these methods still have limitations when dealing with complex semantics, cross-language alignment, or low-resource scenarios.
[0003] In the field of cybersecurity, knowledge graphs are widely used in tasks such as threat intelligence analysis, intrusion detection, malware analysis, and security situational awareness. Since cybersecurity data typically comes from a wide range of sources and in diverse formats, including log information, attack events, and threat intelligence reports, efficiently integrating, correlating, and reasoning about this heterogeneous data presents a key challenge. Knowledge graphs, with their powerful relational modeling capabilities, can structure this dispersed information, revealing potential attack paths, threat sources, and impact scope, thereby improving the efficiency and accuracy of security analysis.
[0004] However, cybersecurity knowledge graphs often involve multiple data sources and intelligence repositories from different organizations, resulting in the same attack behavior or malicious entity being represented differently in different graphs. This makes entity alignment a core issue in knowledge fusion and reasoning, directly impacting the integrity and accuracy of threat intelligence. Efficient entity alignment methods can improve the collaboration between different security systems, reduce information redundancy, and enhance the response to new security threats. Therefore, researching efficient, accurate, and applicable entity alignment methods for cybersecurity scenarios is crucial for building intelligent and automated security defense systems.
[0005] In recent years, large language models (LLMs) have demonstrated powerful semantic understanding and reasoning capabilities, providing a new solution for entity alignment. Combining traditional methods with LLMs can significantly improve the accuracy and robustness of entity alignment. Therefore, this paper proposes a method for entity alignment in the network security knowledge graph based on a large language model. Summary of the Invention
[0006] In order to solve the problems faced by traditional entity alignment methods in the field of network security, such as the difficulty in constructing effective representation of multi-source threat intelligence data, insufficient entity alignment accuracy, and poor interpretability of results, the present invention proposes a knowledge graph entity alignment method and system in the field of network security based on a large language model.
[0007] The technical solution adopted by the present invention to solve the above problems is: the steps of the method of the present invention include:
[0008] Step 1: Prepare data; input the network security knowledge graphs KG1 and KG2 to be matched;
[0009] Step 2: Standardize entity and relationship types. Standardize entity attributes and attribute values in the knowledge graph according to international standards. Use LLM to adjust the expression specifications of entity attribute types and values to ensure the uniformity and consistency of data expression.
[0010] Step 3: Pre-alignment: Use the traditional entity alignment algorithm to perform preliminary representation extraction and pre-alignment on the entities in the knowledge graph. Find the 1, 10, and 20 entities with the highest similarity in KG2, and construct the candidate entity sets respectively and
[0011] Step 4: Multiple rounds of entity alignment;
[0012] Step 5: Knowledge graph alignment.
[0013] Furthermore, the pre-alignment algorithms used in step 3 include: an entity alignment algorithm based on a representation learning model, an entity alignment algorithm based on a graph neural network, and an entity alignment algorithm based on a pre-trained language model of a description text.
[0014] Furthermore, in step 4, for each entity in KG1 Three rounds of entity alignment are performed respectively. The steps of the rth round entity alignment are as follows:
[0015] Step 401: Obtain entity information;
[0016] Use graph traversal algorithm to extract entities to be aligned on the knowledge graph Entity information With structured information and this round of candidate entity sets Each entity Entity information With structured information
[0017] Step 402: Construct prompt words;
[0018] According to the entity to be aligned With each candidate entity Node information info(e), structured information s_info(e), and entity alignment hint words are constructed based on the alignment hint word template
[0019] Step 403: Multi-angle similarity analysis;
[0020] The entities to be aligned With each candidate entity in this round Constructed prompt word set The message is sent to the large language model, which analyzes the message from the perspectives of node name, node information, structural information, and background information based on the prompt word content. and It describes the possibility of the same real entity, outputs the analysis process of each angle and gives a possibility score (1 to 5 points), and finally give a comprehensive score based on all aspects (1-5 points);
[0021] Step 404: Confirm the alignment result of this round;
[0022] After the prompt words in this round of prompt word set are processed, the evaluation results and comprehensive evaluation of each angle in each prompt word are analyzed, and the entity pair with the highest score is selected in the order of comprehensive score and average score of each angle.
[0023] Construct result confirmation prompt words based on confirmation prompt word template Let the big model review the alignment results of this round;
[0024] If the large model returns YES and the comprehensive score If the similarity exceeds the preset threshold θ, the alignment result of this round is recognized and the entity pair describing the same real thing is found. Stop The next round of alignment;
[0025] If the large model returns NO, or the comprehensive score If the threshold value θ is not exceeded, it means that there is no entity in the candidate entity range that matches The corresponding entity returns to step 401 and uses Perform the r+1th round of entity alignment;
[0026] If no entity pair with a comprehensive score exceeding the threshold θ is found after three rounds of entity alignment, it means that the entity in KG1 There is no corresponding entity in KG2, terminate the entity Multi-round entity alignment process.
[0027] Furthermore, in step 5, for each entity in KG1 Perform the multiple rounds of alignment described in step 4 once and generate a Construct a mapping relationship set from KG1 to KG2:
[0028]
[0029] in express When the similarity score of two entities exceeds the threshold, it is considered that the two entities in different knowledge graphs correspond to the same real thing.
[0030] The system of the present invention includes a data preparation module, an entity and relationship type standardization module, a pre-alignment module, a multi-round entity alignment module and a knowledge graph alignment module;
[0031] Data preparation module, used to input the cybersecurity knowledge graph to be aligned;
[0032] The entity and relationship type standardization module is used to standardize the entity attributes and attribute values in the knowledge graph with reference to international standards to ensure the uniformity and consistency of data expression;
[0033] The pre-alignment module is used to extract and pre-align entities in the knowledge graph using traditional entity alignment algorithms, narrowing the scope of entity alignment and improving the efficiency of subsequent multiple rounds of alignment;
[0034] The multi-round entity alignment module performs three rounds of entity alignment for each entity in the knowledge graph to be aligned;
[0035] The knowledge graph alignment module builds a set of mapping relationships between two input knowledge graphs based on all found entity pairs and summarizes the alignment results between knowledge graphs.
[0036] The beneficial effects of the present invention are:
[0037] 1. By introducing a large language model (LLM) and a multi-round alignment strategy, this paper addresses the difficulty of traditional methods in constructing effective representations of multi-source threat intelligence data in the field of network security. Compared with traditional methods that rely on static rules or a single model, this paper can dynamically learn and update entity representations, adapting to the complex and changing entity types and relationship structures in the field of network security.
[0038] 2. This invention combines multi-angle similarity analysis and external knowledge description to significantly improve the accuracy of entity alignment. By conducting comprehensive analysis from multiple angles such as node name, node information, structural information, and background information, and introducing a similarity threshold judgment mechanism, it effectively reduces false matches and missed matches.
[0039] 3. This invention supports configuring entity similarity comparison rules based on entity type, clarifying the weights and priorities of different analysis angles, and solving the problem of poor flexibility in comparison rules of traditional methods. By dynamically adjusting the comparison rules, the system can adapt to the entity alignment requirements in different scenarios.
[0040] 4. This invention uses a large language model to generate a multi-angle analysis process and comprehensive score, providing interpretability of the alignment results. Each step of the analysis process and score results are traceable, making it easier for security analysts to understand and verify the rationality of the alignment results, solving the problem of poor interpretability of traditional methods.
[0041] 5. The modular design of the present invention supports integration with existing network security architectures. At the same time, by combining RAG technology and the external knowledge retrieval capabilities of large language models, the system can continuously expand and update the knowledge base to adapt to new network security threats and intelligence needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flowchart of the present invention;
[0043] Figure 2 This is a schematic diagram of the structure of the knowledge graph entity alignment system in the field of network security based on the large language model of the present invention;
[0044] Figure 3 This is the flowchart of step 4. DETAILED DESCRIPTION
[0045] Specific implementation method 1: Figures 1 to 3 As shown in the figure, the entity alignment method of the knowledge graph in the field of network security based on the large language model includes the following steps:
[0046] Step 1: Prepare data; input the network security knowledge graphs KG1 and KG2 to be matched;
[0047] Step 2: Standardize entity and relationship types. Standardize entity attributes and attribute values in the knowledge graph according to international standards. Use LLM to adjust the expression specifications of entity attribute types and values to ensure the uniformity and consistency of data expression.
[0048] Step 3: Pre-alignment: Use the traditional entity alignment algorithm to perform preliminary representation extraction and pre-alignment on the entities in the knowledge graph. Find the 1, 10, and 20 entities with the highest similarity in KG2, and construct the candidate entity sets respectively and
[0049] The pre-alignment algorithms used include: entity alignment algorithm based on representation learning model, entity alignment algorithm based on graph neural network, and entity alignment algorithm based on pre-trained language model of description text;
[0050] Step 4: Multiple rounds of entity alignment;
[0051] For each entity in KG1 Three rounds of entity alignment are performed respectively. The steps of the rth round entity alignment are as follows:
[0052] Step 401: Obtain entity information;
[0053] Use graph traversal algorithm to extract entities to be aligned on the knowledge graph Entity information With structured information and this round of candidate entity sets Each entity Entity information With structured information
[0054] Step 402: Construct prompt words;
[0055] According to the entity to be aligned With each candidate entity Node information info(e), structured information s_info(e), and entity alignment hint words are constructed based on the alignment hint word template
[0056] Step 403: Multi-angle similarity analysis;
[0057] The entities to be aligned With each candidate entity in this round Constructed prompt word set The message is sent to the large language model, which analyzes the message from the perspectives of node name, node information, structural information, and background information based on the prompt word content. and It describes the possibility of the same real entity, outputs the analysis process of each angle and gives a possibility score (1 to 5 points), and finally give a comprehensive score based on all aspects (1-5 points);
[0058] Step 404: Confirm the alignment result of this round;
[0059] After the prompt words in this round of prompt word set are processed, the evaluation results and comprehensive evaluation of each angle in each prompt word are analyzed, and the entity pair with the highest score is selected in the order of comprehensive score and average score of each angle.
[0060] Construct result confirmation prompt words based on confirmation prompt word template Let the big model review the alignment results of this round;
[0061] If the large model returns YES and the comprehensive score If the similarity exceeds the preset threshold θ, the alignment result of this round is recognized and the entity pair describing the same real thing is found. Stop The next round of alignment;
[0062] If the large model returns NO, or the comprehensive score If the threshold value θ is not exceeded, it means that there is no entity in the candidate entity range that matches The corresponding entity returns to step 401 and uses Perform the r+1th round of entity alignment;
[0063] If no entity pair with a comprehensive score exceeding the threshold θ is found after three rounds of entity alignment, it means that the entity in KG1 There is no corresponding entity in KG2, terminate the entity Multi-round entity alignment process;
[0064] Step 5: Knowledge graph alignment;
[0065] For each entity in KG1 Perform the multiple rounds of alignment described in step 4 once and generate a Construct a mapping relationship set from KG1 to KG2:
[0066]
[0067] in, express When the similarity score of two entities exceeds the threshold, it is considered that the two entities in different knowledge graphs correspond to the same real thing.
[0068] Among them, knowledge graph: KG1 = {E1, R1}, KG2 = {E2, R2},
[0069] Mapping function: M:E1→E2,
[0070] Similarity:
[0071] Multi-wheel alignment:
[0072] Mapping relationship set:
[0073] The fused knowledge graph: KG fusion=(E1∪E2,R1∪R2∪M),
[0074] Preferably, in step 2, the knowledge graph ontology structure and attribute value range can be constructed based on subsequent business needs, and in the normalization process of the network security intelligence knowledge graph, by mapping the incoming data to the defined ontology structure, the standardization and normalization of entity types, relationship types, entity attributes and their attribute values can be systematically achieved.
[0075] Preferably, in step 401, the RAG (Retrieval-Augmented Generation) technology can be combined to retrieve knowledge from the external threat intelligence database through the entity name information to obtain relevant external retrieval knowledge. At the same time, the external knowledge in the pre-trained large model in the field of network security is used to generate the external description of the entity. Form an external knowledge description of the entity.
[0076] Preferably, in step 402, the above external knowledge representation is organically integrated with the node information and structured information of the entity itself, and an enhanced entity alignment prompt word is constructed based on the entity alignment prompt word template.
[0077] Preferably, in step 403, the external knowledge description will be used as an independent analysis dimension to participate in the multi-angle similarity analysis together with the node name, node information, structured information and other angles to improve the accuracy and reliability of entity alignment.
[0078] Preferably, entity similarity comparison rules may be configured based on the aligned entity types in step 402. These rules should clearly define the comparison methods for different similarity analysis angles, as well as the weight and priority of each analysis angle in the comprehensive similarity analysis.
[0079] Preferably, in step 403, entity similarity comparison rules are referred to to perform similarity analysis on entity pairs in the prompt word from different perspectives, and further complete an overall analysis of the comprehensive similarity of the entity pairs.
[0080] Preferably, in step 404, when constructing the result confirmation prompt words, the following information can be integrated for reference by the large language model to support its final judgment on the alignment results of this round:
[0081] 1. Entity information: including the entity to be aligned and alternative entities Node names, attributes, relationships, etc.
[0082] 2. Structured information: This includes the surrounding structure description of the entity to be aligned and the candidate entity (such as adjacent nodes and their attributes, adjacency relationships, etc.) and the overall description of the structural information;
[0083] 3. External knowledge description (if any): external knowledge rag_info(e) retrieved from external threat intelligence databases using RAG technology, and external description extern_description(e) generated using pre-trained large models in the field of network security;
[0084] 4. Entity similarity comparison rules (if any): Similarity analysis rules configured for different entity types, including comparison methods and weightings for each analysis perspective;
[0085] 5. Multi-angle similarity analysis results: including scores from various angles such as node name, node information, structured information, and external knowledge description (if any) and comprehensive rating
[0086] By integrating the above information, construct the result confirmation prompt word For LLM to make final decision.
[0087] Specific implementation method 2: Figure 2 As shown in the figure, the network security knowledge graph entity alignment system based on the large language model includes a data preparation module, an entity and relationship type standardization module, a pre-alignment module, a multi-round entity alignment module, and a knowledge graph alignment module;
[0088] Data preparation module, used to input the cybersecurity knowledge graph to be aligned;
[0089] The entity and relationship type standardization module is used to standardize the entity attributes and attribute values in the knowledge graph with reference to international standards to ensure the uniformity and consistency of data expression;
[0090] The pre-alignment module is used to extract and pre-align entities in the knowledge graph using traditional entity alignment algorithms, narrowing the scope of entity alignment and improving the efficiency of subsequent multiple rounds of alignment;
[0091] The multi-round entity alignment module performs three rounds of entity alignment for each entity in the knowledge graph to be aligned;
[0092] The knowledge graph alignment module builds a set of mapping relationships between two input knowledge graphs based on all found entity pairs and summarizes the alignment results between knowledge graphs.
[0093] The data preparation module supports multiple data sources, including API procurement, platform purchase, acquisition through cooperation with relevant institutions, searching existing open source intelligence knowledge databases, and self-construction through the collection of open source threat intelligence technology, providing basic data support for subsequent entity alignment;
[0094] The entity and relationship type standardization module uses a large language model to adjust the entity attribute types and value expression ranges to eliminate data heterogeneity;
[0095] The pre-alignment module provides a set of candidate entities for multiple rounds of entity alignment by constructing a set of candidate entities.
[0096] The multi-round entity alignment module gradually selects entity pairs that describe the same real entity by obtaining entity information, constructing prompt words, performing multi-angle similarity analysis, and confirming the alignment results, ensuring the accuracy of the alignment results.
[0097] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A knowledge graph entity alignment method in the field of network security based on a large language model, characterized by: The specific steps include: Step 1: Prepare data; input the network security knowledge graphs KG1 and KG2 to be matched; Step 2: Standardize entity and relationship types. Standardize entity attributes and attribute values in the knowledge graph according to international standards. Use LLM to adjust the expression specifications of entity attribute types and values to ensure the uniformity and consistency of data expression. Step 3: Pre-alignment: Use the traditional entity alignment algorithm to perform preliminary representation extraction and pre-alignment on the entities in the knowledge graph. Find the 1, 10, and 20 entities with the highest similarity in KG2, and construct the candidate entity sets respectively and Step 4: Multiple rounds of entity alignment; Step 5: Knowledge graph alignment.
2. The network security knowledge graph entity alignment method based on a large language model according to claim 1 is characterized in that: The pre-alignment algorithms used in step 3 include: entity alignment algorithm based on representation learning model, entity alignment algorithm based on graph neural network, and entity alignment algorithm based on pre-trained language model of description text.
3. The network security knowledge graph entity alignment method based on a large language model according to claim 1 is characterized in that: In step 4, for each entity in KG1 Three rounds of entity alignment are performed respectively. The steps of the rth round entity alignment are as follows: Step 401: Obtain entity information; Use graph traversal algorithm to extract entities to be aligned on the knowledge graph Entity information With structured information and this round of candidate entity sets Each entity Entity information With structured information Step 402: Construct prompt words; According to the entity to be aligned With each candidate entity Node information info(e), structured information s_info(e), and entity alignment hint words are constructed based on the alignment hint word template Step 403: Multi-angle similarity analysis; Entities to be aligned With each candidate entity in this round Constructed prompt word set The message is sent to the large language model, which analyzes the message from the perspectives of node name, node information, structural information, and background information based on the prompt word content. and It describes the possibility of the same real entity, outputs the analysis process of each angle and gives a possibility score (1 to 5 points), and finally give a comprehensive score based on all aspects (1-5 points); Step 404: Confirm the alignment result of this round; After the prompt words in this round of prompt word set are processed, the evaluation results and comprehensive evaluation of each angle in each prompt word are analyzed, and the entity pair with the highest score is selected in the order of comprehensive score and average score of each angle. Construct result confirmation prompt words based on confirmation prompt word template Let the big model review the alignment results of this round; If the large model returns YES and the comprehensive score If the similarity exceeds the preset threshold θ, the alignment result of this round is recognized and the entity pair describing the same real thing is found. Stop The next round of alignment; If the large model returns NO, or the comprehensive score If the threshold value θ is not exceeded, it means that there is no entity in the candidate entity range that matches The corresponding entity returns to step 401 and uses Perform the r+1th round of entity alignment; If no entity pair with a comprehensive score exceeding the threshold θ is found after three rounds of entity alignment, it means that the entity in KG1 There is no corresponding entity in KG2, terminate the entity Multi-round entity alignment process.
4. The network security knowledge graph entity alignment method based on a large language model according to claim 1 is characterized in that: In step 5, for each entity in KG1 Perform the multiple rounds of alignment described in step 4 once and generate a Construct a mapping relationship set from KG1 to KG2: in express When the similarity score of two entities exceeds the threshold, it is considered that the two entities in different knowledge graphs correspond to the same real thing.
5. A knowledge graph entity alignment system in the field of network security based on a large language model, characterized by: It includes data preparation module, entity and relationship type standardization module, pre-alignment module, multi-round entity alignment module and knowledge graph alignment module; Data preparation module, used to input the cybersecurity knowledge graph to be aligned; The entity and relationship type standardization module is used to standardize the entity attributes and attribute values in the knowledge graph with reference to international standards to ensure the uniformity and consistency of data expression; The pre-alignment module is used to extract and pre-align entities in the knowledge graph using traditional entity alignment algorithms, narrowing the scope of entity alignment and improving the efficiency of subsequent multiple rounds of alignment; The multi-round entity alignment module performs three rounds of entity alignment for each entity in the knowledge graph to be aligned; The knowledge graph alignment module builds a set of mapping relationships between two input knowledge graphs based on all found entity pairs and summarizes the alignment results between the knowledge graphs.
Citation Information
Cited By
Network security threat monitoring method, system, device, medium and product
CN122179241A