Network security attack path deduction method and system based on multi-modal knowledge graph
Patent Information
- Application Number
- CN202610897221.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]本申请提供了基于多模态知识图谱的网络安全攻击路径推演方法及系统,旨在解决现有攻击路径推演无法精准还原真实复杂网络攻击链路,导致难以快速定位风险的技术问题,达到提升企业网络安全主动防御能力的技术效果
[0010] By constructing a standardized ontology model for the cybersecurity domain, a three-dimensional evaluation system based on name similarity, attribute overlap, and structural similarity is used to align cross-modal entities. A weighted adjudication mechanism combining source authority, timeliness, and modal reliability is introduced to resolve knowledge conflicts, generating a multimodal cybersecurity knowledge graph free of redundancy and logical contradictions. An incremental learning mechanism is also introduced, using entity fingerprint change detection to merge and update only newly added and changed data, reducing computational overhead and enabling real-time dynamic iteration of the knowledge graph. Based on this, a reverse search strategy is used to accurately identify high-risk attack entry points, and a graph search algorithm is used to generate candidate paths. These paths are then ranked using a multi-dimensional comprehensive confidence evaluation system that considers path integrity, cumulative edge weights, and other factors. Explanatory text with modal source tracing is generated, ultimately improving the credibility and interpretability of the inference results. This helps security personnel quickly identify core network vulnerabilities, deploy targeted protection strategies in advance, and effectively enhance enterprises' proactive cybersecurity defense capabilities.
Smart Images

Figure CN122601335A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security, specifically to a method and system for network security attack path deduction based on multimodal knowledge graphs. Background Technology
[0002] As cyberspace offense and defense continue to escalate, attacks are exhibiting multi-stage, cross-modal, and covert characteristics. Complex attack methods such as APT attacks and supply chain attacks are becoming increasingly common. Traditional cybersecurity attack path deduction technologies are mostly based on a single structured vulnerability database and static topology graph to build a knowledge system. They can only integrate structured vulnerability data and asset ledger information, and cannot effectively integrate multi-source heterogeneous security data such as semi-structured security device logs, unstructured threat intelligence reports, hacker forum texts, and binary characteristics of malicious code. This results in problems such as entity redundancy, attribute conflicts, and limited coverage in the knowledge graph. At the same time, traditional technologies use a full reconstruction method to update the knowledge graph, which has a long update cycle and consumes a lot of computing resources. It is difficult to incorporate newly disclosed zero-day vulnerabilities and new attack methods in real time. Furthermore, path search relies on fixed edge weights and does not take into account dynamic factors such as the exploitation status of vulnerabilities in the wild, the effectiveness of defense measures, and the timeliness of threats. The generated attack paths contain a large number of low-probability redundant links, the confidence assessment is distorted, and the results lack traceable explanatory basis. Security personnel cannot quickly locate core risks and formulate accurate protection measures. Summary of the Invention
[0003] This application provides a method and system for network security attack path inference based on multimodal knowledge graphs. It aims to solve the technical problem that existing attack path inference cannot accurately reconstruct real and complex network attack links, making it difficult to quickly locate risks, and achieve the technical effect of improving the proactive defense capabilities of enterprise network security.
[0004] In view of the above problems, this application provides a method and system for network security attack path inference based on multimodal knowledge graph.
[0005] The first aspect disclosed in this application provides a method for network security attack path deduction based on multimodal knowledge graphs, the method comprising:
[0006] Multimodal cybersecurity data is collected, and knowledge graph modeling and conflict resolution are performed on the multimodal cybersecurity data to generate a multimodal cybersecurity knowledge graph. An incremental learning mechanism is introduced to fuse and update the multimodal cybersecurity knowledge graph, constructing a cybersecurity knowledge update graph. Target assets are acquired, and attack entry point identification and dynamic graph search are performed on the cybersecurity knowledge update graph based on the target assets to generate multiple candidate cybersecurity attack paths. The comprehensive confidence of the multiple candidate cybersecurity attack paths is calculated, and the multiple candidate cybersecurity attack paths are sorted in descending order according to the comprehensive confidence and output with explanatory text generation to obtain a cybersecurity attack path deduction list.
[0007] Another aspect disclosed in this application provides a network security attack path inference system based on multimodal knowledge graphs, the system comprising:
[0008] The system comprises the following modules: a data acquisition module for collecting multimodal cybersecurity data, performing knowledge graph modeling and conflict resolution on the data, and generating a multimodal cybersecurity knowledge graph; an update module for introducing an incremental learning mechanism to fuse and update the multimodal cybersecurity knowledge graph, constructing a cybersecurity knowledge update graph; a search module for acquiring target assets, performing attack entry point identification and dynamic graph search on the cybersecurity knowledge update graph based on the target assets, and generating multiple candidate cybersecurity attack paths; and a calculation module for calculating the comprehensive confidence score of the multiple candidate cybersecurity attack paths, sorting the multiple candidate cybersecurity attack paths in descending order according to the comprehensive confidence score, generating explanatory text, and obtaining a cybersecurity attack path deduction list.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] By constructing a standardized ontology model for the cybersecurity domain, a three-dimensional evaluation system based on name similarity, attribute overlap, and structural similarity is used to align cross-modal entities. A weighted adjudication mechanism combining source authority, timeliness, and modal reliability is introduced to resolve knowledge conflicts, generating a multimodal cybersecurity knowledge graph free of redundancy and logical contradictions. An incremental learning mechanism is also introduced, using entity fingerprint change detection to merge and update only newly added and changed data, reducing computational overhead and enabling real-time dynamic iteration of the knowledge graph. Based on this, a reverse search strategy is used to accurately identify high-risk attack entry points, and a graph search algorithm is used to generate candidate paths. These paths are then ranked using a multi-dimensional comprehensive confidence evaluation system that considers path integrity, cumulative edge weights, and other factors. Explanatory text with modal source tracing is generated, ultimately improving the credibility and interpretability of the inference results. This helps security personnel quickly identify core network vulnerabilities, deploy targeted protection strategies in advance, and effectively enhance enterprises' proactive cybersecurity defense capabilities.
[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0012] Figure 1 A flowchart illustrating the network security attack path deduction method based on multimodal knowledge graph is provided for embodiments of this application;
[0013] Figure 2 A schematic diagram of the structure of a network security attack path inference system based on multimodal knowledge graph is provided for the embodiments of this application.
[0014] Figure labeling: Acquisition module 11, Update module 12, Search module 13, Calculation module 14. Detailed Implementation
[0015] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0016] The overall concept of the technical solution provided in this application is as follows:
[0017] This application provides a method and system for network security attack path inference based on multimodal knowledge graphs. It achieves unified mapping and modeling of multimodal data by pre-setting a standardized network security domain ontology model. A three-dimensional entity similarity evaluation system and a weighted adjudication mechanism are introduced to complete cross-modal entity alignment and conflict resolution. An incremental learning mechanism is combined to achieve real-time dynamic updates of the knowledge graph. Simultaneously, a reverse attack entry point identification strategy and a graph search algorithm based on four-dimensional dynamic edge weights such as vulnerability exploitability and defense measure strength are used to generate candidate attack paths. The paths are then ranked using a multi-dimensional comprehensive confidence evaluation system that considers path integrity and cumulative edge weight values, and explanatory text with traceable evidence is generated. Ultimately, this improves the credibility and interpretability of the inference results, helping security personnel quickly identify network security vulnerabilities and formulate targeted protection strategies, effectively enhancing the proactive network security defense capabilities of enterprises.
[0018] After introducing the basic principles of this application, various non-limiting embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0019] Example 1, as Figure 1 As shown in the embodiments of this application, a method for network security attack path inference based on multimodal knowledge graphs is provided. The method includes:
[0020] S100: Collect multimodal network security data, perform knowledge graph modeling and conflict resolution on the multimodal network security data, and generate a multimodal network security knowledge graph.
[0021] Specifically, the first step is to pre-define an ontology model for the cybersecurity domain. This involves building a standardized knowledge framework within the cybersecurity field and clearly defining all types of knowledge entities to be included, such as vulnerability CVE numbers, asset IP addresses, attack tools, attacker organizations, firewall policies, and the types of relationships between entities. This includes whether vulnerabilities exist on assets, how attack tools exploit vulnerabilities, and how defense measures block attack paths. This ontology serves as the unified benchmark for all subsequent data mapping and knowledge fusion, preventing knowledge confusion caused by differences in the representation of different modalities. Based on this, multimodal cybersecurity data is collected, which includes security data from different forms and sources. This includes structured vulnerability databases, such as NVD and CNNVD data, asset ledger table data, semi-structured security logs, such as firewall logs and intrusion detection system (IDS) logs, threat intelligence report XML / JSON data, as well as unstructured security bulletin texts, hacker forum posts, and malicious code binary signature data.
[0022] Next, these multimodal data are mapped one by one to a preset domain ontology model according to the data structure type. The multimodal network security entity set is then generated. For example, CVE-2025-1234 is extracted from the NVD vulnerability database as a vulnerability entity, and 192.168.1.100 is extracted from the enterprise asset ledger as a server asset entity. Then, the entity set is coded in a directed manner, and the business relationship between entities is transformed into a standard head entity-relationship-tail entity triple form, such as CVE-2025-1234-exists-192.168.1.100, forming a network security entity triple set. At the same time, all the original multimodal data is tagged, and each knowledge entry is attached with a source identifier, collection timestamp, and modality type. These tags will run through all subsequent knowledge processing stages, providing a traceability basis for conflict resolution and confidence calculation. Finally, based on the triple set and multimodal data tags, the initial modeling is completed, and an initial multimodal knowledge graph containing all the information of the original data but not yet subjected to cross-modal fusion and conflict processing is constructed.
[0023] After the initial knowledge graph is constructed, cross-modal entity alignment and merging operations are performed. This involves identifying and integrating knowledge entities describing the same objective object from different modalities. Specifically, a three-dimensional entity similarity evaluation dimension is first constructed, including name similarity, attribute overlap, and structural similarity. Name similarity is calculated using the edit distance algorithm to determine the degree of matching between entity name strings. Attribute overlap is calculated as the proportion of common attributes between two entities to their respective total attributes. Structural similarity measures the similarity of the connection relationships between two entities and other entities in the initial knowledge graph. Subsequently, all knowledge entities in the initial knowledge graph are traversed, and the comprehensive similarity between each pair is calculated to generate a knowledge entity similarity set. A preset similarity threshold of 0.85 is set, and entity pairs with a comprehensive similarity exceeding this threshold are selected as similar knowledge entity pairs. Finally, cross-modal alignment is performed on these entity pairs, merging all their attributes and relationships while retaining effective information, resulting in a multimodal entity-aligned knowledge graph that eliminates entity redundancy.
[0024] A weighted adjudication mechanism is then introduced to resolve conflicts, specifically addressing inconsistencies in attribute descriptions of the same entity across different modalities. First, all entities with inconsistent attribute descriptions are extracted from the aligned knowledge graph, forming a multimodal conflict entity set. Then, differentiated weights are assigned to attributes of different modalities based on the authority, accuracy, and timeliness of the data source. For example, the weight of the NVD official vulnerability database is set to 0.9, internal enterprise log data to 0.8, third-party threat intelligence platforms to 0.7, and unofficial information from hacker forums to 0.3. A weighted summation is used to calculate the comprehensive score for each conflicting attribute, and the attribute with the highest score is identified as the target conflicting entity attribute. Finally, the knowledge graph is revised based on the determined target conflicting entity attribute set, deleting contradictory attributes and relationships to generate a unified, redundant, and conflict-free final multimodal cybersecurity knowledge graph.
[0025] S200: Incremental learning mechanism is introduced to fuse and update the multimodal cybersecurity knowledge graph, and a cybersecurity knowledge update graph is constructed.
[0026] Specifically, the incremental learning mechanism is first defined as a machine learning method that learns and integrates only newly added or changed data while fully preserving existing knowledge. After introducing the incremental learning mechanism, multimodal real-time updated data is collected, covering security logs generated by firewalls, intrusion detection systems, and endpoint detection and response devices; vulnerability announcements recently released by the National Information Security Vulnerability Sharing Platform and the US National Vulnerability Database; the latest threat intelligence reports pushed by security vendors; newly captured malicious code binary samples; and real-time posts from hacker forums and dark web trading platforms. Then, change detection is performed on this real-time data based on the constructed multimodal network security knowledge graph. This is achieved by generating entity fingerprints, which are unique identifiers calculated from the entity's core attributes, relationships, and modal labels. The system compares the entities, relationships, and attributes in the real-time data with their corresponding entries in the existing data graph to accurately identify four types of changes: new entities (e.g., the recently disclosed CVE-2026-0001 remote code execution vulnerability, a newly deployed database server with the IP address 192.168.1.200), new relationships (e.g., CVE-2026-0001 exists at 192.168.1.200), modified attributes (e.g., the CVSSv3.1 score of a vulnerability being increased from an initial 7.5 to 9.8), and deleted entities (e.g., old office servers that have been permanently taken offline). All changes are then packaged into a structured incremental data change package.
[0027] Next, incremental data change packages are used to fuse and update the original multimodal cybersecurity knowledge graph. This process involves first performing standardized preprocessing on the incremental data, consistent with the initial knowledge graph construction phase. This includes mapping to a preset cybersecurity domain ontology model, generating head entity-relationship-tail entity triples, adding source identifiers, collection timestamps, and modality type labels. Then, cross-modal entity alignment is performed between entities in the incremental data and entities in the existing graph. This is achieved by using the three-dimensional evaluation dimensions of name similarity, attribute overlap, and structural similarity. If the similarity between an incremental entity and an existing entity exceeds a preset threshold, the newly added attributes and relationships of the incremental entity are merged into the existing entity. If it is a completely new entity that is not matched, it and its related triples are directly added to the original graph. After fusion, an incremental cybersecurity knowledge graph containing incremental knowledge is obtained.
[0028] Finally, consistency verification and correction updates are performed on the incremental cybersecurity knowledge graph. This involves a comprehensive check for logical inconsistencies in the merged graph, such as conflicting attribute values for the same entity, incorrect blocking relationships between defense measures and attack paths, and missing associations between assets and vulnerabilities. For conflicts discovered during verification, the weighted adjudication mechanism from the initial stage is reused. Weights are allocated based on the authority and timeliness of the data source for comprehensive adjudication. For example, if the vulnerability exploitation difficulty labeled by a third-party threat intelligence platform in the incremental data conflicts with the official description from NVD, the official NVD data with higher weight is used for correction, and conflicting attributes and relationships are deleted. Ultimately, a complete, logically consistent cybersecurity knowledge update graph containing the latest cybersecurity situation is constructed.
[0029] S300: Obtain the target asset, and based on the target asset, perform attack entry point identification and dynamic graph search in the network security knowledge update graph to generate multiple candidate network security attack paths.
[0030] Specifically, the first step is to acquire the target assets, which are the core network resources that the enterprise needs to protect. These typically include core database servers that store sensitive data, application servers that carry critical business operations, and enterprise intranet domain controllers. By accurately matching key fields such as asset IP address, asset unique identifier, and asset name in the network security knowledge update graph, the full set of related information of the target asset is extracted. This includes its operating system version, open TCP / UDP ports, installed system patches, deployed endpoint protection software, network zone, and connection relationships with other assets. This information is the basic premise for subsequent attack path deduction.
[0031] This process involves identifying attack entry points. Attack entry points refer to nodes where attackers may breach the enterprise network boundary and gain initial access to the internal network. These include web servers exposed to the public internet, unauthorized open remote desktop protocol ports, VPN gateways with configuration flaws, and uncontrolled IoT devices connected to the internal network. This embodiment employs a reverse search strategy to identify attack entry points. Starting with the target asset, it traces back through the knowledge graph to all upstream nodes that can reach the target asset, generating a set of potential attack entry point entities. Subsequently, it applies preset attack entry point judgment criteria to filter this set. The judgment criteria comprehensively consider the node's network exposure, vulnerability risk level, configuration compliance, and correlation with historical attacks. For example, nodes accessible from the public internet and with unpatched high-risk vulnerabilities (CVSS score ≥ 7.0), remote access nodes with weak or blank passwords, and nodes that have been recorded as attack attempts by the intrusion detection system are identified as feasible attack entry points. Office terminals accessible only from the internal network and with all patches installed and configured in accordance with security standards are excluded. This results in a concise and high-value set of feasible attack entry point entities.
[0032] Next, a dynamic graph search is performed based on the set of feasible attack entry points. This involves dynamically calculating the weight of each knowledge graph edge based on real-time network security situation and data characteristics. A higher weight indicates a higher feasibility for an attacker to complete the attack steps through that edge. First, a path edge weight evaluation factor is constructed, including vulnerability exploitability, defense strength, modal credibility, and temporal proximity. Vulnerability exploitability is quantified by combining the vulnerability's CVSS exploitability sub-score, the existence of publicly available exploit code, and its in-the-wild exploitation status. Edges with mature exploits and large-scale in-the-wild exploitation have significantly higher weights. Defense strength is assessed by evaluating the firewalls deployed along the attack path. The ability of security devices such as firewalls, intrusion prevention systems, and web application firewalls to block corresponding attack behaviors will be considered. If a firewall on a certain path has enabled protection rules for a specific vulnerability, the weight of that edge will be significantly reduced. Modal credibility is calculated based on the source tag attached to each relationship in the knowledge graph. Relationship edges from the NVD official vulnerability database and internal enterprise security logs have higher credibility than third-party threat intelligence, while edges corresponding to unofficial information from hacker forums have the lowest credibility. Temporal proximity assigns higher weight to newer security information. For example, a zero-day vulnerability disclosed a week ago has a higher weight than an older vulnerability disclosed a year ago, because attackers are more inclined to exploit the latest vulnerabilities that have not yet been widely patched.
[0033] After dynamically calculating edge weights, an A* heuristic search algorithm, inspired by the target asset, is employed. Starting from each feasible attack entry entity, the algorithm prioritizes paths with high edge weights and fewer hops to the target asset while traversing knowledge graph nodes, filtering out paths with more than a preset threshold of hops. Paths exceeding 8-10 hops are typically excluded because such attack paths are highly improbable in real-world scenarios and easily detected. The search process meticulously records the relationships between all entities and attack steps along each path, ultimately generating multiple logically complete candidate cybersecurity attack paths that closely resemble real-world attack scenarios.
[0034] S400: Calculate the overall confidence level of the multiple candidate network security attack paths, sort the multiple candidate network security attack paths in descending order according to the overall confidence level, and generate explanatory text to obtain a network security attack path deduction list.
[0035] Specifically, the first step is to calculate the overall confidence level of each candidate network security attack path. This involves scientifically evaluating the probability that the attack path will be actually exploited by an attacker in a real network environment using multi-dimensional quantitative indicators. In this embodiment, the confidence level is calculated using four dimensions: path integrity, cumulative edge weights, modal coverage, and time decay correction. The preset weights for these four dimensions are 40%, 20%, 20%, and 20%, respectively. The final overall confidence level ranges from 0 to 1, with higher values indicating a greater likelihood of an attack. Path integrity measures whether the attack path covers all necessary attack stages from the attack entry point to the target asset without logical breakpoints, and is quantified as 0 to 1. The numerical values represent the complete path coverage of the entire attack process: initial access → code execution → privilege escalation → lateral movement → data access. A score of 1 is awarded for each complete path, with 0.2 points deducted for each missing necessary attack step. The cumulative edge weight is the result of normalizing the dynamic weights of all knowledge graph edges in the path. These edge weights are the attack feasibility weights calculated during the dynamic graph search phase based on factors such as vulnerability exploitability and defense strength. For example, a path with 3 edges and weights of 0.9, 0.8, and 0.7 has a cumulative value of 2.4, which normalizes to 0.8. Modality coverage measures the diversity and reliability of the knowledge sources supporting the path, calculated as follows:
[0036] Number of different modes involved in the path / Total number of mode types × Weighted sum of the confidence levels of each mode;
[0037] Normalized to between 0 and 1, the path coverage from structured vulnerability databases, semi-structured security logs, and unstructured threat intelligence is significantly higher than that relying solely on a single third-party intelligence source. Time decay correction is used to reflect the time-sensitive nature of cybersecurity threats, employing an exponential decay function to correct for the collection time of all knowledge within the path. ,in The correction coefficient is calculated based on the number of days since the knowledge collection time was last estimated. The more recently disclosed the vulnerability and the more recently generated the security log, the higher the correction coefficient. Finally, the average of all knowledge correction coefficients in the path is taken as the score for this dimension. After multiplying the scores of the four dimensions by their corresponding weights and summing them, the comprehensive confidence of each candidate path can be obtained.
[0038] After calculating the overall confidence score of all candidate paths, the output is sorted in descending order. That is, all paths are sorted from highest to lowest overall confidence score. If the difference in confidence score between two paths is less than 0.05, the path with fewer attack steps (≤5 hops) is prioritized. This is because attack chains with fewer steps are easier for attackers to exploit quickly and are harder to detect. If the number of hops is the same, the path containing high-risk or critical vulnerabilities with a CVSS score ≥9.0 is prioritized. This sorting mechanism presents the most threatening attack paths to security personnel first.
[0039] Subsequently, explanatory text is generated using a step-by-step source-tracing generation method. For each attack node and relationship in each attack path, explanatory text with modal source references is generated. That is, the data source, modal type, and collection time of each knowledge conclusion are clearly marked. For example, for the step where the attacker uses the CVE-2026-3456 remote code execution vulnerability to invade the public Nginx web server, the explanatory text will indicate that the basic information of the vulnerability comes from the NVD official vulnerability database, the web server asset information comes from the enterprise asset ledger, and the relationship between the vulnerability and the asset comes from the enterprise vulnerability scanning report. At the same time, the explanatory text will also summarize the core risk points of the path and the potential harm after the attack is successful, so that security personnel can clearly understand the basis of the deduction conclusion.
[0040] The final generated list of network security attack paths is in structured table format, containing 10 fields: Rank (sorted in descending order of overall confidence score), Overall confidence score (rounded to two decimal places), Unique ID of the attack path, Attack entry entity, Target asset entity, Attack path steps (arrows connecting the entities to the attack), Path risk level (≥0.8 for extremely high risk, 0.6-0.8 for high risk, 0.4-0.6 for medium risk, <0.4 for low risk), Key risk points, List of modal source citations, and Preliminary protection recommendations. For example, the first-ranked entry is: Rank 1, Overall confidence score 0.92, Attack path IDP001, Attack entry point 192.168.1.10, Public Nginx network. Web server, target asset 192.168.2.50, core user database. Attack path steps: 192.168.1.10 → obtain webshell using CVE-2026-3456 → move laterally to 192.168.1.20 (internal file server) via SMB protocol → log in to the domain controller using a weak password from the domain administrator → access 192.168.2.50. Risk level: extremely high. Key risk points: The public web server has an unpatched 9.8-point high-risk remote code execution vulnerability; the domain administrator uses a weak password and has not enabled multi-factor authentication. Modal sources: NVD official vulnerability database, enterprise asset ledger, vulnerability scan report, domain controller security log. Preliminary protection recommendations: immediately patch the CVE-2026-3456 vulnerability, force the domain administrator to change to a strong password and enable MFA, and deploy a web application firewall to block malicious requests.
[0041] Furthermore, the method provided in this application embodiment, which performs knowledge graph modeling and conflict resolution on the multimodal network security data to generate a multimodal network security knowledge graph, includes: presetting a network security domain ontology model, wherein the network security domain ontology model includes knowledge entity types and entity relationship types; mapping the multimodal network security data to the network security domain ontology model for knowledge graph modeling to construct an initial multimodal knowledge graph; performing cross-modal entity alignment and merging on the initial multimodal knowledge graph to obtain a multimodal entity-aligned knowledge graph; and introducing a weighted adjudication mechanism to resolve conflicts in the multimodal entity-aligned knowledge graph to generate a multimodal network security knowledge graph.
[0042] Specifically, the first step is to pre-define a cybersecurity ontology model. This is a standardized semantic framework for expressing cybersecurity knowledge, constructed by cybersecurity experts in conjunction with international standards and actual enterprise security needs. It adopts a three-tiered structure: core layer, extension layer, and instance layer. The core layer defines the most fundamental knowledge categories in the cybersecurity field, including five major categories of knowledge entity types and twelve core entity relationship types. The knowledge entity types are further divided into: asset type, subdivided into servers, terminal devices, network devices, databases, and application systems; vulnerability type, subdivided into remote code execution vulnerabilities, SQL injection vulnerabilities, and privilege escalation vulnerabilities; threat type, subdivided into attack tools, malicious code, attack behaviors, and attack tactics; defense type, subdivided into firewalls, intrusion detection systems, web application firewalls, patches, and security policies; and attacker type, subdivided into individual hackers, hacker organizations, and APT groups. The specific relationship types include vulnerabilities existing in assets, attack tools exploiting vulnerabilities, attack behaviors exploiting vulnerabilities, and defense measures blocking attack behaviors. Attackers use attack tools, assets contain assets, and related vulnerabilities are also included. The core layer's definition strictly follows internationally accepted cybersecurity standards to ensure the universality and compatibility of knowledge representation. The extension layer supplements specific entity types and relationship types based on the security characteristics of the enterprise's industry. The instance layer is used to store entity instances and relationship instances extracted from specific multimodal data. The ontology model is constructed through an iterative optimization process that primarily uses expert definition and secondarily uses data-driven approaches. First, security experts manually define the core layer and the basic extension layer. Then, through the analysis and mining of historical multimodal security data, uncovered entity types and relationship types are automatically discovered and added to the ontology model after expert review and verification. At the same time, the ontology model is updated quarterly according to the latest cybersecurity standards and global threat landscape to ensure its timeliness and comprehensiveness.
[0043] Based on this, multimodal data mapping operations are performed on the collected multimodal cybersecurity data. This involves transforming raw data in different formats into knowledge elements conforming to ontology model specifications according to the data's structure type. For structured data, including NVD vulnerability databases, CNNVD vulnerability databases, and enterprise asset ledger tables, entities and relationships can be directly extracted through field mapping. First, a dedicated parser extracts structured fields, then maps them to the ontology model. For unstructured data, including security bulletin texts, hacker forum posts, and malware analysis reports, named entity recognition and relationship analysis are performed using a BERT pre-trained language model fine-tuned for cybersecurity. Extraction techniques are used to automatically identify cybersecurity entities and semantic relationships between entities in text. After entity and relationship extraction, all knowledge elements are oriented and encoded to transform business relationships between entities into standard head-entity-relationship-tail-entity triples. Each triple is then tagged with a source identifier, collection timestamp, and modality type. These tags will be used throughout all subsequent knowledge processing stages to provide a basis for resolving conflicts and calculating attack path confidence. Finally, based on the triple set and multimodal data tags, preliminary modeling is completed, constructing an initial multimodal knowledge graph that contains all information from the original data but has not yet undergone cross-modal fusion and conflict resolution.
[0044] Subsequently, cross-modal entity alignment and merging operations are performed, which involves identifying and integrating knowledge entities describing the same objective object in different modalities of data. Specifically, a three-dimensional entity similarity evaluation system is first constructed, including name similarity, attribute overlap, and structural similarity. Name similarity is calculated by combining string edit distance and semantic similarity. String edit distance is used to measure the literal matching degree of entity name strings, while semantic similarity is calculated by using a pre-trained cybersecurity domain model to calculate the cosine similarity of the semantic vectors of entity names. Attribute overlap is calculated as the proportion of attributes shared by two entities to the total number of attributes of each entity. Structural similarity is calculated using the SimRank algorithm to measure the similarity between two entities in the initial knowledge graph. The similarity of neighbor structures is used to determine the likelihood that two entities are the same entity if their connections with other entities are similar. Then, all knowledge entities in the initial graph are traversed, and the overall similarity between each pair is calculated, with name similarity accounting for 40%, attribute overlap for 35%, and structural similarity for 25%. This generates a knowledge entity similarity set. A preset similarity threshold of 0.85 is set, and entity pairs with overall similarity exceeding this threshold are selected as similar knowledge entity pairs. Finally, these similar entity pairs are cross-modal aligned and merged. During the merging process, all attributes and relationships of the two entities are retained, while duplicate attributes and relationships are deleted, resulting in a multimodal entity-aligned knowledge graph that eliminates entity redundancy.
[0045] A weighted adjudication mechanism is then introduced to resolve conflicts, specifically addressing contradictions in attribute descriptions or inter-entity relationship descriptions of the same entity across different modalities. Specifically, this involves first extracting all entities and relationships with inconsistent attribute values or contradictory relationship descriptions from the aligned multimodal entity alignment knowledge graph, forming a multimodal conflict entity set. Then, differentiated weights are assigned to attributes and relationships of different modalities based on the authority, accuracy, timeliness, and reliability of the data source and modality type. The authority weight is set as follows: NVD / CNNVD official vulnerability database 0.9, internal enterprise security logs 0.8, threat intelligence from mainstream domestic security vendors 0.7, international third-party threat intelligence 0.6, and unofficial information from hacker forums 0.3. The timeliness weight is calculated using an exponential decay function.
[0046] ;
[0047] in The weight of data is determined by the number of days since the data collection, with newer data receiving higher weights. The reliability weights for different modal types are set as follows: 0.9 for structured data, 0.8 for semi-structured data, and 0.7 for unstructured data. These three weights are multiplied to obtain the comprehensive weight for each conflicting attribute or relationship. A weighted summation is then used to calculate the comprehensive score for each conflicting attribute or relationship. The attribute or relationship with the highest score is determined as the final credible knowledge content. For example, regarding the CVSS score conflict mentioned above, the comprehensive weight of the official NVD data is 0.9 × 0.99 (assuming it was published 10 days ago × 0.9 = 0.8019), while the comprehensive weight of the third-party threat intelligence data is 0.7 × 0.95 (assuming it was published 30 days ago × 0.7 = 0.4655). Therefore, the official NVD score of 9.8 is ultimately used as the CVSS score for this vulnerability. Finally, based on the determined credible knowledge content, the multimodal entity alignment knowledge graph is corrected, contradictory attributes and relationships are deleted, and a unified, non-redundant, and logically conflict-free final multimodal cybersecurity knowledge graph is generated.
[0048] Furthermore, the method provided in this application embodiment maps the multimodal network security data to the network security domain ontology model for knowledge graph modeling to construct an initial multimodal knowledge graph, including: mapping the multimodal network security data to the network security domain ontology model according to the data structure type to obtain a multimodal network security entity set; performing directed encoding on the multimodal network security entity set to obtain a network security entity triple set, wherein the network security entity triple set is in the form of head entity-relationship-tail entity; performing tagging processing on the multimodal network security data to obtain multimodal network security data tags, wherein the multimodal network security data tags include source identifier, collection timestamp, and modality type; and performing knowledge graph modeling based on the network security entity triple set and the multimodal network security data tags to construct an initial multimodal knowledge graph.
[0049] Specifically, the collected multimodal network security data is first divided into three categories according to its data structure type, and differentiated multimodal data mapping operations are performed. The data structure type refers to the degree of standardization of data organization and storage. Structured data refers to well-organized data with a fixed table structure that can be directly queried through a database, including official vulnerability databases of NVD / CNNVD, enterprise asset ledgers, firewall policy configuration tables, etc. This type of data uses a field-level direct mapping method, mapping each column in the data table to the entity attributes in the preset network security domain ontology model one by one. The product name and version fields are mapped to the name attributes of asset entities, and the relationship between vulnerabilities and assets is automatically extracted. Semi-structured data refers to data with some structured features but no unified table structure as a whole, including firewall logs, IDS / EDR alarm logs, threat intelligence reports in XML / JSON format, etc. This type of data first extracts key structured fields through a dedicated parser and then maps them to the ontology model. For example, the source IP address, destination IP address, access port, access time, action, allow / block fields are parsed from the firewall logs and mapped to the attributes of attacker entities, asset entities, and port entities, respectively, and the relationship between attacker access to asset ports is generated.
[0050] Unstructured data refers to free text or binary data without a fixed structure, including vulnerability announcements released by security vendors, hacker forum posts, malicious code analysis reports, security incident investigation reports, etc. This type of data uses a BERT pre-trained language model fine-tuned based on the cybersecurity field. First, it uses named entity recognition technology to identify core entities such as vulnerabilities, assets, attack tools, and attackers in the text. Then, it uses relation extraction technology to extract semantic relationships between entities. Through the above classification and mapping, a multimodal cybersecurity entity set covering all modalities is finally obtained.
[0051] After entity extraction is completed, a directed encoding operation is performed on the entity set. That is, according to the causal logic of the attack behavior and the dependencies between entities, the associations between entities are transformed into a standard head entity-relationship-tail entity triple form. The head entity and tail entity correspond to the mapped knowledge entities, and the relationship corresponds to the semantic association between entities. The core of directed encoding is to clarify the directionality of the relationship, which is highly consistent with the unidirectional progressive characteristics of network attacks. At the same time, a globally unique identifier is assigned to each triple for easy subsequent tracing and management. Finally, a set of network security entity triples containing all entities and their relationships is formed.
[0052] Next, all raw multimodal data and their corresponding triples are tagged. Each triple is attached with a multimodal cybersecurity data tag containing a source identifier, a collection timestamp, and a modality type. The source identifier records the original source of the data, providing a basis for weighted adjudication in the subsequent conflict resolution phase. The collection timestamp is accurate to the second-level UTC time and records the time when the data was generated, providing a basis for time decay correction in subsequent incremental update change detection and attack path confidence calculation. The modality type is used to label the structured category to which the data belongs, including structured data, semi-structured log data, unstructured text data, and binary feature data, providing a basis for subsequent knowledge credibility assessment. All tags are bound and stored with their corresponding triples to ensure the traceability of knowledge.
[0053] Finally, knowledge graph modeling is performed based on the set of network security entity triples and multimodal data tags to construct an initial multimodal knowledge graph. Specifically, an attribute graph model is used, mapping each knowledge entity to a node in the graph. Node attributes include all inherent attributes of the entity and corresponding modality tag information. Relationships in each triple are mapped to directed edges in the graph, with edge attributes including relation type, source identifier, collection timestamp, and modality type. For example, the triple "CVE-2026-4567-exists-Windows Server 2022" will generate two nodes: the vulnerability node "CVE-2026-4567" and the asset node "Windows Server". The model includes 2022 and a directed edge from the vulnerability node to the asset node. The edge attributes include the relation type, the source identifier NVD official vulnerability database, the collection timestamp 2026-06-10T08:00:00Z, and the modality type structured data. The initial multimodal knowledge graph obtained after modeling completely retains all the information of the original data and includes all extracted entities and relations. However, cross-modal entity merging and conflict handling have not yet been performed, so there is some entity redundancy and logical contradiction.
[0054] Furthermore, in the method provided in this application embodiment, cross-modal entity alignment and merging are performed on the initial multimodal knowledge graph to obtain a multimodal entity-aligned knowledge graph, including: constructing an entity similarity evaluation dimension, which includes name similarity, attribute overlap, and structural similarity; traversing each knowledge entity in the initial multimodal knowledge graph according to the entity similarity evaluation dimension to calculate similarity, thereby obtaining a knowledge entity similarity set; filtering and extracting knowledge entities that exceed a preset similarity threshold based on the knowledge entity similarity set to obtain a set of similar knowledge entity pairs; and performing cross-modal entity alignment and merging on the set of similar knowledge entity pairs to obtain a multimodal entity-aligned knowledge graph.
[0055] Specifically, a three-dimensional entity similarity assessment dimension is first constructed, comprising name similarity, attribute overlap, and structural similarity. These three dimensions comprehensively characterize the degree of similarity between entities from three levels: literal identification, inherent attributes, and graph topological relationships, respectively. Name similarity measures the matching degree of entity names at both the literal and semantic levels. Addressing the widespread problem of mixed use of abbreviations, aliases, and numbers in the cybersecurity field, a weighted calculation method combining literal and semantic similarity is adopted. Literal similarity is calculated using the Levenshtein edit distance algorithm to determine the degree of character difference between two entity name strings. The formula is:
[0056] Levenshtein(s1,s2) = min{the minimum number of insertion, deletion, and replacement operations required to transform s1 into s2};
[0057] Then, the results are normalized to the 0-1 range by dividing by the maximum length of the two strings. Semantic similarity is calculated by converting entity names into 768-dimensional semantic vectors using a BERT-based pre-trained model fine-tuned on 1 million cybersecurity vulnerability announcements, threat intelligence, and log data. The cosine similarity of the two vectors is then calculated, resulting in name similarity = literal similarity × 0.6 + semantic similarity × 0.4. Attribute overlap measures the degree of matching of shared attributes between two entities. A weighted Jaccard similarity algorithm is used. First, based on the cybersecurity ontology model, entity attributes are divided into core attributes, including the CVE number of the vulnerability, the IP address of the asset, the MD5 hash value of the attack tool, and the family identifier of the malware, etc., and non-core attributes, including the release time of the vulnerability, the hardware model of the asset, and the author information of the attack tool. The weight of core attributes is set to 0.8, and the weight of non-core attributes is set to 0.2.
[0058] Attribute overlap = (sum of intersection weights of core attributes + sum of intersection weights of non-core attributes) / (sum of union weights of core attributes + sum of union weights of non-core attributes);
[0059] If two entities have any one core attribute that completely matches, the attribute overlap is directly assigned a value of 0.95. Structural similarity measures the topological similarity between two entities in the knowledge graph. Based on the fundamental graph theory assumption that similar entities have similar neighbors, an improved SimRank algorithm is used. It only calculates the weighted Jaccard similarity of the first-order neighbor set of the two entities. The weight of the neighbor node is dynamically determined according to its relationship type with the target entity, according to:
[0060] Structural similarity = Σ (product of the weights of common neighbor nodes) / Σ (sum of the weights of all neighbor nodes).
[0061] After completing the construction of the evaluation dimensions, a block-based traversal and distributed parallel computing strategy is adopted to calculate the similarity of all entities in the initial multimodal knowledge graph. To avoid the exponential computational complexity of O(n²), the entities are first divided into blocks according to the top-level type based on the preset cybersecurity domain ontology model, including vulnerability blocks, asset blocks, attack tool blocks, defense measure blocks, etc. Similarity calculation is only performed within the same type of entity block. Then, within each block, it is further subdivided according to the core attribute prefix, including vulnerability blocks divided by CVE number year, asset blocks divided by Class C IP network segment, attack tool blocks divided by the first 4 digits of MD5 hash, etc., reducing the computational complexity to an approximately linear level of O(n). Subsequently, the three-dimensional similarity of all entity pairs is calculated in parallel using the Spark distributed computing framework within each sub-block. Then, the comprehensive similarity is calculated according to the preset weight formula of name similarity × 0.4 + attribute overlap × 0.35 + structural similarity × 0.25, generating a knowledge entity similarity set containing the comprehensive similarity of all entity pairs.
[0062] Next, based on the comprehensive similarity threshold of 0.85 verified by a large number of experiments, the similarity set is filtered, and entity pairs with a comprehensive similarity ≥ 0.85 are extracted into a set of similar knowledge entity pairs. At the same time, duplicate undirected entity pairs are filtered out, and only one set and meaningless self-loop entity pairs are retained.
[0063] Subsequently, cross-modal entity alignment and merging operations were performed on the set of similar knowledge entity pairs. During the alignment process, a globally unique master entity ID was assigned to each similar entity group, and a mapping table from all entity IDs to the master entity ID was established. During the merging process, the principles of full retention and deduplication were strictly followed. All attributes and relationships of all entities within the group were merged into the master entity, and completely duplicated attribute values and undirected duplicate relationship edges were deleted. At the same time, the source identifier, collection timestamp, and modal type label corresponding to each attribute and relationship were fully retained. For example, the CVE-2021-44228 Log4j vulnerability ATT&CK T1203, which exploits the vulnerability to execute code, was merged into a single master entity. The master entity contains all attributes of the three entities, including CVSSv3.1 score of 9.8, affected versions 2.0-beta9 to 2.15.0, attack type remote code execution, disclosure date of 2021-12-09, and relationships, including those existing in Apache. Log4j assets were exploited by the Log4jShell tool and blocked by WAF rule R20211210. The source tags of each attribute and relationship were retained, including the NVD official vulnerability database, Alibaba Cloud Threat Intelligence Platform, and enterprise core IDS logs. Finally, a multimodal entity alignment knowledge graph was generated that eliminated entity redundancy and achieved deep cross-modal knowledge association.
[0064] Furthermore, the method provided in this application embodiment introduces a weighted adjudication mechanism to resolve conflicts in the multimodal entity alignment knowledge graph and generate a multimodal network security knowledge graph, including: obtaining a set of multimodal conflicting entities in the multimodal entity alignment knowledge graph; introducing a weighted adjudication mechanism to assign weights and make weighted adjudications on each modality source attribute in the set of multimodal conflicting entities to determine a set of target conflicting entity attributes; and resolving conflicts in the multimodal entity alignment knowledge graph based on the set of target conflicting entity attributes to generate a multimodal network security knowledge graph.
[0065] Specifically, the first step is to obtain a set of multimodal conflicting entities. This involves filtering out entities from the knowledge graph that has completed cross-modal entity alignment, where the same entity has at least two different attribute values or the same entity pair has at least two contradictory relationships. Conflict types are mainly divided into attribute conflicts and relation conflicts. Attribute conflicts refer to inconsistent descriptions of the same attribute of the same entity by different modal data. Relation conflicts refer to opposite descriptions of the relationship between the same pair of entities by different modal data. By traversing the attribute sets of all entities and the relation sets of all directed edges in the aligned graph, and comparing all values of the same attribute with all relations of the same entity pair, all conflicting entities and relations can be extracted to form a set of multimodal conflicting entities.
[0066] Next, a weighted decision-making mechanism is introduced. Instead of a simple majority rule, it assigns differentiated weights to information from different modalities based on differences in data credibility. The most credible knowledge content is determined through weighted calculation. The core of this mechanism is the construction of a comprehensive weighting system encompassing three dimensions: source authority, timeliness, and modal reliability. The weight calculation for each dimension uses a quantitative formula and has been verified with extensive cybersecurity data. The source authority weight is pre-set with a fixed benchmark value based on the data publisher's professional capabilities, data accuracy, and industry recognition. Specifically: NVD / CNNVD official vulnerability database 0.9, CISA (Cybersecurity and Infrastructure Security Agency) announcements 0.88, enterprise internal security device logs 0.8, official threat intelligence from leading domestic security vendors 0.75, internationally renowned third-party threat intelligence platforms 0.7, unofficial information disclosed on hacker forums 0.3, and anonymous network source information 0.1. The timeliness weight reflects the dynamic decay characteristics of cybersecurity knowledge and is calculated using an exponential decay function, with the formula:
[0067] Timeliness weight = ;
[0068] in The formula represents the number of days between the data collection time and the current conflict resolution time, ensuring that newer data has a higher weight. The modal reliability weight is set based on the degree of data structuring and the proportion of manual verification. Structured data, due to its standardized format and low error rate, has a weight of 0.9; semi-structured data, due to certain parsing errors, has a weight of 0.8; unstructured data, due to its difficulty in semantic understanding and susceptibility to ambiguity, has a weight of 0.7; and binary feature data, due to its strong uniqueness and high accuracy, has a weight of 0.95.
[0069] After calculating the weights for the three dimensions, a weighted decision is made for each conflicting attribute or relationship. The specific process is as follows: First, all source information for the conflicting item is collected, and the comprehensive weight of each piece of information is calculated.
[0070] Overall weight = Source authority weight × Timeliness weight × Modal reliability weight;
[0071] Then, for each candidate attribute value or candidate relation, the comprehensive weights of all information supporting that value / relationship are summed to obtain the comprehensive score of that candidate value; finally, the candidate value with the highest comprehensive score is selected as the final credible result for that conflict term.
[0072] Finally, based on the target conflict entity attribute set obtained from the adjudication, the multimodal entity alignment knowledge graph is conflict-resolved. Specific operations include: for attribute conflicts, updating the corresponding attribute value of the entity to the adjudicated target value, while retaining all original attribute values and their source tags as historical versions for subsequent tracing and re-adjudication; for relationship conflicts, deleting contradictory relationship edges with lower overall scores, retaining the trusted relationship edges with the highest scores; if the difference in overall scores among multiple relationships is less than 0.05, all relationships are retained and marked as pending verification for manual review by security personnel; after completing all conflict corrections, a global consistency check is performed on the graph to ensure that there are no logical contradictions such as the same entity having mutually exclusive attribute defense measures that simultaneously block and allow the same attack behavior, ultimately generating a unified, non-redundant, and logically conflict-free multimodal network security knowledge graph.
[0073] Furthermore, the method provided in this application embodiment introduces an incremental learning mechanism to fuse and update the multimodal cybersecurity knowledge graph, constructing a cybersecurity knowledge update graph. This includes: introducing an incremental learning mechanism to collect multimodal real-time updated data; performing change detection on the multimodal real-time updated data based on the multimodal cybersecurity knowledge graph to obtain incremental data change packages; using the incremental data change packages to perform incremental entity fusion on the multimodal cybersecurity knowledge graph to obtain a cybersecurity knowledge incremental graph; and performing consistency verification, correction, and updates on the cybersecurity knowledge incremental graph to construct the cybersecurity knowledge update graph.
[0074] Specifically, the incremental learning mechanism is defined as a highly efficient update method that selectively learns and integrates only newly added or changed data while fully preserving existing knowledge systems and relationships. After introducing the incremental learning mechanism, multimodal real-time updated data is collected through a combination of streaming real-time acquisition and scheduled batch retrieval. Differentiated acquisition strategies are adopted based on the update characteristics of different data sources: for security logs generated by internal enterprise security devices, including firewalls, IDS, EDR, and SIEM platforms, Kafka message queues are used to achieve millisecond-level streaming acquisition, capturing data such as abnormal network access, attack alerts, and endpoint behavior in real time; for official vulnerability databases such as NVD, CNNVD, and CISA, and threat intelligence from security vendors... The reporting platform enables minute-level real-time subscriptions via API interfaces, providing immediate access to newly disclosed vulnerability information, attack tool characteristics, and APT group activity intelligence. For unofficial data sources such as hacker forums, dark web trading platforms, and security communities, distributed crawlers are used for hourly targeted crawling to collect information such as leaked vulnerability exploit code, attack method discussions, and black market transactions. For enterprise asset change data, including newly launched servers, decommissioned old equipment, and network topology adjustments, daily scheduled synchronization is achieved through integration with the CMDB configuration management database, while also supporting manual submission of asset change information by security personnel. All collected real-time data undergoes preliminary cleaning to remove duplicate, invalid, and incorrectly formatted data, preparing it for subsequent change detection.
[0075] Subsequently, change detection is performed on the cleaned real-time data based on the existing multimodal cybersecurity knowledge graph. This involves accurately comparing the content differences between the real-time data and the existing graph to identify all knowledge entries that need updating. This relies on entity fingerprinting technology, which generates a unique, irreversible fingerprint identifier for each entity in the knowledge graph. The fingerprint is derived from the hash value of the entity's core attributes. The specific change detection process is as follows: First, the real-time data undergoes preprocessing operations identical to those in the initial graph construction phase, including mapping to the cybersecurity domain ontology model, generating head entity-relationship-tail entity triples, adding source identifiers, collection timestamps, and modality type labels. Next, the entity fingerprint generated for each entity in the real-time data is calculated and compared with the existing entity fingerprint database in the knowledge graph. A step-by-step comparison is performed; based on the comparison results, the changes are categorized into four types: first, newly added entities, i.e., entities whose fingerprints are not matched in the fingerprint database; second, newly added relationships, i.e., entities whose fingerprints already exist, but the relationships between entities do not appear in the graph; third, modified attributes, i.e., entities whose fingerprints match, but the non-core attribute values of the entities have changed; and fourth, deleted entities, i.e., entities that have been permanently taken offline, scrapped, or no longer included in the protection scope, as confirmed through CMDB synchronization or manual marking. Finally, all identified changes are packaged according to the entity dimension to generate a structured incremental data change package. Each change entry includes the change type, unique entity identifier, comparison of content before and after the change, data source tag, and collection timestamp, ensuring that every change is traceable.
[0076] Next, incremental entity fusion is performed on the original multimodal cybersecurity knowledge graph using incremental data change packages. First, all triples in the incremental data change packages undergo ontology consistency verification to ensure that their entity types and relation types fully conform to the preset cybersecurity domain ontology model. Triples that do not conform to the specifications are automatically corrected or marked for manual review. Then, cross-modal entity alignment is performed, using the previously constructed three-dimensional comprehensive similarity evaluation system: name similarity × 0.4 + attribute overlap × 0.35 + structural similarity × 0.25. The incremental entities are compared with entities in the existing graph for similarity calculation. Entity pairs with a comprehensive similarity ≥ 0.85 are considered the same entity, and the new incremental entity is... Added attributes and relationships are merged into existing entities. During the merging process, all source tags and timestamps are fully preserved to provide a basis for subsequent conflict resolution. For incremental entities with a comprehensive similarity of <0.85, they are determined to be new entities, and they and their related triples are directly added to the original knowledge graph. For attribute-type changes, the attribute values of the corresponding entities are directly updated, and the old attribute values are retained as historical versions. For entity-type changes, a logical deletion method is used to mark the entity as offline and disconnect all its relationships with other entities, but the entity's historical information is retained for traceability analysis. After completing the above fusion operations, a temporary graph containing all incremental knowledge is obtained, namely the cybersecurity knowledge incremental graph.
[0077] Finally, consistency verification and correction updates are performed on the incremental cybersecurity knowledge graph. The consistency verification mainly covers three types of contradictions: first, attribute conflicts, i.e., multiple different values for the same attribute of the same entity after fusion; second, relationship conflicts, i.e. contradictory relationships between the same pair of entities; and third, logical contradictions, i.e., relationships that do not conform to business logic. For all the contradictions found during the verification, the weighted adjudication mechanism in the initial construction phase is reused. A comprehensive weight is calculated based on the authority, timeliness, and modal reliability of the data source, and the knowledge content with the highest comprehensive score is selected as the final credible result. For highly controversial contradictions with a weight difference of less than 0.05, they are marked as pending manual review and pushed to the security operations platform for professional judgment. After all contradiction corrections are completed, the graph is globally indexed and the entity fingerprint database is synchronized, ultimately constructing a complete, logically consistent cybersecurity knowledge update graph that includes the latest cybersecurity situation.
[0078] Furthermore, the method provided in this application embodiment generates multiple candidate network security attack paths, including: identifying attack entry points based on the target asset in the network security knowledge update graph, and performing a reverse search for a set of potential attack entry point entities; pre-setting attack entry point judgment conditions, using the attack entry point judgment conditions to judge and filter the set of potential attack entry point entities, and determining a set of feasible attack entry point entities; and performing a dynamic graph search on the network security knowledge update graph based on the set of feasible attack entry point entities to generate multiple candidate network security attack paths.
[0079] Specifically, the process begins by identifying user-specified target assets—core resources that the enterprise needs to protect, including Oracle databases storing sensitive user data, application servers carrying core transactions, and enterprise intranet domain controllers. Attack entry points are identified within a real-time updated network security knowledge graph. An attack entry point is a node that allows attackers to breach network boundaries and gain initial access to the internal network. This embodiment employs a reverse search strategy for entry point identification. Starting with the target asset, the system traverses backwards along all directed edges in the knowledge graph pointing to the target asset. The traversed relationships are strictly limited to five categories related to attack propagation: accessed, controlled, exploited, and trusted. Each upstream node is added to the potential attack entry point entity set until all reachable network boundary nodes are traversed or the preset maximum number of reverse hops is reached (typically 10 hops). For example, starting with the core database 192.168.2.50, a reverse search will sequentially find intranet file servers, domain controllers, and office terminals with access relationships. Further, it will find public web servers, VPN gateways, remote desktop gateways, and other boundary nodes that can access these internal nodes, ultimately forming a set of potential entry points containing only nodes that can reach the target asset.
[0080] After collecting potential entry points, the set of potential attack entry point entities is screened using preset multi-dimensional quantitative attack entry point judgment criteria. These criteria comprehensively consider the attack feasibility and risk level of nodes, including four quantitative dimensions, the weights of which have been verified by numerous real attack cases: 1. Network exposure level (weight 0.4): If the asset's IP belongs to a public address range or has port mapping configured to the public network without an access control list, the score is 1; if it is only accessible from the internal network, the score is 0. 2. Vulnerability risk level (weight 0.3): If the asset has an unpatched high-risk vulnerability with a CVSSv3.1 score ≥ 7.0, the score is 1; if it has a medium-risk vulnerability with a score of 5.0-6.9, the score is 0.5; if there are no unpatched vulnerabilities, the score is 0. 3. Configuration compliance (weight 0.2): If the asset has configuration defects such as weak passwords, empty passwords, failure to enable multi-factor authentication, or open high-risk ports such as 3389 / 22 without access restrictions, the score is 0.3 for each defect, with a maximum score of 1. 4. Historical... Historical attack correlation, weighted at 0.1, is calculated by weighting the asset based on whether it has been recorded by IDS / EDR for attack attempts or alerts targeting the same vulnerability within the past 30 days. If so, the score is 1; otherwise, it is 0. The comprehensive risk score of each potential entry point is calculated by weighted summation. Nodes with a comprehensive score ≥ 0.6 are identified as feasible attack entry points. For example, a public Nginx server has an exposure score of 1, an unpatched remote code execution vulnerability with a CVSS score of 9.8 (vulnerability risk score of 1), no SSH multi-factor authentication enabled (configuration compliance score of 0.3), and has been recorded with 3 attack attempts in the past 7 days (historical correlation score of 1). The comprehensive score is 1×0.4 + 1×0.3 + 0.3×0.2 + 1×0.1 = 0.86, which is identified as a high-priority feasible attack entry point. Office terminals that are only accessed from the internal network, have been fully patched, and have compliant configurations usually have a comprehensive score below 0.3 and are directly excluded. This results in a concise and high-value set of feasible attack entry point entities.
[0081] Next, a dynamic graph search is performed based on the set of feasible attack entry points to generate candidate attack paths. This involves dynamically calculating the attack feasibility weight of each knowledge graph edge based on the real-time network security situation. A higher weight indicates a greater probability that the attacker will complete the corresponding attack steps through that edge. First, a four-dimensional edge weight evaluation system is constructed, incorporating vulnerability exploitability, defense strength, modal credibility, and temporal proximity. The dynamic weight calculation formula for each edge is as follows:
[0082] Dynamic weight = vulnerability exploitability × 0.4 + (1 - defense strength) × 0.3 + modal credibility × 0.2 + temporal proximity × 0.1;
[0083] The vulnerability exploitability score is calculated using a combination of CVSS exploitability sub-scores, accounting for 50% of the score. The presence of publicly available exploit code (POC / EXP) adds 0.3 points, and the presence of large-scale in-the-wild exploits adds another 0.2 points, with a maximum of 1. The strength of defensive measures is quantified based on whether the corresponding attack behavior protection rules are enabled on the firewalls, WAFs, IPSs, and other security devices deployed along the path. Enabling the latest rules results in a strength of 0.8, deployment without updated rules results in 0.4, and no deployment results in 0. Modal credibility directly reuses the source weights of relation edges in the knowledge graph. Temporal proximity is calculated using an exponential decay function, with the formula: Δt represents the number of days the knowledge corresponding to this relationship is collected, ensuring that the edge weights are higher for newly disclosed vulnerabilities and newly generated logs.
[0084] After completing the dynamic calculation of edge weights, the A* heuristic search algorithm, with the target asset as the inspiration, is adopted. Starting from each feasible attack entry entity, the evaluation function of the algorithm is f(n) = g(n) + h(n), where g(n) is the sum of the cumulative edge weights from the attack entry to the current node, and h(n) is the minimum number of hops from the current node to the target asset multiplied by 0.8, i.e., the global average edge weight, as the heuristic value. Each time, the node with the largest f(n) value is expanded first, and a double pruning rule is set: first, if the path hop count exceeds 8 hops, the expansion is terminated, because the probability of an excessively long attack chain occurring in a real scenario is extremely low and it is easily detected; second, if duplicate nodes appear in the path, the expansion is terminated. The path is immediately discarded to avoid creating loop paths. During the search process, all entities and attack steps along each path are fully recorded. For example, 192.168.1.10, the public Nginx server → obtains a Webshell using CVE-2026-7890 → moves laterally to 192.168.1.20, the internal file server, using the SMB protocol → logs into the domain controller using a weak password from the domain administrator → accesses 192.168.2.50, the core database. Finally, all complete logical paths from feasible attack entry points to target assets are collected, forming a set of multiple candidate network security attack paths.
[0085] Furthermore, in the method provided in this application embodiment, dynamic graph search is performed on the network security knowledge update graph based on the feasible attack entry entity set to generate multiple candidate network security attack paths, including: constructing path edge weight evaluation factors, which include vulnerability exploitability, defense measure strength, modal credibility, and temporal proximity; taking the target asset as the heuristic target, starting from each attack entry entity in the feasible attack entry entity set, performing edge weight evaluation and dynamic heuristic graph search on the network security knowledge update graph based on the path edge weight evaluation factors to generate multiple candidate network security attack paths.
[0086] Specifically, the first step is to identify attack entry points based on the user-specified target assets, namely the core resources that the enterprise needs to protect, including Oracle databases storing sensitive user data, application servers carrying core transactions, and enterprise intranet domain controllers. Attack entry points are identified in a real-time updated network security knowledge graph. An attack entry point is a node that an attacker can breach the network boundary and gain initial access to the internal network. This implementation uses a reverse search strategy to perform entry point identification. Starting from the target asset, it traverses backwards along all directed edges in the knowledge graph that point to the target asset. The types of relationships traversed are strictly limited to five categories related to attack propagation: accessed, controlled, exploited, and trusted. Each upstream node is added to the potential attack entry point entity set until all reachable network boundary nodes are traversed or the preset maximum number of reverse hops is reached, usually 10 hops. Finally, a set of potential entry points containing only nodes that can reach the target asset is formed.
[0087] After collecting potential entry points, the set of potential attack entry point entities is screened using preset multi-dimensional quantitative attack entry point judgment criteria. These criteria comprehensively consider the attack feasibility and risk level of nodes, including four core quantitative dimensions, the weights of which have been verified through numerous real attack cases: First, network exposure level, weight 0.4: if the asset IP belongs to the public address range or has port mapping configured to the public network without an access control list, the score is 1; if accessible only from the internal network, the score is 0. Second, vulnerability risk level, weight 0.3: if the asset has an unpatched high-risk vulnerability with a CVSSv3.1 score ≥ 7.0, the score is 1; if it has a medium-risk vulnerability with a score of 5.0-6.9, the score is 0.5; if there are no unpatched vulnerabilities, the score is 0. Third, configuration compliance, weight 0.2. If an asset has configuration defects such as weak passwords, empty passwords, failure to enable multi-factor authentication, or open high-risk ports such as 3389 / 22 without access restrictions, it will receive 0.3 points for each defect, with a maximum score of 1. Fourth is the historical attack correlation, with a weight of 0.1. If the asset has been recorded by IDS / EDR for attack attempts or alerts targeting the same vulnerability in the past 30 days, it will receive a score of 1; otherwise, it will receive a score of 0. The comprehensive risk score of each potential entry point is calculated by weighted summation. Nodes with a comprehensive score ≥ 0.6 are identified as feasible attack entry points and are classified as high-priority feasible attack entry points. Office terminals that are only accessible from the internal network, have been fully patched, and have compliant configurations usually have a comprehensive score below 0.3 and will be directly excluded. Finally, a concise and high-value set of feasible attack entry point entities is obtained.
[0088] Next, a dynamic graph search is performed based on the set of feasible attack entry points to generate candidate attack paths. This differs from traditional static graph searches that use fixed edge weights. It dynamically calculates the attack feasibility weight of each knowledge graph edge based on the real-time network security situation. A higher weight indicates a greater probability that an attacker will complete the corresponding attack steps through that edge. First, a four-dimensional edge weight evaluation system is constructed, including vulnerability exploitability, defense strength, modal credibility, and temporal proximity. The dynamic weight calculation formula for each edge is as follows:
[0089] Dynamic weight = vulnerability exploitability × 0.4 + (1 - defense strength) × 0.3 + modal credibility × 0.2 + temporal proximity × 0.1;
[0090] The vulnerability exploitability score is calculated using a combination of CVSS exploitability sub-scores, accounting for 50% of the score. The presence of publicly available exploit code (POC / EXP) adds 0.3 points, and the presence of large-scale in-the-wild exploits adds another 0.2 points, with a maximum of 1. The strength of defensive measures is quantified based on whether the corresponding attack behavior protection rules are enabled on the firewalls, WAFs, IPSs, and other security devices deployed along the path. Enabling the latest rules results in a strength of 0.8, deployment without updated rules results in 0.4, and no deployment results in 0. Modal credibility directly reuses the source weights of relation edges in the knowledge graph. Temporal proximity is calculated using an exponential decay function, with the formula... Δt represents the number of days the knowledge corresponding to this relationship is collected, ensuring that the edge weights are higher for newly disclosed vulnerabilities and newly generated logs.
[0091] After completing the dynamic calculation of edge weights, the A* heuristic search algorithm, which uses the target asset as the inspiration, is adopted. Starting from each feasible attack entry entity, the evaluation function of the algorithm is f(n) = g(n) + h(n), where g(n) is the sum of the cumulative edge weights from the attack entry to the current node, and h(n) is the minimum number of hops from the current node to the target asset multiplied by 0.8, i.e., the global average edge weight, as the heuristic value. Each time, the node with the largest f(n) value is expanded first. At the same time, a double pruning rule is set: first, if the number of hops in the path exceeds 8, the expansion is terminated, because the probability of an excessively long attack chain occurring in a real scenario is extremely low and it is easily detected; second, if a duplicate node appears in the path, the path is immediately discarded to avoid generating loop paths. During the search process, the relationship between all entities and attack steps traversed by each path is completely recorded. Finally, all complete logical paths from feasible attack entry to the target asset are collected to form a set of multiple candidate network security attack paths.
[0092] Furthermore, in the method provided in this application embodiment, calculating the comprehensive confidence of the multiple candidate network security attack paths includes: obtaining a confidence calculation dimension, wherein the confidence calculation dimension includes path integrity, cumulative edge weight value, modality coverage, and time decay correction; and performing a confidence weighted calculation on the multiple candidate network security attack paths based on the confidence calculation dimension to obtain the comprehensive confidence of the multiple candidate network security attack paths.
[0093] Specifically, it is necessary to first clarify that the comprehensive confidence level is a standardized quantitative indicator with a value range of 0 to 1. The higher the value, the higher the probability of the attack path and the higher the threat level. This embodiment adopts a four-dimensional evaluation system that has been trained and verified by hundreds of thousands of real attack and defense cases. The preset weights of the four dimensions are path integrity 40%, cumulative edge weight 20%, modality coverage 20%, and time decay correction 20%, to ensure that the evaluation results are consistent with the real attack behavior patterns.
[0094] First, path integrity is calculated. This dimension measures whether the attack path covers all core attack stages from initial access to the target asset without logical breakpoints. Core attack stages are defined according to the MITRE ATT&CK framework as five essential steps: initial access, code execution, privilege escalation, lateral movement, and target access. The absence of any one of these steps will prevent the attack from closing the loop. Quantification uses a deduction system: full coverage of all five core stages earns a maximum score of 1 point, with 0.2 points deducted for each missing core stage. If the path contains logical jumps, an additional 0.1 points are deducted. For example, a path that only includes two nodes (public web server → core database) and lacks the three core stages of code execution, privilege escalation, and lateral movement has a path integrity score of 1 - 0.2 × 3 = 0.4. However, a path that fully covers the entire process of obtaining a webshell using CVE-2026-7890 → escalating privileges to the server administrator → lateral movement to the file server via the SMB protocol → stealing domain administrator credentials → logging into the domain controller to access the core database has a path integrity score of 1 point.
[0095] Next, the cumulative edge weights are calculated. This dimension is a comprehensive quantification of the feasibility of each attack step in the path. The dynamic weights of each relation edge generated during the dynamic graph search phase are directly reused, including four-dimensional weights: vulnerability exploitability, defense strength, modal credibility, and temporal proximity. To eliminate the influence of path hop count on the cumulative value, a normalization method is adopted, namely:
[0096] Cumulative edge weight = Sum of dynamic weights of all edges on the path ÷ Number of hops on the path;
[0097] Then, modality coverage is calculated. This dimension measures the diversity and reliability of the knowledge sources supporting the attack path. The higher the diversity of knowledge sources, the smaller the impact of a single data source error on the path's credibility. The specific calculation consists of two steps: First, the modality diversity score is calculated by counting the number of different modality types involved in the path, including four categories: structured data, semi-structured data, unstructured text data, and binary feature data. This score is then divided by the total number of modality types, 4, to obtain the diversity score. Second, the average modality credibility is calculated by extracting the source credibility weight corresponding to each relation edge in the path, i.e., 0.9 from the NVD / CNNVD official vulnerability database and 0.8 from enterprise internal security device logs. The following data is used: Threat intelligence from leading domestic security vendors (0.75), unofficial information from hacker forums (0.3), and the average weight of all sources is taken. The final modality coverage is calculated as: Modality Diversity Score × 0.5 + Average Modality Credibility × 0.5. For example, if a path uses three modalities—NVD vulnerability database (structured, 0.9), enterprise EDR terminal logs (semi-structured, 0.8), and QiAnXin threat intelligence (unstructured, 0.75)—the modality diversity score is 3 ÷ 4 = 0.75, and the average modality credibility is (0.9 + 0.8 + 0.75) ÷ 3 ≈ 0.817. Therefore, the modality coverage is 0.75 × 0.5 + 0.817 × 0.5 ≈ 0.783.
[0098] Next, time decay correction is calculated. This dimension reflects the timeliness of cybersecurity threats; that is, the more recently disclosed a vulnerability or the more recently generated security incident, the higher the risk of the corresponding attack path, because the vulnerability's patching rate is lowest and the attacker's willingness to exploit it is strongest at this time. An exponential decay function is used for quantification, with the formula: Single-step time correction coefficient = , where Δt is the number of days between the knowledge collection time corresponding to the attack step and the current inference, and the final time decay correction score is the average of the time correction coefficients of all steps in the path.
[0099] After completing the sub-quantification of the four dimensions, a confidence-weighted calculation is performed. The scores of each dimension are multiplied by their corresponding preset weights and then summed to obtain the final comprehensive confidence of each candidate path. For example, if a candidate path has a path integrity score of 1, a cumulative edge weight value of 0.858, a modality coverage of 0.783, and a time decay correction of 0.99, its comprehensive confidence = 1×0.4+0.858×0.2+0.783×0.2+0.99×0.2≈0.4+0.172+0.157+0.198≈0.927, which is an extremely high-risk attack path. At the same time, a special correction rule is set. If the difference in comprehensive confidence between two paths is less than 0.05, the path with fewer hops is given higher confidence. Attack chains with fewer steps are easier to execute quickly and are harder to detect. If the number of hops is the same, the path containing a severe vulnerability with a CVSSv3.1 score ≥9.0 is given higher confidence.
[0100] In summary, the network security attack path deduction method based on multimodal knowledge graphs provided in this application has the following technical effects:
[0101] By constructing a standardized ontology model for the cybersecurity domain, a three-dimensional evaluation system based on name similarity, attribute overlap, and structural similarity is used to align cross-modal entities. A weighted adjudication mechanism combining source authority, timeliness, and modal reliability is introduced to resolve knowledge conflicts, generating a multimodal cybersecurity knowledge graph free of redundancy and logical contradictions. An incremental learning mechanism is also introduced, using entity fingerprint change detection to merge and update only newly added and changed data, reducing computational overhead and enabling real-time dynamic iteration of the knowledge graph. Based on this, a reverse search strategy is used to accurately identify high-risk attack entry points, and a graph search algorithm is used to generate candidate paths. These paths are then ranked using a multi-dimensional comprehensive confidence evaluation system that considers path integrity, cumulative edge weights, and other factors. Explanatory text with modal source tracing is generated, ultimately improving the credibility and interpretability of the inference results. This helps security personnel quickly identify core network vulnerabilities, deploy targeted protection strategies in advance, and effectively enhance enterprises' proactive cybersecurity defense capabilities.
[0102] Example 2, based on the same inventive concept as the network security attack path deduction method based on multimodal knowledge graphs in the aforementioned examples, such as... Figure 2 As shown in the figure, this application provides a network security attack path inference system based on multimodal knowledge graph, the system including:
[0103] The acquisition module 11 is used to acquire multimodal network security data, perform knowledge graph modeling and conflict resolution on the multimodal network security data, and generate a multimodal network security knowledge graph. The update module 12 is used to introduce an incremental learning mechanism to fuse and update the multimodal network security knowledge graph and construct a network security knowledge update graph. The search module 13 is used to acquire target assets, perform attack entry identification and dynamic graph search on the network security knowledge update graph based on the target assets, and generate multiple candidate network security attack paths. The calculation module 14 is used to calculate the comprehensive confidence of the multiple candidate network security attack paths, sort the multiple candidate network security attack paths in descending order according to the comprehensive confidence and generate explanatory text to obtain a network security attack path inference list.
[0104] Furthermore, the acquisition module 11 is also used to perform the following steps: preset a cybersecurity domain ontology model, which includes knowledge entity types and entity relationship types; map the multimodal cybersecurity data to the cybersecurity domain ontology model to perform knowledge graph modeling and construct an initial multimodal knowledge graph; perform cross-modal entity alignment and merging on the initial multimodal knowledge graph to obtain a multimodal entity-aligned knowledge graph; introduce a weighted adjudication mechanism to resolve conflicts in the multimodal entity-aligned knowledge graph and generate a multimodal cybersecurity knowledge graph.
[0105] Furthermore, the acquisition module 11 is also used to perform the following steps: mapping the multimodal network security data to the network security domain ontology model according to the data structure type to obtain a multimodal network security entity set; performing directed encoding on the multimodal network security entity set to obtain a network security entity triple set, wherein the network security entity triple set is in the form of head entity-relationship-tail entity; performing tagging processing on the multimodal network security data to obtain multimodal network security data tags, wherein the multimodal network security data tags include source identifier, acquisition timestamp and modality type; and performing knowledge graph modeling based on the network security entity triple set and the multimodal network security data tags to construct an initial multimodal knowledge graph.
[0106] Furthermore, the acquisition module 11 is also used to perform the following steps: constructing an entity similarity evaluation dimension, which includes name similarity, attribute overlap, and structural similarity; traversing each knowledge entity in the initial multimodal knowledge graph according to the entity similarity evaluation dimension to calculate similarity, thereby obtaining a knowledge entity similarity set; filtering and extracting knowledge entities that exceed a preset similarity threshold based on the knowledge entity similarity set to obtain a set of similar knowledge entity pairs; and performing cross-modal entity alignment and merging on the set of similar knowledge entity pairs to obtain a multimodal entity aligned knowledge graph.
[0107] Furthermore, the acquisition module 11 is also used to perform the following steps: acquiring the set of multimodal conflicting entities in the multimodal entity alignment knowledge graph; introducing a weighted adjudication mechanism to assign weights and make weighted adjudications on the modal source attributes in the set of multimodal conflicting entities to determine the set of target conflicting entity attributes; and resolving conflicts in the multimodal entity alignment knowledge graph based on the set of target conflicting entity attributes to generate a multimodal network security knowledge graph.
[0108] Furthermore, the update module 12 is also used to perform the following steps: introducing an incremental learning mechanism to collect multimodal real-time update data; performing change detection on the multimodal real-time update data based on the multimodal cybersecurity knowledge graph to obtain an incremental data change package; using the incremental data change package to perform incremental entity fusion on the multimodal cybersecurity knowledge graph to obtain a cybersecurity knowledge incremental graph; performing consistency verification, correction, and update on the cybersecurity knowledge incremental graph to construct a cybersecurity knowledge update graph.
[0109] Furthermore, the search module 13 is also used to perform the following steps: based on the target asset, identify attack entry points in the network security knowledge update graph, and perform reverse search on a set of potential attack entry point entities; preset attack entry point judgment conditions, and use the attack entry point judgment conditions to judge and filter the set of potential attack entry point entities to determine a set of feasible attack entry point entities; perform dynamic graph search on the network security knowledge update graph based on the set of feasible attack entry point entities to generate multiple candidate network security attack paths.
[0110] Furthermore, the search module 13 is also used to perform the following steps: constructing path edge weight evaluation factors, which include vulnerability exploitability, defense measure strength, modal credibility, and temporal proximity; taking the target asset as the heuristic target, starting from each attack entry entity in the feasible attack entry entity set, performing edge weight evaluation and dynamic heuristic graph search on the network security knowledge update graph based on the path edge weight evaluation factors, and generating multiple candidate network security attack paths.
[0111] Furthermore, the calculation module 14 is also used to perform the following steps: obtaining the confidence calculation dimension, which includes path integrity, cumulative edge weight, modality coverage, and time decay correction; performing confidence weighted calculation on the multiple candidate network security attack paths based on the confidence calculation dimension to obtain the comprehensive confidence of the multiple candidate network security attack paths.
[0112] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for network security attack path deduction based on multimodal knowledge graphs, characterized in that, The method includes: Collect multimodal network security data, perform knowledge graph modeling and conflict resolution on the multimodal network security data, and generate a multimodal network security knowledge graph; An incremental learning mechanism is introduced to fuse and update the multimodal cybersecurity knowledge graph, thereby constructing a cybersecurity knowledge update graph; The target asset is acquired, and attack entry points are identified and dynamic graph search is performed on the network security knowledge update graph based on the target asset to generate multiple candidate network security attack paths; Calculate the overall confidence level of the multiple candidate network security attack paths, sort the multiple candidate network security attack paths in descending order according to the overall confidence level, and generate explanatory text to obtain a network security attack path deduction list.
2. The network security attack path deduction method based on multimodal knowledge graph as described in claim 1, characterized in that, The multimodal cybersecurity data is subjected to knowledge graph modeling and conflict resolution to generate a multimodal cybersecurity knowledge graph, including: A preset cybersecurity domain ontology model is provided, which includes knowledge entity types and entity relationship types. The multimodal network security data is mapped to the network security domain ontology model for knowledge graph modeling, and an initial multimodal knowledge graph is constructed. Cross-modal entity alignment and merging are performed on the initial multimodal knowledge graph to obtain a multimodal entity-aligned knowledge graph; A weighted adjudication mechanism is introduced to resolve conflicts in the multimodal entity alignment knowledge graph, generating a multimodal network security knowledge graph.
3. The network security attack path deduction method based on multimodal knowledge graph as described in claim 2, characterized in that, The multimodal network security data is mapped to the network security domain ontology model for knowledge graph modeling, constructing an initial multimodal knowledge graph, including: The multimodal network security data is mapped to the network security domain ontology model according to the data structure type to obtain a multimodal network security entity set; The multimodal network security entity set is oriented and encoded to obtain a network security entity triple set, which is in the form of head entity-relationship-tail entity; The multimodal network security data is tagged to obtain multimodal network security data tags, which include source identifier, collection timestamp, and modality type; Knowledge graph modeling is performed based on the set of network security entity triples and the multimodal network security data tags to construct an initial multimodal knowledge graph.
4. The network security attack path deduction method based on multimodal knowledge graph as described in claim 2, characterized in that, The initial multimodal knowledge graph is subjected to cross-modal entity alignment and merging to obtain a multimodal entity-aligned knowledge graph, including: Construct entity similarity evaluation dimensions, which include name similarity, attribute overlap, and structural similarity; The similarity is calculated by traversing each knowledge entity in the initial multimodal knowledge graph according to the entity similarity evaluation dimension to obtain a knowledge entity similarity set; Based on the knowledge entity similarity set, knowledge entities that exceed the preset similarity threshold are filtered and extracted to obtain a set of similar knowledge entity pairs; Cross-modal entity alignment and merging are performed on the set of similar knowledge entity pairs to obtain a multimodal entity-aligned knowledge graph.
5. The network security attack path deduction method based on multimodal knowledge graph as described in claim 2, characterized in that, A weighted adjudication mechanism is introduced to resolve conflicts in the multimodal entity-aligned knowledge graph, generating a multimodal cybersecurity knowledge graph, including: Obtain the set of multimodal conflicting entities in the multimodal entity alignment knowledge graph; A weighted adjudication mechanism is introduced to assign weights and make weighted adjudications on the modal source attributes in the multimodal conflict entity set to determine the target conflict entity attribute set; Based on the target conflict entity attribute set, conflict resolution is performed on the multimodal entity alignment knowledge graph to generate a multimodal network security knowledge graph.
6. The network security attack path deduction method based on multimodal knowledge graph as described in claim 1, characterized in that, An incremental learning mechanism is introduced to fuse and update the multimodal cybersecurity knowledge graph, constructing a cybersecurity knowledge update graph, including: An incremental learning mechanism is introduced to collect multimodal real-time updated data. Based on the multimodal network security knowledge graph, change detection is performed on the multimodal real-time updated data to obtain incremental data change packages. The incremental data change package is used to perform incremental entity fusion on the multimodal cybersecurity knowledge graph to obtain an incremental cybersecurity knowledge graph; The incremental network security knowledge graph is subjected to consistency verification, correction and update to construct an updated network security knowledge graph.
7. The network security attack path deduction method based on multimodal knowledge graph as described in claim 1, characterized in that, Multiple candidate network security attack paths are generated, including: Based on the target assets, attack entry points are identified in the cybersecurity knowledge update graph, and a set of potential attack entry point entities is searched in reverse. Preset attack entry point determination conditions are used to determine and filter the set of potential attack entry point entities to identify a set of feasible attack entry point entities. Based on the set of feasible attack entry points, a dynamic graph search is performed on the network security knowledge update graph to generate multiple candidate network security attack paths.
8. The network security attack path deduction method based on multimodal knowledge graph as described in claim 7, characterized in that, Based on the set of feasible attack entry points, a dynamic graph search is performed on the network security knowledge update graph to generate multiple candidate network security attack paths, including: Construct path edge weight evaluation factors, which include vulnerability exploitability, defense strength, modal credibility, and temporal proximity; Using the target asset as an inspiration, starting from each attack entry entity in the set of feasible attack entry entities, the network security knowledge update graph is evaluated for edge weights and subjected to dynamic heuristic graph search based on the path edge weight evaluation factors to generate multiple candidate network security attack paths.
9. The network security attack path deduction method based on multimodal knowledge graph as described in claim 1, characterized in that, Calculate the overall confidence level of the multiple candidate network security attack paths, including: Obtain the confidence calculation dimensions, which include path integrity, cumulative edge weights, modality coverage, and time decay correction; The confidence scores of the multiple candidate network security attack paths are weighted and calculated based on the aforementioned confidence score calculation dimension to obtain the overall confidence score of the multiple candidate network security attack paths.
10. A network security attack path inference system based on multimodal knowledge graph, characterized in that... The system is used to implement the network security attack path deduction method based on multimodal knowledge graph as described in any one of claims 1 to 9, the system comprising: The data acquisition module is used to collect multimodal network security data, perform knowledge graph modeling and conflict resolution on the multimodal network security data, and generate a multimodal network security knowledge graph. The update module is used to introduce an incremental learning mechanism to fuse and update the multimodal cybersecurity knowledge graph, and construct a cybersecurity knowledge update graph. The search module is used to acquire target assets, and based on the target assets, to perform attack entry point identification and dynamic graph search in the network security knowledge update graph to generate multiple candidate network security attack paths. The calculation module is used to calculate the comprehensive confidence level of the multiple candidate network security attack paths, sort the multiple candidate network security attack paths in descending order according to the comprehensive confidence level, and generate explanatory text to obtain a network security attack path inference list.