A Link Prediction Method for Temporal Network Security Knowledge Graph Based on Polar Coordinate System

By mapping the five-tuple information of the network security knowledge graph to the polar coordinate system and using the polar coordinate system to represent time information, the problems of incomplete knowledge graph and inconsistent time format are solved, and more accurate entity link prediction effect is achieved, and network security protection is enhanced.

CN115456175BActive Publication Date: 2025-06-24HARBIN ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211058684.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-06-24
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of incomplete knowledge graphs, especially in the field of network security. Traditional knowledge graph link prediction methods cannot accurately represent the dependencies between entities and analyze the uncertainty of their dependencies.

Method used

The timing network security knowledge graph link prediction method based on polar coordinate system is adopted. By mapping the five-tuple information into the polar coordinate system, the time information is represented as the scaling and rotation of the entity, and the embedding of different time-constrained entities is distinguished using coefficients and angles.

Benefits of technology

It achieves more accurate entity link prediction effects, solves the problems of inconsistent time formats of knowledge graphs and repeated embedding, and enhances the basic needs of network security protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115456175B_ABST
    Figure CN115456175B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of network security knowledge representation and knowledge graph completion, and particularly relates to a method for predicting links in a temporal network security knowledge graph based on a polar coordinate system. The present invention unifies the time format of the knowledge graph into a start time and an end time, and represents the time change through the scaling and rotation of entities in the polar coordinate system by means of a temporal knowledge graph embedding model, thus solving the problems of inconsistent time formats and repeated embeddings in the knowledge graph; adopts a mapping embedding model in the polar coordinate system, and uses coefficients and angles to distinguish the embeddings of entities with different time constraints, so as to avoid generating similar entities with time constraints in a single dimension embedding, and solve the similarity problem of temporal embedding. The present invention maps the security events into polar coordinate vectors by splitting them into five-tuples, enabling the model to capture the interaction information between entities and relationships more fully, thereby achieving a more accurate entity link prediction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security knowledge representation and knowledge graph completion, and specifically relates to a method for predicting links in a time-series network security knowledge graph based on a polar coordinate system. Background Art

[0002] With the in-depth development of technologies such as semantic search and intelligent question answering, the knowledge graph, as a structured semantic knowledge base, has become an important form of knowledge representation. When constructing a knowledge graph, entities and relationships in unstructured data are first extracted, and then the mutual relationships between entities are stored and represented in a structured manner. Now, over time, large-scale knowledge graphs have been established in various industries. Due to the huge amount of data in the graph, manual verification can no longer guarantee the integrity of the knowledge graph, and link prediction has also become one of the important tasks for knowledge graph completion.

[0003] Link prediction is the task of predicting the missing relationships between entities in an existing knowledge graph, that is, based on the known entities and existing association relationships between entities in the network structure of the knowledge graph, to predict the possibility of a link existing between two entities that are not connected by an edge in the network structure. Link prediction is a promising and widely studied task, aiming to solve the problem of incomplete knowledge graphs. Due to the possible incompleteness of data sources and omissions in the knowledge acquisition process, the authenticity and integrity of existing knowledge graphs still need to be improved. Research shows that 69% - 99% of the entities in most knowledge graphs have missing knowledge, and the missing knowledge is mainly related attribute information triples. The implementation of knowledge graph link prediction technology can provide a basis for restoring missing information and incorrect information, and further provide application support for network applications in many aspects such as intelligent search, intelligent question answering, and personalized recommendation. It is of great significance for knowledge discovery, knowledge navigation, and knowledge fusion in the expert domain knowledge base under big data knowledge. However, in large-scale knowledge graphs, the structure is complex, and the entity association relationships in the open world are constantly changing. Many structured knowledges are only valid within a specific range, resulting in uncertainty in the entity association relationships. Traditional knowledge graph link prediction methods cannot accurately represent the dependence relationships between knowledge graph entities and analyze the uncertainty of their dependence relationships. At the same time, new knowledge needs to be continuously added to the knowledge graph, and this information may come from rapidly generated and evolving data in various fields. Therefore, the construction of the knowledge graph needs to reflect new knowledge in real time. To solve this problem, the time-series knowledge graph model embeds time information into triples, enabling the model to obtain good performance on knowledge graphs involving time-related relationships, which has an important impact on the accuracy of link prediction results.

[0004] In a temporal knowledge graph, time information is considered a fundamental element of entities. Compared with static knowledge graphs, facts in temporal knowledge graphs are generally more realistic expressions of entities in the real world. At the same time, reasoning about entities with time constraints can yield more accurate results, such as ICEWS, GDELT, YAGO3, and Wikidata. These temporal knowledge graphs store a large amount of entity information with timestamps, generally saved in two formats: one is to regard the time condition as the moment of the entity's action; the other is to represent the entity's attributes with a five-tuple, that is, in addition to the traditional triple describing the entity's attributes, it also includes the entity's timestamp information. The transfer distance model can transform the problem of measuring the rationality of triples in the vectorized knowledge graph into the problem of measuring the distance between the head entity and the tail entity. The key to this method is how to design the scoring function, which is often designed as a function that utilizes the relationship to transfer the head entity to the tail entity reasonably. However, in terms of the similarity of the transfer distance model, different entities with timestamps can produce similar embedding results in a single dimension. Most models can only identify entities after mixing time information, but cannot identify entities and time separately. In addition, most models ignore the integrity and relevance of the time set, resulting in the use of a unified time unit for data sets with different time spans.

[0005] Cybersecurity data has strong time-dependent attributes. When constructing a local cybersecurity knowledge graph, it is necessary to model the internal network devices, device attributes, and device vulnerabilities to form a cybersecurity knowledge graph. The data of the cybersecurity knowledge graph can be composed of time-dependent event or device information links. Summary of the Invention

[0006] The purpose of the present invention is to provide a link prediction method for a temporal cybersecurity knowledge graph based on a polar coordinate system that can complete the link prediction task for a certain security event knowledge graph data set with time utility, thereby achieving the purpose of knowledge graph completion and meeting the basic requirements of cybersecurity protection.

[0007] A link prediction method for a temporal cybersecurity knowledge graph based on a polar coordinate system includes the following steps:

[0008] Step 1: Extract complete five-tuple information and entity information from the cybersecurity knowledge graph, and preprocess the obtained data to form a set of five-tuples (h, r, t, [τ s , τ e ) based on events;

[0009] Among them, h and t respectively represent the head entity and the tail entity in the entity set, h, t belong to the entity set E; r ∈ R represents that the relationship between the head entity and the tail entity belongs to the relationship set R; [τ s , τ eRepresents the time span of the entity, τ s Represents the start time, τ e Represents the end time of the fact;

[0010] Step 2: Input the set of five-tuples, and uniformly initialize all entities, relationships, and timestamps in vector form, and form the vector form of the five-tuples (h, r, t, [τ s , τ e );

[0011] Step 3: Map the extracted entity relationships to the polar coordinate system, and divide the training model into a magnitude part and an angular part;

[0012] Define the magnitude part component to be represented in vector form as (h m , r m , t m , [τ s,m , τ e,m ), and add conditional constraints through the mapping function to convert time-independent entities into time-constrained entities, and use the relationship r m as a scaling transformation between two time-constrained entities, and finally obtain a score through the reward function ;

[0013] Define the angular part component to be represented in vector form as (h a , r a , t a , [τ s,a , τ e,a ), and distinguish time-constrained entities with the same modulus through the time constraints h τ,a = (h a + τ s,α ) mod 2π, t τ,a = (t a + τ e,a ) mod 2π, and finally obtain a score through the reward function f τ,a (h τ,a , r a , t τ,a ) = ||sin((h τ,a + r a - t τ,a ) / 2‖1;

[0014] The scoring function is:

[0015]

[0016] where ε is the input five-tuple; is used to adjust the proportion of the embedding of the two part components;

[0017] Step 4: Determine the loss function using the core idea in the knowledge representation learning method, continuously perform gradient updates, and obtain the final relationship information vector;

[0018] Step 4.1: Calculate the loss function:

[0019]

[0020] where, ε is a positive sample, ε′ i is the i-th negative five-tuple; γ represents a fixed margin; σ is the sigmoid function; is the positive sample set; is the negative sample set;

[0021] Step 4.2: Calculate the link prediction probability p(ε′ j |ε′ i );

[0022]

[0023] where, α is temperature sampling;

[0024] Step 4.3: Repeat Steps 3 to 4, continuously update the scoring function of the five-tuples to obtain a probability distribution until all five-tuples in the search task are traversed, and complete the model training;

[0025] Step 5: Put the missing facts (h, r, t 缺失 , [t s , τ e ) into the model, and the model outputs the probabilities of all candidate entities to complete the current fact according to the existing entity set, and selects candidate entities with appropriate probabilities according to different requirements.

[0026] Furthermore, the entity set described in Step 1 includes equipment-type entities and security element-type entities: the equipment-type entities include manufacturing equipment, industrial control systems, detection equipment, test and experimental equipment, and safety protection equipment; the security element-type entities include operating systems, security logs, control equipment models, control management software, network vulnerability attacks, and communication protocol types; the relationship set is divided into connection relationships and event-type relationships, including equipment connection status, triggering backdoors, worm trojans, overflows, brute force attacks and weak passwords, service disruptions, information leakage, and denial of service.

[0027] The beneficial effects of the present invention are as follows:

[0028] The present invention can complete the link prediction task of the cybersecurity knowledge graph based on past cybersecurity events, unify the time format of the knowledge graph into start time and end time, and represent the time change as the scaling and rotation of entities in the polar coordinate system through the temporal knowledge graph embedding model, solving the problems of inconsistent time format and repeated embedding in the knowledge graph; adopt the mapping embedding model in the polar coordinate system, use coefficients and angles to distinguish the embeddings of entities with different time constraints, so as to avoid generating similar time-constrained entities in a single dimension embedding, in order to solve the similarity problem of temporal embedding; by analyzing the impact of different time embeddings on entities and the accuracy of link prediction for different time units, select the smallest time unit of each dataset as the final result of the experimental results. The present invention maps the security events into polar coordinate vectors by splitting them into five-tuples, enabling the model to capture the interaction information between entities and relationships more fully, thereby achieving a more accurate entity link prediction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic flowchart of the present invention.

[0030] Figure 2 It is an instance diagram of the security graph device node adopted in the embodiment of the present invention.

[0031] Figure 3 It is an instance diagram of the security information of the security graph adopted in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] The present invention will be further described below with reference to the accompanying drawings.

[0033] The present invention aims to implement a knowledge graph suitable for cybersecurity protection and solve the problem of generating similar time-constrained entities in a single dimension embedding. In the traditional cybersecurity knowledge graph completion task, entity information can be embedded into the knowledge graph to predict entity link relationships, but the impact of time information in security events on entity relationships is ignored, resulting in the link prediction effect of the knowledge graph involving time information not meeting the basic requirements of cybersecurity protection. The present invention proposes a method for predicting links in a temporal cybersecurity knowledge graph based on the polar coordinate system, which can complete the link prediction task for a certain dataset of security event knowledge graphs with time utility, thereby achieving the purpose of knowledge graph completion and meeting the basic requirements of cybersecurity protection.

[0034] Figure 1 It is a schematic flowchart of a method for predicting links in a temporal knowledge graph based on the polar coordinate system according to the present invention. The specific implementation of this method includes:

[0035] Step 1: Extract complete five-tuple information and entity information from the network security knowledge graph, preprocess the obtained data, and form a set of five-tuples (h, r, t, [τ s , τ e ) based on events, such as (192.168.0.1, Denial of Service attack, 192.168.0.100, 2022-1-1, 2022-1-2), which describes the event that "the device with the internal network IP of 192.168.0.1 launched a Denial of Service attack on the device with the internal network IP of 192.168.0.100 between January 1, 2022 and January 2, 2022".

[0036] Among them, h and t represent the head entity and the tail entity in the entity set respectively. There are h, t belonging to the entity set E, and r ∈ R indicating that the relationship between the head entity and the tail entity belongs to the relationship set R. [τ s , τ e represents the time span of the entity, τ s represents the start time, and τ e represents the end time of the fact.

[0037] The entity set mainly includes device entities and security element entities: Device entities mainly include typical industrial control devices such as manufacturing equipment, industrial control systems, detection devices, test and experimental equipment, and security protection equipment, as well as security elements such as operating systems, security logs, control device models, control management software, network vulnerability attacks, and communication protocol types.

[0038] The relationship set can be divided into connection relationships and event relationships: mainly including device connection status, triggering backdoors, worm trojans, overflows, brute force attacks and weak passwords, service interference, information leakage, and denial of service. As Figure 2 , Figure 3 shown, using the IP address as the device node and adding the basic device information and security information as sub-nodes to form a device security graph.

[0039] Step 2: Input the set of five-tuples, and uniformly initialize all entities, relationships, and timestamps in vector form, and form the vector form of the five-tuples (h, r, t, [τ s , τ e ).

[0040] Step 3: Map the extracted entity relationships to the polar coordinate system, and divide the training model into a modulus part and an angular part.

[0041] Define the component of the modulus part to be represented in vector form as (h m , r m , t m , [τ s,m , τ e,m),Through the mapping function Add conditional constraints to convert time-independent entities into time-constrained entities, and regard the relationship r m As a scaling transformation between two time-constrained entities, and finally obtain a score through the reward function .

[0042] Define the angular part component to be represented vectorially as (h a , r a , t a , [τ s,a , τ e,a ), distinguish time-constrained entities with the same modulus through the time constraint h τ,a =(h a +τ s,a ) mod 2π, t τ,a =(τ a +τ e,a ) mod 2π, and finally obtain a score through the reward function f τ,a (h τ,a , r a , t τ,a ) = ||sin((h τ,a +r a -t τ,a ) / 2‖1.

[0043] The scoring function is composed of the modulus part and the angular part, and the formula is where ε is the input five-tuple, used to adjust the embedding ratio of the two part components.

[0044] Step 4: Use the core idea in the knowledge representation learning method to determine the loss function, continuously perform gradient updates, and obtain the final relationship information vector.

[0045] Step 4.1: Define the loss function according to the scoring function f τ (ε) in Step 3:

[0046]

[0047] where, ε is the positive sample, ε′ i is the i-th negative five-tuple, γ represents the fixed margin, σ is the sigmoid function, is the positive sample set, is the negative sample set.

[0048] Step 4.2: Calculate to obtain the link prediction probability, where α is the temperature sampling.

[0049] Step 4.3: Repeat Step 3 to Step 4 to continuously update the scoring function of the five-tuples and obtain the probability distribution until all the five-tuples in the search task are traversed, completing the model training.

[0050] Step 5: Put the missing fact (h, r,?, [τ s , τ e ) example (192.168.0.1, Denial of Service attack,?, 2022-1-1, 2022-1-2) into the model, and the model will output the probabilities of all candidate entities to complete the current fact according to the existing entity set, and select the candidate entities with appropriate probabilities according to different requirements.

[0051] Compared with the existing technologies, the beneficial effects of the present invention are as follows: It can complete the link prediction task of the cybersecurity knowledge graph according to past cybersecurity events, unify the time format of the knowledge graph into start time and end time, represent the time change as the scaling and rotation of entities in the polar coordinate system through the temporal knowledge graph embedding model, and solve the problems of inconsistent time formats and repeated embeddings in the knowledge graph; adopt the mapping embedding model in the polar coordinate system, use coefficients and angles to distinguish the embeddings of different time-constrained entities to avoid generating similar time-constrained entities in a single dimension embedding, so as to solve the similarity problem of temporal embedding; by analyzing the influence of different time embeddings on entities and the accuracy of link prediction for different time units, select the minimum time unit of each data set as the final result of the experimental result. By splitting security events into five-tuples and mapping them into polar coordinate vectors, the present invention enables the model to capture the interaction information between entities and relationships more fully, thereby achieving a more accurate entity link prediction effect.

[0052] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting links in a temporal network security knowledge graph based on polar coordinates, characterized in that, It includes the following steps: Step 1: Extract complete five-tuple information and entity information from the network security knowledge graph, preprocess the obtained data, and form a set of event-based five-tuples (h, r, t, [τ s , τ e ) set; Among them, h and t respectively represent the head entity and the tail entity in the entity set, where h, t ∈ E; r ∈ R indicates that the relationship between the head entity and the tail entity belongs to the relationship set R; [τ s , τ e represents the time span of the entity, τ s represents the start time, and τ e represents the end time of the fact; The entity set includes device entities and security element entities: The device entities include manufacturing equipment, industrial control systems, detection equipment, test and experimental equipment, and safety protection equipment; The security element entities include operating systems, security logs, control device models, control management software, network vulnerability attacks, and communication protocol types; The relationship set is divided into connection relationships and event relationships, including device connection status, triggered backdoors, worm trojans, overflows, brute force attacks and weak passwords, service disruptions, information leakage, and denial of service; Step 2: Input the set of five-tuples, and uniformly initialize all entities, relationships, and timestamps in vector form, and form the vector form of the five-tuples (h, r, t, [τ s , τ e ) Step 3: Map the extracted entity relationships to the polar coordinate system and divide the training model into a modulus part and an angular part; Define the partial components of the modulus length to be represented vectorially as (h m , r m , t m , [τ s,m , τ e,m ). By means of the mapping function Add conditional constraints to convert time-independent entities into time-constrained entities, and regard the relationship r m as a scaling transformation between two time-constrained entities. Finally, obtain a score through the reward function ; Define the angular part component to be represented vectorially as (h a , r a , t a , [τ s,a , τ e,a ). Distinguish time - constrained entities with the same modulus through the time constraints h τ,a = (h a + τ s,a ) mod 2π, t τ,a = (t a + T e,a ) mod 2π. Finally, obtain the score through the reward function f τ,a (h τ,a , r a , t τ,a ) = ‖sin((h τ,a + r a - t τ,a ) / 2||1; The scoring function is: where ε is the input five-tuple; used to adjust the ratio of the embedding of the two partial components; Step 4: Use the core idea in the knowledge representation learning method to determine the loss function, continuously update the gradient, and obtain the final relationship information vector; Step 4.1: Calculate the loss function: Among them, ε is a positive sample, and ε' i is the i-th negative five-tuple; γ represents a fixed boundary; σ is the sigmoid function; is the positive sample set; is the negative sample set; Step 4.2: Calculate the link prediction probability p(ε′ j |ε′ i ); where α is the temperature sampling; Step 4.3: Repeat steps 3 to 4, continuously update the scoring function of the quintuple, obtain the probability distribution, until all quintuples in this search task are traversed, and complete the model training; Step 5: Put the missing facts (h, r, t 缺失 , [τ s , τ e ) into the model, and the model outputs the probabilities of all candidate entities to complete the current fact based on the existing entity set, and selects candidate entities with appropriate probabilities according to different requirements.

Citation Information

Patent Citations

  • Water conservancy literature recommendation method and system based on automatic completion of knowledge graph

    CN113239210A

  • Time sequence knowledge graph completion method and system based on time graph convolutional network

    CN114780739A