Host side attack detection model construction method and system based on attack subgraph
By building a host-side attack detection model based on attack subgraph, using reverse traceability algorithm and representation learning technology, high-order attack logic information is extracted and fused, the problem of low detection accuracy in the existing technology is solved, and more efficient attack detection is achieved.
Patent Information
- Application Number
- CN202510523764.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
AI Technical Summary
The existing host-side attack detection methods cannot fully mine the high-order attack logic features in the system log, resulting in low detection accuracy and difficulty in effectively detecting attack behavior in encrypted traffic.
A host-side attack detection model based on attack subgraph is constructed, a network security behavior knowledge graph is constructed by obtaining system logs, a reverse traceability algorithm based on similarity is used to extract the attack subgraph, and a DeepWalk and GCN models are used for representation learning, and the attack subgraph is fused with the knowledge graph to perform attack type detection.
It improves the accuracy of host-side attack detection, can capture key attack information, narrow the attack range, and reduce the impact of normal behavior on detection results.
Smart Images

Figure CN120342708A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of host intrusion detection, and more specifically, relates to a method and system for constructing a host-side attack detection model based on attack subgraphs. Background Art
[0002] As a device that individuals and organizations generally own and use in the information age, hosts may store everything from personal privacy to national secrets. Therefore, hosts have also become important targets for network attacks. To defend against external attacks, enterprises and operators will deploy a series of security defense devices at the network boundary, such as firewalls, intrusion prevention systems (IPS), etc. However, some attacks can still bypass the firewall and intrude into the internal network. After the attacker enters the internal network, the communication network transmission usually uses traffic encryption technology, and traditional network devices also have certain difficulties in processing encrypted traffic, making it difficult to detect and prevent attackers on the network side after the attacker enters the internal network. The operating system will record various activities that occur on the host in the form of log files. After the attacker penetrates the network protection and invades the personal host, behavioral traces will be left on the host and recorded in the host system log, which makes it possible to detect threats on the host side.
[0003] The host system log records the behaviors at the operating system kernel level such as files, registries, processes, sockets, etc., which can be used to detect harmful activities such as worms, ransoms, and mining. However, the system log is presented in text form, which has problems such as being difficult for computers to understand and excessive storage consumption. At the same time, existing detection methods cannot fully mine the key information in the log, resulting in low efficiency and poor accuracy in host-side attack investigations based on system logs.
[0004] Therefore, how to improve the accuracy of host-side attack detection is an urgent problem to be solved. Summary of the Invention
[0005] Aiming at the defects of the existing technology, the purpose of this application is to provide a method and system for constructing a host-side attack detection model based on attack subgraphs, which can effectively improve the accuracy of host-side attack detection.
[0006] To achieve the above purpose, in the first aspect, this application provides a method for constructing a host-side attack detection model based on attack subgraphs, including the following steps:
[0007] S10, obtain and construct a network security behavior knowledge graph according to the system log;
[0008] S20, given a malicious event, and use a similarity-based reverse tracing algorithm to conduct an attack investigation on the malicious event, and extract an attack subgraph containing high-order attack logic;
[0009] S30. Embed and represent the network security behavior knowledge graph through representation learning technology;
[0010] S40. Integrate the embedded and represented knowledge graph with the attack subgraph, add high-order attack logic information to the embedded and represented knowledge graph, and then perform attack type detection.
[0011] Advantages of this application: The method for constructing a host-side attack detection model based on attack subgraphs provided by this application extracts attack subgraphs containing high-order attack logic by conducting attack investigations on malicious alert events detected by the host system. Then, it embeds and represents the network security behavior knowledge graph through representation learning technology, and integrates the complete graph representation with the attack subgraph representation. This enables the model to capture key attack information related to attack behaviors, narrow the attack scope, reduce the impact of normal behaviors on attack prediction results, and effectively improve the accuracy of the host-side attack detection model.
[0012] As a further preference, step S10 is specifically:
[0013] Graph structure modeling of log data, extract system entities, events, and their attribute information from system logs, establish connections between entities through events, and organize the activities on the host represented by the logs in the form of a graph;
[0014] Entity and event attribute screening, construct an attribute set, and determine the attributes most relevant to attack investigation and attack detection;
[0015] Integrate graph structure modeling and attribute screening to construct a network security behavior knowledge graph with complete semantics.
[0016] As a further preference, step S20 is specifically:
[0017] Given malicious events detected by security software;
[0018] According to the malicious events, perform backward tracking through malicious similarity features and dependency weights to find the starting point on which the event depends;
[0019] Obtain a forward subgraph through forward traversal from this starting point, and at the same time obtain a backward subgraph through backward traversal from the malicious event;
[0020] Union the forward subgraph and the backward subgraph to obtain the attack subgraph where the malicious event is located, thereby extracting high-order attack logic information of malicious attacks.
[0021] As a further preference, the malicious similarity features include time similarity, node degree similarity, and data volume similarity.
[0022] As a further preference, step S30 is specifically:
[0023] Apply DeepWalk random walk to learn the local sequence semantic information of entities and relationships in the knowledge graph, and obtain the vector representations of entities and relationships containing local sequence semantic information;
[0024] Use the sequence semantic vectors learned by DeepWalk as the initial input vectors of the GCN model, and utilize the GCN model to capture the adjacency structure information between entities and relationships, generating the vector representations of entities and relationships that contain both sequence semantic and adjacency structure information.
[0025] As a further preference, step S40 is specifically as follows:
[0026] Perform separate pooling operations on the embedded knowledge graph and the attack subgraph to obtain the complete graph representation vector and the attack logic subgraph vector with attack information, and perform weighted fusion on the two to obtain the knowledge graph feature vector of network security behavior for attack detection;
[0027] Detect the attack type by training a multi-layer perceptron classifier to obtain the prediction probabilities of each attack type.
[0028] As a further preference, in step S40, the attack types include malicious programs, mining, Trojans, malware, extortion, system viruses, worms, and backdoors.
[0029] As a further preference, after step S40, it further includes:
[0030] Guide the model to be trained through multi-class cross-entropy loss and contrastive learning InfoNCE loss.
[0031] As a further preference, a cost-sensitive learning strategy is added to the cross-entropy loss.
[0032] In a second aspect, the present application provides a host-side attack detection model construction system based on an attack subgraph, including:
[0033] A knowledge graph construction module, used to obtain and construct a network security behavior knowledge graph according to system logs;
[0034] An attack subgraph extraction module, used to given a malicious event, and adopt a similarity-based reverse tracing algorithm to conduct an attack investigation on the malicious event, and extract an attack subgraph containing high-order attack logic;
[0035] An embedding representation module, used to perform embedding representation on the network security behavior knowledge graph through representation learning techniques;
[0036] An attack detection module is used to fuse the knowledge graph with embedded representations and the attack subgraph, so that high-order attack logic information is added to the knowledge graph with embedded representations, and then attack type detection is performed.
[0037] It can be understood that the beneficial effects of the second aspect above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of the Drawings
[0038] Figure 1 is a flowchart of a method for constructing a host-side attack detection model based on an attack subgraph provided by an embodiment of the present application;
[0039] Figure 2 is a schematic diagram of reverse traceability provided by a specific embodiment of the present application;
[0040] Figure 3 is an overall flowchart of a method for constructing a host-side attack detection model based on an attack subgraph provided by a specific embodiment of the present application;
[0041] Figure 4 is a LDGCN framework diagram provided by a specific embodiment of the present application. Detailed Description of the Embodiments
[0042] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] It should be understood that in the description of the present application, the meaning of the term "several" is at least one, for example, one, two, etc., unless otherwise specifically defined; the meaning of the term "multiple" is two or more, unless otherwise specifically defined; the terms "first" and "second" etc. are used to distinguish different objects, rather than to describe a specific order of the objects; the term "and / or" includes any and all combinations of one or more of the related listed items.
[0044] In addition, the reference to "an embodiment" throughout this specification; the appearance of "in an embodiment", "an example" or similar language means that the specific features, structures or characteristics described in connection with that embodiment are included in at least one embodiment of the present application. Therefore, the phrases "in an embodiment;" and "in an embodiment" and similar language that appear throughout this specification may or may not all refer to the same embodiment.
[0045] Through research, it is found in this application that traditional attack type detection methods ignore high-order attack logic features such as "tactics, techniques, and procedures" that are most directly related to attack behaviors. High-order attack logic is the key logical step for successful execution of attacks and is the information in the host that is most relevant to attack behaviors. There is a problem of insufficient utilization of malicious information in the host.
[0046] In addition, most traditional attack type detection methods equally treat various system activities and entities on the host, and all activities on the host are used in attack detection. This processing method not only increases the computational complexity but also loses key information due to insufficient attention to malicious events, reducing the detection accuracy. Attack behaviors on the host are generally hidden in normal behaviors and wait for opportunities. In addition, to reduce the risk of being discovered, it will minimize its own activity range and minimize its association with other entities and activities in the system. Therefore, the activities that actually execute malicious behaviors only account for a very small part in the host.
[0047] Aiming at the problem of low attack detection accuracy caused by insufficient utilization of key attack information in traditional attack detection, this application designs an attack detection model that integrates high-order attack logic to improve the attack detection effect. Specifically, this application conducts an attack investigation on malicious alert events detected by the host system, traces the origin of the attack, extracts attack subgraphs containing high-order attack logic, locates key attack information, and designs a knowledge graph representation learning and attack detection model for attack prediction. In the representation learning stage, LDGCN (Logical Random Walk Graph Convolutional Network) first learns the sequential semantic information of entities and relationships in the knowledge graph through DeepWalk, and then uses the sequential semantics learned by DeepWalk as the initial vector representation of the GCN model (Graph Convolutional Network). The GCN learns the structural features of entities and relationships. This application reconstructs the adjacency matrix of the knowledge graph by treating relationships as nodes and establishing edges between the connected entities, so that the GCN can capture the adjacency information of relationships during learning; in the attack detection stage, pooling is performed on the complete network security behavior knowledge graph and the attack subgraph respectively, and the pooled vector representations are fused, explicitly adding key attack information to the global knowledge graph representation, and finally attack prediction is performed through a classifier. In addition, the research content of this application also reduces the impact of the sample imbalance problem through a cost-sensitive learning strategy and enhances the generalization ability of the model.
[0048] As Figure 1 shown, the method for constructing a host-side attack detection model based on attack subgraphs provided by this application includes steps S10 to S40, which are described in detail as follows:
[0049] Step S10, obtain and construct a network security behavior knowledge graph according to system logs.
[0050] Considering that knowledge graphs model entities and events in the form of graph structures and can effectively capture the associated information between entities. Knowledge graphs describe various facts in the world in triples, and their structured knowledge representation is also closer to the form in which humans learn and remember knowledge. Drawing on the concept of knowledge graphs, this application models the system entities on the host and the system activities between them as a cybersecurity behavior knowledge graph, which is stored in the form of a graph structure. This can not only remove duplicate texts, reduce data redundancy, and lower the storage consumption brought by log data, but also effectively and intuitively reveal the system activities on the host.
[0051] In this embodiment, step S10 can specifically be:
[0052] Step S11, graph structure modeling of log data. Extract system entities, events, and their attribute information from system logs, establish connections between entities through events, and organize the activities on the host represented by the logs in the form of a graph.
[0053] Step S12, entity and event attribute screening. Construct an attribute set and determine the attributes most relevant to attack investigation and attack detection.
[0054] Step S13, fuse the graph structure attribute information to construct a cybersecurity behavior knowledge graph with complete semantics.
[0055] Step S20, given a malicious event, and use a similarity-based reverse tracing algorithm to conduct an attack investigation on this malicious event, and extract an attack subgraph containing high-order attack logic.
[0056] In this embodiment, step S20 can specifically be:
[0057] Step S21, given a malicious event detected by security software.
[0058] Step S22, according to the malicious event, conduct reverse tracking through malicious similarity features and dependency weights to find the starting point on which this event depends.
[0059] Step S23, conduct forward traversal through this starting point to obtain a forward subgraph, and at the same time conduct backward traversal through the malicious event to obtain a backward subgraph.
[0060] Step S24, take the union of the forward subgraph and the backward subgraph to obtain the attack subgraph in which this malicious event is located, so as to extract the high-order attack logic information of this malicious attack.
[0061] Step S30, perform embedding representation on the cybersecurity behavior knowledge graph through representation learning techniques.
[0062] In this embodiment, step S30 can specifically be:
[0063] Step S31: Apply DeepWalk random walk to learn the local sequence semantic information of entities and relationships in the knowledge graph, and obtain the vector representations of entities and relationships containing local sequence semantic information.
[0064] Step S32: Use the sequence semantic vectors learned by DeepWalk as the initial input vectors of the GCN model, and utilize the GCN model to capture the adjacency structure information between entities and relationships, generating entity and relationship vectors that contain both sequence semantics and adjacency structure information.
[0065] Step S40: Integrate the embedded knowledge graph with the attack subgraph, add high-order attack logic information to the embedded knowledge graph, and then perform attack type detection.
[0066] In this embodiment, Step S40 can specifically be:
[0067] Step S41: Perform separate pooling operations on the embedded knowledge graph and the attack subgraph to obtain the complete graph representation vector and the attack logic subgraph vector with attack information, and perform weighted fusion on the two to obtain the network security behavior knowledge graph feature vector for attack detection.
[0068] Step S42: Detect the attack type by training a multi-layer perceptron classifier to obtain the prediction probabilities of each attack type.
[0069] In Step S40, the attack types include virus attacks such as malicious programs, mining, Trojans, malware, ransoms, system viruses, worms, and backdoors.
[0070] The method for constructing a host-side attack detection model based on an attack subgraph provided in this embodiment extracts an attack subgraph containing high-order attack logic by conducting an attack investigation on malicious alert events detected by the host system, and then performs an embedded representation on the network security behavior knowledge graph through representation learning technology, and integrates the complete graph representation with the attack subgraph representation, enabling the model to capture key attack information related to attack behaviors, narrow the attack scope, reduce the impact of normal behaviors on the attack prediction results, and effectively improve the accuracy of the attack detection model.
[0071] Based on the same inventive concept, the present application also provides a system for constructing a host-side attack detection model based on an attack subgraph, including a knowledge graph construction module, an attack subgraph extraction module, an embedded representation module, and an attack detection module.
[0072] Among them, the knowledge graph construction module is used to obtain and construct a network security behavior knowledge graph based on system logs.
[0073] The attack subgraph extraction module is used to given a malicious event, and adopt a similarity-based reverse tracing algorithm to conduct an attack investigation on the malicious event, and extract an attack subgraph containing high-order attack logic.
[0074] The embedding representation module is used to perform an embedding representation on the network security behavior knowledge graph through representation learning techniques.
[0075] The attack detection module is used to fuse the embedded knowledge graph with the attack subgraph, so that the high-order attack logic information is added to the embedded knowledge graph, and then the attack type detection is performed.
[0076] Specifically, the functions of the modules provided in this embodiment can refer to the detailed introduction in the foregoing method embodiment, and will not be elaborated here.
[0077] Next, a specific embodiment is used to elaborate in detail on the method for constructing a host-side attack detection model based on attack subgraphs provided in this application.
[0078] System audit logs collect system-level audit events from the operating system kernel and provide interaction events between different system entities. Graph data has shown good application prospects in describing the connections between all things in the world. Knowledge graphs and related technologies and applications have developed rapidly, and graph neural networks have been widely used in fields such as social network analysis, recommendation systems, and subgraph classification.
[0079] Modeling the text-formatted log data according to the graph structure to construct a network security behavior knowledge graph, using nodes to represent system entities and edges to represent system events, can clearly and intuitively describe the entities and system activities on the host. At the same time, taking the attack type detection task as a graph classification task can apply the technology of graph neural networks to attack type detection, providing a new technical method for attack investigation on the host side.
[0080] Aiming at the problem that traditional attack detection methods do not make full use of the key attack information in the network security behavior knowledge graph, this application extracts attack subgraphs according to malicious events, finds high-order attack logic related to attack behaviors, narrows the attack scope, and reduces the impact of normal behaviors on the attack prediction results, so as to improve the accuracy of the attack detection model.
[0081] Specifically, the method for constructing a host-side attack detection model based on attack subgraphs provided in this application includes the following steps:
[0082] Step (1): Construction of the network security behavior knowledge graph.
[0083] This step introduces the construction of the network security behavior knowledge graph, preparing data for subsequent research. The specific research content is as follows: 1. Graph structure modeling of log data, extracting system entities, events, and their attribute information from system logs, establishing connections between entities through events, and organizing the activities on the host represented by the logs in the form of a graph; 2. Entity and event attribute screening, constructing an attribute set, determining the attributes most relevant to attack investigation and attack detection, and extracting attribute information according to this standard subsequently to unify log information from different sources; 3. Construction of the network security behavior knowledge graph, using triples as the basic unit, and constructing the final network security behavior knowledge graph with the entities, events, and attributes determined in the previous two steps.
[0084] Graph structure modeling: This step completes the graph structure construction of the network security behavior knowledge graph. The main work involved is to extract entities, relationships, and triples from log texts. System logs consist of individual system activities and their descriptions, including event types, event subjects, event objects, event times, system users, current directories, event levels, process IDs, etc. For each system activity, three parts can be extracted: the event subject and object, the specific event description information, and the process ID that generated the activity. In the graph structure, two nodes and one edge can be extracted, that is, two entities and one relationship. It should be noted that in addition to the event subject and object, each system activity also includes a process entity because any system activity needs to be executed through a process. In the finally modeled network security behavior knowledge graph, there will be a process entity acting as a transit between activity entities, and activity entities refer to the subjects or objects of events such as files or Sockets. The log data used in this method is XML or Json formatted logs, so the extraction of entities and events can be completed through XML or Json parsing.
[0085] Entity and relationship screening: System logs contain attribute information of entities and relationships, which describe the characteristics of entities and relationships. However, some attributes are irrelevant to attack detection and network defense, and the attribute information contained in logs from different sources is inconsistent. Therefore, we need to screen the entity and relationship attributes, remove redundant attributes, and retain the attributes related to attack detection to construct a network security behavior knowledge graph with a unified format and concise attributes.
[0086] To more accurately describe entity types, the content of this study further classifies entity types. Entities are divided into three categories: files, processes, and IPs. The reason for this is that different entity types have different attribute information. For example, files have path attributes, while IPs have process ID attributes, and the event types associated with them are also different. Events related to file entities are file creation, registry modification, etc., and events related to IP entities may be Socket connections, etc. Further subdividing entity types can better classify and manage attribute and event information.
[0087] According to the extracted attribute information, the extracted entities and relationship attributes are divided into two categories. One category is numerical attributes that can be used for calculation, including the in-degree and out-degree of nodes, the start time and end time of events, which are important information related to attack investigation; the other category is semantic attributes that represent relationship semantics. These attributes can make the constructed knowledge graph more readable and help security personnel and computers understand the content of the knowledge graph for attack investigation and defense. The specific attribute information is shown in Tables 1 and 2.
[0088] For entity attributes, the file entity is identified by the absolute path and file name, the process entity is identified by the pid, and the IP entity is identified by the IP address. In addition, each entity contains an Event set that includes events related to the entity, and there is also an EntityType attribute that represents the type of the entity, as shown in Table 1.
[0089] Table 1 Representative attributes of system entities
[0090]
[0091] For relationship attributes, Operation represents the specific event of the relationship. The event can be represented as an edge in the graph and the two nodes it connects. There are additional attributes on this edge to describe the event. It includes the start time, end time, head, and tail entity types, as shown in Table 2.
[0092] Table 2 Representative attributes of relationships
[0093]
[0094] Among them, the bold attributes in the table are numerical attributes used for attack tracing, and the other attributes are semantic attributes of entities and events used to describe entities and events, making the constructed network security behavior knowledge graph and the detected attack sequences more readable, and helping analysts reasonably infer the attack sequences semantically to prevent attack behaviors.
[0095] Structure and Semantic Fusion: The complete cybersecurity behavior knowledge graph contains the attribute semantic information of entities and relationships. After the graph structure modeling and attribute screening in the previous two steps, the graph structure and related semantic information of the knowledge graph have been obtained. Next, the two need to be fused to endow the nodes and edges with relevant attribute information and construct a cybersecurity behavior knowledge graph with complete semantics. The specific operation is to add the attribute information of entities and relationships to the graph on the basis of the unchanged graph structure. The overall process is shown in Algorithm 1.
[0096]
[0097] Step (2): Attack Subgraph Extraction.
[0098] For attacks of the same type, the specific behavior sets generated on different hosts are different, but the "tactics, techniques, procedures" required for successful attack execution are directly related to their ultimate attack goals, and their potential attack logics are the same. For example, the attack logic of a certain virus is "move files → hide → execute scripts → steal data → send". These key logical steps are necessary for the attacker to complete the attack and are independent of the host state and will be executed on any host. This part of the key logical steps is called the high-order attack logic, which exists in the form of a subgraph in the graph and is the context environment and key attack features highly related to the attack behavior. Utilizing these features can provide key information related to the attack behavior for attack detection.
[0099] The extraction of high-order attack logic needs to take malicious events as the entry point. After constructing the cybersecurity behavior knowledge graph, this step uses a similarity-based reverse traceability algorithm to conduct attack investigations on malicious events and restore the attack path and attack subgraph. Given a malicious POI event e s , first, according to the malicious POI event, perform reverse tracking through malicious similarity features and dependency weights to find the starting point on which the event depends. Then, perform forward traversal through the starting point to obtain a forward subgraph, and at the same time, perform backward traversal through the POI malicious event to obtain a backward subgraph. Consider the union of the forward subgraph and the backward subgraph as the context environment involved in this malicious event, that is, the attack subgraph in which this malicious event is located, so as to extract the high-order attack logic information of this malicious attack, as Figure 2 shown.
[0100] The events and nodes in the subgraph are highly related to the malicious POI event e s , and it can be considered that the nodes and events in the subgraph jointly execute this attack behavior, as Figure 2 shown.
[0101] It is found through research that the smaller the time difference between an event and a malicious event, the higher the similarity between it and the malicious event, and the more likely it is to be a malicious event. The closer the data volume transmitted by it is to that of the malicious POI event, the more likely it is to be a malicious event. Malicious entities generally have the following characteristics: very small in-degree and very large out-degree, that is, fewer incoming edges and more outgoing edges. Fewer incoming edges mean that the entity only has passive interactions with the previous one or several entities in the attack path and has no interaction with other entities in the system. This feature can reduce the risk of itself being discovered. Multiple outgoing edges mean that the entity has active interactions with multiple entities in the system, generating multiple system behaviors, but only one behavior is malicious. This feature can reduce the risk of its malicious behavior being discovered. Therefore, this application uses time similarity, data volume similarity, and node degree similarity to calculate the similarity between all events in the graph and malicious events. The calculation formula is as follows:
[0102] Time similarity:
[0103]
[0104] In the formula, t e and represent the timestamp values of event e and the malicious POI event. The smaller the difference, the higher the time correlation. f T(e) =ln(1 + 1e10) represents the time correlation of the malicious POI event, which has the highest time correlation.
[0105] At the same time, for an event e(u, v), where u is the source node and v is the aggregation node, the calculation formula for node degree similarity is as follows:
[0106] Node degree similarity:
[0107]
[0108] In the formula, OutDegree(v) and InDegree(v) represent the out-degree and in-degree of the node. OutDegree(v) and InDegree(v) represent the out-degree and in-degree of the aggregation node v in a certain edge e(u, v).
[0109] Data volume similarity:
[0110]
[0111] In the formula, s e and represent the data volumes transmitted by event e and the malicious event POI, represents the similarity between the two, The smaller the difference, the greater the correlation of the data volume. α is a small positive number representing the data volume similarity of POI events, such as e -4 .
[0112] After that, the discriminant feature projection scheme based on Linear Discriminant Analysis (LDA) is used to calculate the feature weights and perform feature aggregation. Linear Discriminant Analysis is a supervised dimensionality reduction technique that can help identify which features are most useful for classification, thus enabling effective feature selection, while reducing the number of features and retaining the information that can best distinguish different classes.
[0113] In this embodiment, aggregating the weights of time, node degree, and data volume correlation is regarded as a dimensionality reduction operation, converting the original features of the three dimensions of time and node degree into features of one dimension. By finding the optimal linear decision boundary to maximize the difference between critical edges and non-critical edges, where critical edges are the edges that have performed malicious behaviors and non-critical edges are the edges that have performed normal activities. Since LDA is supervised, before performing weight aggregation, the KMeans++ algorithm is first used to cluster the edges in the dependency graph, generating an initial label {0, 1} for each edge, where 0 represents an edge that has performed normal activities and 1 represents an edge that has performed malicious activities. KMeans++ is an improvement of the unsupervised clustering algorithm Kmeans. Kmeans has randomness in the selection of the initial clustering center, may fall into local optimality, and has unstable effects. KMeans++ has optimized the selection of the initial clustering center, with a faster convergence speed and fewer iteration times. The Kmeans++ process is shown in Algorithm 2.
[0114]
[0115] After obtaining the initial labels, the weights of the two malicious similarity features are obtained through LDA, and weighted summation is performed to synthesize the malicious similarity information in the three aspects of time, node degree, and data volume. The feature aggregation formula is as follows:
[0116]
[0117] In the formula, and are the aggregation weights of the three malicious similarity features obtained through LDA.
[0118] After aggregating the similarity features of the three dimensions, for the convenience of calculation and analysis, the weights of the edges are normalized so that the out-edge weights of each node are between [0, 1], and the sum of the weights of all out-edges is equal to 1. The calculation formula for weight normalization is as follows:
[0119]
[0120] In the formula, outgoingEdge(u) represents the set of all outgoing edges of node u.
[0121] Dependency Propagation: The purpose of this step is to obtain the starting point of malicious POI events through dependency propagation. The malicious similarity feature of the edge is assigned to the node through iterative operations and defined as the dependency propagation weight, which represents the degree of dependence of the malicious POI event on the entity. In this embodiment, the dependency propagation weight of the node in the malicious POI event is set to 1.0, and then the reverse propagation of the node dependency propagation weight is performed through the malicious similarity feature of the edge to update the dependency propagation weights of other nodes. The process of dependency reverse propagation is as follows: for a node u, the malicious similarity features of the child node and the edge are assigned to u by taking the weighted sum of the dependency propagation weight of its child node and the edge weight. The dependency reverse propagation formula is as follows:
[0122]
[0123] In the formula, DI u represents the dependency propagation weight of node u, W e(u,v) represents the normalized malicious feature similarity of edge e(u,v), outgoingNode(u) represents the set of outgoing edge nodes of node u, and DI v represents the dependency propagation weight of the outgoing edge nodes of v. Through this propagation method, it can be ensured that the weight of any node does not exceed the maximum weight of its child nodes, and at the same time, the weight of any node does not exceed the weight of the nodes in the POI event.
[0124] Attack Subgraph Extraction: After the dependency propagation is completed, the entry nodes are sorted according to the node dependency propagation weights, and the entry node with the highest ranking is selected as the starting point of the attack behavior. The entry node is the node without incoming edges, and this node can be regarded as the entry point leading to the occurrence of malicious POI events. Then, the attack subgraph containing high-order attack logic is extracted by generating the forward subgraph and the backward subgraph. The forward subgraph is obtained by forward traversing along the outgoing edges from the selected entry node; the backward subgraph is obtained by backward traversing along the incoming edges from the malicious POI event. Then, the two subgraphs are merged to obtain the union of the two subgraphs, and it is output as the attack subgraph related to the malicious POI event. For malicious attacks, the completion of the attack target depends on the host behavior on this attack subgraph. The attack subgraph represents the high-order attack logic of this attack and is the key attack information in the network security behavior knowledge graph. In the following work, the attack detection method LDGCN that fuses high-order attack logic is used to perform attack detection on the network security behavior knowledge graph.
[0125] Step (3): Graph Embedding Framework.
[0126] The attack type detection method integrating high-order attack logic proposed in this embodiment is as follows Figure 3 as shown. In terms of the representation learning of the knowledge graph, considering that the attack behaviors and normal behaviors on the host can be regarded as event chains or subgraphs with order, and DeepWalk has good local sequence information capture ability, this embodiment first applies DeepWalk random walk to learn the local sequence semantic information between entities and relationships, and obtains the entity and relationship vector representations containing local sequence semantic information. The DeepWalk model can capture the dynamic order in which entities and relationships occur in the network security behavior knowledge graph. Further, there is not only an order semantic association between entities and relationships, but also adjacent structural information. The graph convolutional neural network GCN can accurately capture the structural information and can process labeled data. This study uses the sequence semantic vectors learned by DeepWalk as the initial input vectors of the GCN model, and uses GCN to capture the adjacent structural information between entities and relationships, and generates entity and relationship vectors that contain both sequence semantics and adjacent structural information. At the attack type detection level, in order to make full use of the high-order attack logic information in the network security behavior knowledge graph, this study performs separate pooling operations on the complete knowledge graph and the attack subgraph to obtain the complete graph representation vector and the attack logic subgraph vector with attack information, and weights and fuses the two to obtain the network security behavior knowledge graph feature vector for attack detection. This method can not only extract complete host behavior information, but also explicitly add attack logic information to the network security behavior knowledge graph feature vector to prevent the attack information from being submerged, thereby improving the model's ability to detect attack types. Finally, a multi-layer perceptron classifier is used to perform attack detection on the network security behavior knowledge graph feature vector.
[0127] Further, before classification by the multi-layer perceptron classifier, LDGCN weights and fuses the two to obtain the final vector representation for attack detection, and finally predicts the attack type through the classifier. The attack types here include virus attacks such as malicious programs, mining, Trojans, malware, ransoms, system viruses, worms, backdoors, etc. At the same time, this embodiment also uses a cost-sensitive learning strategy to reduce the impact brought by the problem of sample imbalance. Among them, the application and embedded representation of the attack subgraph are the key points of this embodiment.
[0128] Step (4): Global information and attack subgraph information encoding.
[0129] This step introduces the knowledge graph representation learning and the attack detection model LDGCN of this method. The main purpose of this part is to generate high-quality entity and relationship vectors.
[0130] Graph Representation Learning: Incorporating subgraph information containing high-order attack logic into graph representation learning can generate high-quality graph embedding vectors. At the same time, the path information in the graph also represents the sequence of system events and contains certain sequential semantics. Therefore, in this section, an attack type detection method LDGCN that considers the sequence and structural semantics between attack information entities and relationships is designed for the representation learning and attack detection of network security knowledge graphs.
[0131] LDGCN first uses the Deepwalk algorithm to capture the sequential semantic information between nodes. The algorithm process of DeepWalk is shown in Algorithm 4. DeepWalk learns the vector representations of entities and relationships based on the co-occurrence relationship between entities and relationships in the graph. It performs sequential sampling through random walks, starting from each entity and conducting multiple random walks, with each walk generating a sequence of a fixed length. After obtaining a sufficient number of entity and relationship access sequences, vector representation learning is carried out using the skip-gram model. The principle is to given a center word, predict the context words within the window, and train the model by maximizing the probability of the context words for each word. After training, each word will obtain a corresponding low-dimensional vector, reflecting the position and semantic information of the word in the sequence.
[0132]
[0133]
[0134] Structural Information Learning: After the sequential semantic learning is completed, LDGCN further learns the adjacency structure information of entities and relationships through the GCN model. Specifically, LDGCN uses the output vector of Deepwalk as the initial vector of GCN, and captures the adjacency structure information of entities and relationships in the network security behavior knowledge graph through GCN. The adjacency information aggregation formula of the GCN model is:
[0135]
[0136] In the formula, W (l) represents the hidden weight of the l-th layer, H (l) is the vector representation of entities and relationships. It should be noted that H (l) contains not only the embedding vectors of entities but also the embedding vectors of relationships. In this embodiment, the relationship is regarded as a node, and the adjacency matrix of the knowledge graph is reconstructed by establishing edges between the connected entities, so that GCN can capture the adjacency information of relationships during learning. is the adjacency matrix of the knowledge graph after regarding the relationship as a node, is the symmetrically normalized degree matrix. In this embodiment, GCN is set to two layers to prevent overfitting of GCN.
[0137] Global information and attack sub-graph information acquisition: After two layers of information propagation of GCN, the nodes have captured the information of their second-order neighbors. At this time, pooling operations are performed on the complete network security behavior knowledge graph and the extracted attack sub-graph respectively to obtain the global representation vector and the attack logic sub-graph vector. The formula for the pooling operation is as follows:
[0138]
[0139] and
[0140]
[0141] where V is the set of entities and relationships in the network security behavior knowledge graph, s V is the set of entities and relationships in the attack logic sub-graph, α v is the learnable weight of entities and relationships, H g is the global representation vector of the network security knowledge graph, and H s is the attack logic sub-graph vector.
[0142] Key attack information fusion: After obtaining the entity and relationship vectors containing structural information through GCN, the vector representations of the complete network security behavior knowledge graph and the attack sub-graph are weighted and summed through weight aggregation to obtain a graph vector representation containing dual semantics. When aggregating weights, a larger learnable weight is assigned to the attack logic sub-graph vector to enhance the strength of the attack logic information in the graph representation vector and endow the model with a certain learning ability to adaptively adjust the weight parameters according to the classification results. The formula for vector fusion addition is as follows:
[0143] H = α g *H g + α s *H s
[0144] where α g is the global vector weight, and α s is the learnable weight of the attack logic sub-graph vector.
[0145] Attack type prediction: After obtaining the graph vector representation with dual semantics, a multi-layer perceptron (MLP) classifier is trained to detect the attack type. The input of the MLP is H, and the MLP non-linearly converts the graph representation vector into a prediction probability. The formula for the MLP is as follows:
[0146] z l = W l *H l-1 + b l
[0147] a l = σ(z l )
[0148] Wherein, W and b are respectively the linear transformation matrix and bias of the l-th layer, σ is the sigmod activation function, which performs a non-linear mapping on the graph vector. In the last layer of the MLP, the vector z is mapped to the same dimension as the attack category, and the predicted probability of each category is obtained through softmax. The specific formula is as follows:
[0149]
[0150] Wherein, p i is the probability of category i, z i and z j are the i-th and j-th dimensional elements of the vector, and n is the vector dimension.
[0151] Step (5): Model optimization.
[0152] After obtaining the probabilities, the multi-class cross-entropy loss and the contrastive learning InfoNCE loss are used to guide the training of the model. The cross-entropy loss is used to train the classification ability of the model so that it can map the input to the correct label; the contrastive learning loss further improves the quality of the graph embedding vector and the classification ability of the model by increasing the similarity between samples of the same class and decreasing the similarity between samples of different classes. At the same time, we add a cost-sensitive learning strategy to the cross-entropy loss to reduce the problem caused by sample imbalance. This enables the model to give more weights to small samples during the training process. The cross-entropy loss formula with cost-sensitive learning is as follows:
[0153]
[0154] Wherein, cost is a list, and cost[i] is the weight of the i-th category. n all represents the total number of samples, n class and M are both the number of categories, n i represents the number of samples of the i-th category, p ic is the predicted category probability of the model, y ic is the actual category of the sample. This strategy can increase the weight of the graph of the small sample category in the cross-entropy loss.
[0155] In the contrastive learning module, in this embodiment, InfoNCE is used as our loss function. InfoNCE effectively guides the model to learn the essential features of the data by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. The definition of the InfoNCE loss function is as follows:
[0156]
[0157] In the formula, q represents the representation of a given knowledge graph of network security behaviors, and k + represents the positive sample of q, that is, the sample belonging to the same attack category as q, and k i represents the negative sample of q, that is, the sample belonging to a different attack category from q. τ is the temperature parameter, which is used to adjust the output scale of the model. If the temperature τ is low, the probability distribution output by the model is more "sharp", and only samples with high similarity will get high probability values, while the probabilities of samples with low similarity will be suppressed very low. If the temperature τ is high, the probability distribution will be more "smooth", allowing more samples to have high probabilities. The purpose of contrastive learning is to make the graph representation vectors between samples of the same category close to each other, and the graph representations between samples of different categories far away from each other. The final loss of the model is:
[0158] Loss = L CE + γ * Con loss
[0159] In the formula, γ is the addition weight of the cross-entropy loss and the contrastive learning loss. The overall process of this part is as Figure 3 shown.
[0160] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present application, and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for constructing a host - side attack detection model based on attack sub - graphs, characterized in that, It includes the following steps: S10. Obtain and construct a network security behavior knowledge graph based on system logs; S20. Given a malicious event, and use a similarity-based reverse tracing algorithm to conduct an attack investigation on the malicious event, and extract an attack subgraph containing high-order attack logic; S30. Embed and represent the network security behavior knowledge graph through representation learning techniques; S40. Integrate the embedded knowledge graph with the attack subgraph, so that high-order attack logic information is added to the embedded knowledge graph, and then conduct attack type detection.
2. The method for constructing a host-side attack detection model based on an attack subgraph according to claim 1, wherein Specifically, step S10 is as follows: Graph structure modeling of log data, extract system entities, events and their attribute information from system logs, establish connections between entities through events, and organize the activities on the host represented by the logs in the form of a graph; Entity and event attribute screening, construct an attribute set, and determine the attributes related to attack investigation and attack detection; Integrate the graph structure modeling and attribute screening to construct a network security behavior knowledge graph with complete semantics.
3. The method for constructing a host-side attack detection model based on an attack subgraph according to claim 1, wherein Specifically, step S20 is as follows: Given a malicious event detected by security software; According to the malicious event, conduct reverse tracking through malicious similarity features and dependency weights to find the starting point on which the event depends; Obtain a forward subgraph through forward traversal from the starting point, and at the same time obtain a backward subgraph through backward traversal from the malicious event; Take the union of the forward subgraph and the backward subgraph to obtain the attack subgraph where the malicious event is located, so as to extract high-order attack logic information of malicious attacks.
4. The method for constructing a host - side attack detection model based on an attack sub - graph according to claim 3, wherein, The malicious similarity features include time similarity, node degree similarity, and data volume similarity.
5. The method for constructing a host - side attack detection model based on an attack sub - graph according to claim 3, wherein, Specifically, step S30 is as follows: Apply DeepWalk random walk to learn the local sequence semantic information of entities and relationships in the knowledge graph, and obtain vector representations of entities and relationships containing local sequence semantic information; Use the sequence semantic vectors learned by DeepWalk as the initial input vectors of the GCN model, and use the GCN model to capture the adjacency structure information between entities and relationships, and generate entity and relationship vectors that contain both sequence semantics and adjacency structure information.
6. The method for constructing a host - side attack detection model based on an attack sub - graph according to claim 3, wherein, Specifically, step S40 is as follows: Conduct separate pooling operations on the embedded knowledge graph and the attack subgraph to obtain a complete graph representation vector and an attack logic subgraph vector with attack information, and perform weighted fusion on the two to obtain a network security behavior knowledge graph feature vector for attack detection; Conduct attack type detection by training a multi-layer perceptron classifier to obtain the prediction probabilities of each attack type.
7. The method for constructing a host-side attack detection model based on an attack subgraph according to claim 1, wherein In step S40, the attack types include malicious programs, mining, Trojans, malware, ransoms, system viruses, worms, and backdoors.
8. The method for constructing a host-side attack detection model based on an attack subgraph according to claim 3, wherein, After step S40, it further includes: Guide the model to be trained through multi-class cross-entropy loss and contrastive learning InfoNCE loss.
9. The method for constructing a host - side attack detection model based on an attack sub - graph according to claim 8, wherein A cost-sensitive learning strategy is added to the cross-entropy loss.
10. A host - side attack detection model construction system based on attack sub - graphs, characterized in that, It includes: A knowledge graph construction module for obtaining and constructing a network security behavior knowledge graph based on system logs; An attack subgraph extraction module, which is used to given a malicious event and adopt a similarity-based reverse traceability algorithm to conduct an attack investigation on the malicious event, and extract an attack subgraph containing high-order attack logic; An embedding representation module, which is used to perform embedding representation on the network security behavior knowledge graph through representation learning techniques; An attack detection module, which is used to fuse the embedded knowledge graph with the attack subgraph, so that high-order attack logic information is added to the embedded knowledge graph, and then perform attack type detection.