Graph model data-based attack restoration method for advanced persistent threats

Through a method based on graph model data, combined with graph neural network and graph optimization algorithm, the attack process problem that the existing technology is difficult to detect and restore advanced persistent threats is solved, and more efficient and accurate attack process restoration is achieved.

CN120186031APending Publication Date: 2025-06-20XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214643.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

It is difficult to effectively detect and restore attack processes in advanced persistent threats (APTs), especially in the case of zero-day vulnerabilities and complex threat nodes.

Method used

The attack restoration method based on graph model data is adopted, and the data traceability diagram is constructed by extracting data from the system audit log of the network system, and the graph neural network model group is used to restore the attack process by combining graph pruning algorithm and graph traversal algorithm.

Benefits of technology

It improves the accuracy and completeness of attack restore, enhances the readability of the extracted attack flowchart, can more effectively capture the information of the attacker's intrusion path, and reduces the cost of attack restore.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186031A_ABST
    Figure CN120186031A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of attack restoration, in particular to a graph model data-based attack restoration method for advanced persistent threats, which comprises the following steps of: constructing a data traceability graph, performing data cleaning and reduction on the data traceability graph, extracting a benign data sub-graph and an attack data sub-graph, and storing the benign data sub-graph and the attack data sub-graph in a neo4j graph database; constructing a training set based on the benign data sub-graph, and constructing a detection set based on the attack data sub-graph; building a model group consisting of a plurality of graph neural networks, and inputting training data in the training set into the model group for iterative training to obtain a trained model group; inputting to-be-detected data in the detection set into the model group to obtain detection results of all nodes, and storing the detection results in an X sequence; and extracting node information according to the X sequence, and performing attack reduction in the Neo4j graph database by using graph pruning and graph traversal algorithms to obtain a corresponding attack flow chart. According to the method, the accuracy and integrity of attack restoration are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of attack restoration, and particularly to an attack restoration method for advanced persistent threats based on graph model data. Background Art

[0002] A cyber attack refers to any type of offensive action against a computer information system, infrastructure, computer network, or personal computer device. Nowadays, attackers tend to carry out intrusion attacks on the hosts of important network systems in large enterprises and government departments. Attackers often use APT (Advanced Persistent Threats) to achieve the goal of invading the target network. Such attacks usually use some undetected zero-day vulnerabilities to achieve the goal, so they have strong concealment and deception. Most traditional attack detection methods are based on nodes for threat detection, and it is difficult to capture the attacker's intrusion path and threat process from a large number of system logs, making it difficult for system administrators to formulate effective defense strategies to prevent the recurrence of threats.

[0003] Currently, some research teams have proposed some methods for extracting intrusion flowcharts based on system audit information. For example, the method based on path-level detection uses threat nodes to construct a sequence training model, uses the nodes to be detected to construct a sequence for detection, and finally gives an attack restoration graph based on the optimized graph data. However, this method depends on constructing an abnormal sequence based on abnormal nodes and depends on customized graph model data. Essentially, it is a supervised learning method and is difficult to effectively detect zero-day vulnerabilities. In addition, this method also has limitations on the number and format of abnormal threat nodes. When the number of threat nodes is too large or the abnormal situation is complex, it is difficult to play a role. The graph model data of advanced persistent threats is relatively complex and has a large order of magnitude. The data unrelated to threat nodes and attack paths accounts for a large proportion. It is difficult to restore the attack process using the method based on path-level detection. Summary of the Invention

[0004] In view of this, embodiments of the present application propose an attack restoration method for advanced persistent threats based on graph model data, which combines abnormal nodes and the restoration of attack flowcharts, improves the accuracy and integrity of attack restoration, and greatly improves the readability of the extracted attack flowcharts.

[0005] In a first aspect, an embodiment of the present application proposes a method for attacking reduction of advanced persistent threats based on graph model data. The method includes: extracting data from the system audit logs of a network system to construct a data traceability graph, performing data cleaning and data reduction on the data traceability graph, extracting a benign data sub-graph and an attack data sub-graph containing advanced persistent threat intrusions, and saving them to a neo4j graph database; constructing a training set based on the benign data sub-graph and constructing a detection set based on the attack data sub-graph; building a model group composed of multiple graph neural networks, inputting the training data in the training set into the model group, and some nodes in the training data serve as training nodes to enter the first graph neural network for training. After a batch of training is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, obtaining a trained model group; inputting the data to be detected in the detection set into the trained model group, obtaining the detection results of all nodes to be detected in the data to be detected, and storing them in the X sequence; extracting node information according to the X sequence, and using a graph pruning algorithm and a graph traversal algorithm to perform an attack reduction operation in the neo4j graph database to obtain an attack flow chart corresponding to the data to be detected.

[0006] Optionally, constructing a training set based on the benign data sub-graph and constructing a detection set based on the attack data sub-graph includes: using the benign data sub-graph as training data, obtaining all node data from the benign data sub-graph and processing it into a training format, selecting some nodes from all nodes and marking them as training nodes, and annotating each training node with a type label for characterizing the node type. Based on all the annotated training data, a training set is formed; using the attack data sub-graph as the data to be detected, and annotating each node in the data to be detected with a type label for characterizing the node type. Based on all the annotated data to be detected, a detection set is formed.

[0007] Optionally, multiple graph neural networks in the model group are all graph attention networks. In the attention mechanism of the graph attention network, there is an attention coefficient between each node and its neighbor nodes. Let the current node be node i, and a neighbor node of the current node be node j. The attention coefficient between node i and node j is expressed by the formula: where represents the projection matrix of node i, represents the projection matrix of node j, a(·) represents a function for calculating the correlation between two nodes, and e ijdenotes the attention coefficient between node i and node j; to simplify the calculation without affecting the detection accuracy, the graph attention network calculates the second-order neighbor nodes of the node to be detected. The second-order neighbor nodes of the node to be detected refer to the nodes that are two edges away from the node to be detected. There is one node between the second-order neighbor nodes and the node to be detected; the graph attention network first takes the adjacent nodes of the node to be detected as the first-order neighbor nodes of the node to be detected, calculates the attention coefficient between the node to be detected and the first-order neighbor nodes to update the feature value of the node to be detected, then takes the adjacent nodes of the first-order neighbor nodes except the node to be detected as the second-order neighbor nodes of the node to be detected, calculates the attention coefficient between the first-order neighbor nodes and the second-order neighbor nodes to update the feature value of the node to be detected, and finally calculates the attention coefficient between the updated node to be detected and the updated first-order neighbor nodes to update the feature value of the node to be detected again; for a set of input node features, the graph attention network first executes the self-attention mechanism on each node respectively, then uses the attention mechanism to calculate within the range of the second-order neighbor nodes of each node to update the feature value of each node, and finally uses a softmax function with an output dimension of 1 to calculate the predicted type of each node; in the model group training stage, the loss function is calculated using the predicted type and the type label for training; in the model group detection stage, the accuracy of the model group is judged by comparing the predicted type and the type label; in the actual application stage, the predicted type is output as the type judgment result of the node.

[0008] Optionally, the training data in the training set is input into the model group. Some nodes in the training data serve as training nodes and enter the first graph neural network for training. After a batch of training is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, obtaining a trained model group, including: sending the training nodes into the first graph neural network for training. The detection result of the graph neural network for the training nodes is the probability of classifying the training nodes into various types. Set the parameter γ as the ratio between the largest probability and the second-largest probability. The larger γ is set, the higher the requirement for the detection result, resulting in more training nodes that cannot be correctly classified. The smaller γ is set, the lower the requirement for the detection result, resulting in more training nodes that can be correctly classified; after a batch of training is completed, the correctly classified training nodes are removed, and the remaining nodes that are not correctly classified are sent into the next graph neural network for training until all training nodes are correctly classified, obtaining a trained model group.

[0009] Optionally, inputting the data to be detected in the detection set into the trained model group to obtain the detection results of all nodes to be detected in the data to be detected, including: inputting the data to be detected in the detection set into the trained model group, and each node to be detected in the data to be detected enters each graph neural network in turn for classification: in the case where all nodes to be detected are correctly classified in all graph neural networks, mark the nodes to be detected as benign nodes, and in the case where at least one node to be detected is not correctly classified in at least one graph neural network, mark the nodes to be detected as threat nodes.

[0010] Optionally, extract node information according to the X sequence, and use the graph pruning algorithm and the graph traversal algorithm to perform an attack restoration operation in the neo4j graph database to obtain the attack flow chart corresponding to the data to be detected, including: starting from a node selected from the X sequence to extract first-order neighbor nodes, and for the threat nodes among them, also extract first-order neighbor nodes and repeat the graph pruning algorithm. For the benign nodes among them, detect the first-order neighbor nodes of the repeated types. For the edge nodes, only keep one node of the same type. For the non-edge nodes, detect whether there are nodes in the X sequence among their first-order neighbor nodes. If so, continue to repeat the graph pruning algorithm for the threat nodes; where the edge nodes are the nodes with only one in-degree, and the non-edge nodes are the nodes with two or more in-degrees; for the threat nodes, count their path types, sort them according to the proportion size. When it is detected that there are nodes in the X sequence and the graph pruning algorithm has ended, starting from a threat node in the X sequence, first use the graph pruning algorithm to extract first-order neighbor nodes. When it is impossible to connect to the main subgraph, add the path with the highest proportion among the threat nodes and the nodes connected to it to the subgraph, and continue to query from this node. If this node is an edge node, backtrack and re-query the path with the second highest proportion, and so on, until the two subgraphs are connected into one subgraph, and finally obtain the attack flow chart corresponding to the data to be detected.

[0011] Optionally, the neo4j graph database is a high-performance NOSQL graph database, an embedded, disk-based, Java persistence engine with full transaction characteristics, and can be regarded as a high-performance graph engine, with all the characteristics of a mature database; the neo4j graph database stores structured data on the network rather than in tables. The py2neo library is a third-party library based on python. Using the py2neo library can conveniently and simply operate the neo4j database, and when saving the benign data subgraph and the attack data subgraph to the neo4j graph database, it is implemented through the operation of the py2neo library.

[0012] An attack restoration method for advanced persistent threats based on graph model data proposed in this application extracts an attack flow chart based on a node-level model group, which can accurately capture the effective information of the attacker's intrusion path. Considering that in the scenario of intrusion detection, the number of nodes is unbalanced, when performing anomaly detection, if only a single graph neural network is used for detection, it is very likely that the detection will be incorrect because the nodes cannot be correctly classified and are classified as threat nodes. The model group can effectively learn hidden features and improve the detection accuracy. The feature extraction process of the graph neural network makes full use of the neighbor information of the nodes, achieving the purpose of maximizing the detection efficiency. When performing attack restoration, the graph pruning algorithm and the graph traversal algorithm are used to extract attack information, removing redundant nodes and edges, and only retaining the key information in the attack process, effectively improving the efficiency of attack restoration and reducing the cost of attack restoration. The neo4j graph database is used to store the benign data subgraph and the attack data subgraph. Using the graph pruning algorithm and the graph traversal algorithm to perform attack restoration operations in the neo4j graph database can make full use of the high performance of the neo4j graph database. All in all, this application combines the anomaly node and the attack flow chart restoration, improving the accuracy and integrity of the attack restoration, and greatly improving the readability of the extracted attack flow chart.

[0013] In a second aspect, an embodiment of this application proposes an attack restoration system for advanced persistent threats based on graph model data. The system includes: a data extraction and preprocessing module, configured to extract data from the system audit logs of the network system to construct a data traceability graph, perform data cleaning and data reduction on the data traceability graph, extract a benign data subgraph and an attack data subgraph containing advanced persistent threat intrusions, and store them in the neo4j graph database; a data set model group construction module, configured to construct a training set based on the benign data subgraph, construct a detection set based on the attack data subgraph, and build a model group composed of multiple graph neural networks; a model group training module, configured to input the training data in the training set into the model group. Some nodes in the training data enter the first graph neural network for training. When a batch of training is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, obtaining a trained model group; a detection module, configured to input the data to be detected in the detection set into the trained model group, obtain the detection results of all nodes to be detected in the data to be detected, and store them in the X sequence; an attack restoration module, configured to extract node information according to the X sequence, and use the graph pruning algorithm and the graph traversal algorithm to perform attack restoration operations in the neo4j graph database to obtain the attack flow chart corresponding to the data to be detected.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute an attack restoration method for advanced persistent threats based on graph model data as described in the first aspect above.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement an attack restoration method for advanced persistent threats based on graph model data as described in the first aspect above.

[0016] It can be understood that the beneficial effects of the second to fourth aspects above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of an attack restoration method for advanced persistent threats based on graph model data provided in an embodiment of the present application;

[0019] Figure 2 is a visualization diagram of an attack restoration method for advanced persistent threats based on graph model data provided in an embodiment of the present application;

[0020] Figure 3 is a schematic diagram of graph model data stored in a neo4j graph database provided in an embodiment of the present application;

[0021] Figure 4 is a schematic diagram of the optimization process of an attack flowchart provided in an embodiment of the present application;

[0022] Figure 5 is a schematic diagram of a graph pruning algorithm provided in an embodiment of the present application;

[0023] Figure 6 is a schematic diagram of a graph traversal algorithm provided in an embodiment of the present application;

[0024] Figure 7It is the restored attack flow chart provided in an embodiment of the present application;

[0025] Figure 8 It is a schematic diagram of the simulation results of the theia dataset, cadets dataset, and trace dataset provided in an embodiment of the present application;

[0026] Figure 9 It is a schematic diagram of the simulation results of the fivedirections dataset, SC-1 dataset, and SC-2 dataset provided in an embodiment of the present application;

[0027] Figure 10 It is a schematic structural diagram of an advanced persistent threat attack restoration system based on graph model data provided in another embodiment of the present application;

[0028] Figure 11 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be elaborated in detail below with reference to the accompanying drawings. In various embodiments of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of various embodiments is only for convenient description and should not constitute any limitation on the specific implementation manners of the present application. Various embodiments can be combined and cross-referenced with each other on the premise of no conflict.

[0030] To solve the technical problem that the existing attack restoration method based on path-level detection is difficult to restore the attack process containing advanced persistent threats, an embodiment of the present application proposes an attack restoration method for advanced persistent threats based on graph model data, which is applied to an electronic device. The electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described by taking the server as an example. The implementation details of the attack restoration method for advanced persistent threats based on graph model data proposed in this embodiment are specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing the solution.

[0031] The specific process of the attack restoration method for advanced persistent threats based on graph model data proposed in this embodiment can be as Figure 1 shown, and its visualized representation is as Figure 2 shown, and specifically includes:

[0032] Step 101: Extract data from the system audit log of the network system to construct a data traceability graph. Perform data cleaning and data reduction on the data traceability graph, extract a benign data sub-graph and an attack data sub-graph containing advanced persistent threat (APT) intrusions, and save them to the Neo4j graph database.

[0033] In a specific implementation, the server first needs to extract data from the system audit log of the network system to construct a data traceability graph. Next, it needs to perform preprocessing including data cleaning and data reduction on the data traceability graph to extract a benign data sub-graph and an attack data sub-graph containing APT intrusions, and save them to the Neo4j graph database.

[0034] In one example, the server uses the Camflow tool to extract data from the system audit log of the network system. The extracted data can be called graph model data, and its data structure is as Figure 3 shown.

[0035] It should be noted that this embodiment is faced with attack behaviors using APTs, and it is necessary to assume that the attacker has the following characteristics.

[0036] Firstly, it is concealment. The attacker will not directly show their attack activities, but rather hide their activities like a real intrusion, trying to mix their behaviors with a large amount of benign background data, which makes the victim's network system behave like a benign system.

[0037] Secondly, it is the use of zero-day vulnerabilities. A zero-day vulnerability that has not been discovered by the victim in advance often causes great damage. This embodiment assumes that the attacker will use zero-day vulnerabilities to attack the network system. Therefore, the model group in this embodiment does not use any prior experience regarding zero-day vulnerabilities for training.

[0038] Finally, it has attack characteristics. No matter how the attacker conceals themselves, in order to complete a malicious act, the intrusion activity will definitely show characteristics different from benign activities. These characteristics will be manifested in the extracted data traceability graph, making the attacker nodes show local structural characteristics different from normal nodes.

[0039] It should be noted that the Neo4j graph database is a high-performance NoSQL (Not Only SQL) graph database, an embedded, disk-based, fully transactional Java persistence engine, which can be regarded as a high-performance graph engine with all the characteristics of a mature database. The Neo4j graph database stores structured data in a network rather than in tables. The py2neo library is a third-party library based on Python. Using the py2neo library can easily and simply operate the Neo4j database. When saving the benign data subgraph and the attack data subgraph into the Neo4j graph database, it is achieved through the operation of the py2neo library.

[0040] Step 102: Construct a training set based on the benign data subgraph and a detection set based on the attack data subgraph.

[0041] In a specific implementation, the graph neural network in the model group uses a forward propagation algorithm to aggregate information from node ancestors. Its input is a set of node features, and then a new set of node features is generated as the output. Since the model group needs to meet the requirement of detecting zero-day vulnerabilities, this embodiment designs and uses a set of feature extraction methods and label assignment methods that do not require abnormal nodes. That is to say, the training data in the training set of the model group all comes from the benign data subgraph, while the attack data subgraph is used in the detection stage of the model group to form the detection set.

[0042] In an example, the server uses the benign data subgraph as training data, obtains all node data from the benign data subgraph and processes it into a training format. Some nodes are selected from all nodes and marked as training nodes, and type labels for characterizing node types are assigned to each training node. The training set is composed of all the labeled training data.

[0043] In an example, the server uses the attack data subgraph as the data to be detected, and assigns type labels for characterizing node types to each node in the data to be detected. The detection set is composed of all the labeled data to be detected.

[0044] Step 103: Build a model group composed of multiple graph neural networks, input the training data in the training set into the model group. Some nodes in the training data enter the first graph neural network for training as training nodes. When a batch of training is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, and the trained model group is obtained.

[0045] In a specific implementation, the server needs to build a model group composed of multiple graph neural networks (at least two), and input the training data in the training set into the model group. Some nodes in the training data enter the first graph neural network as training nodes for training. After a batch of training is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, and a trained model group is obtained.

[0046] In one example, the multiple graph neural networks in the model group are all graph attention networks. The graph attention network can handle various graph-related tasks and has achieved success in many fields. In this embodiment, its graph processing ability is used to learn the structural information of nodes in the data traceability graph, so as to detect the threats of abnormal nodes.

[0047] In the attention mechanism of the graph attention network, there is an attention coefficient between each node and its neighbor nodes. Let the current node be node i, and a neighbor node of the current node be node j. The attention coefficient between node i and node j is expressed by the formula as:

[0048]

[0049] Among them, represents the projection matrix of node i, represents the projection matrix of node j, a(·) represents a function for calculating the correlation between two nodes, and e ij represents the attention coefficient between node i and node j.

[0050] Theoretically, the application of this attention mechanism can calculate all points on the graph, but in practical applications, a masked attention mechanism is often used to sample a smaller amount of node data for calculation. The detection of abnormal nodes often only involves nodes close to the node. To simplify the calculation and not affect the detection accuracy, the graph attention network calculates the second-order neighbor nodes of the node to be detected. The second-order neighbor nodes of the node to be detected refer to the nodes that are two edges away from the node to be detected, and there is one node between the second-order neighbor nodes and the node to be detected.

[0051] The graph attention network first takes the adjacent nodes of the node to be detected as the first-order neighbor nodes of the node to be detected, calculates the attention coefficient between the node to be detected and the first-order neighbor nodes to update the feature value of the node to be detected, then takes the adjacent nodes of the first-order neighbor nodes except the node to be detected as the second-order neighbor nodes of the node to be detected, calculates the attention coefficient between the first-order neighbor nodes and the second-order neighbor nodes to update the feature value of the node to be detected, and finally calculates the attention coefficient between the updated node to be detected and the updated first-order neighbor nodes to update the feature value of the node to be detected again.

[0052] For a set of input node features, the graph attention network first performs a self-attention mechanism on each node separately, then uses the attention mechanism to calculate within the range of the second-order neighbor nodes of each node to update the feature values of each node, and finally uses a softmax function with an output dimension of 1 to calculate the predicted type of each node. In the model group training stage, the server calculates the loss function using the predicted type and the type label for training. In the model group detection stage, the server determines the accuracy of the model group by comparing the predicted type and the type label. In the actual application stage, the model group outputs the predicted type as the type judgment result of the node.

[0053] During the training process, considering that nodes of the same type may play different roles, it is difficult to use a single model to classify nodes. In the scenario of intrusion detection, the number of nodes is unbalanced (for example, when there are no frequent network connections in a host, there may be thousands of process nodes and very few remote nodes). Therefore, when performing anomaly detection, it is necessary to make a clear division of the node features, dividing the node features into dominant features and hidden features. The hidden features represent the special functions of the nodes. If only a single model is used for detection, nodes that cannot be correctly classified will be misclassified as threat nodes, resulting in detection errors. Training a model group can effectively learn these hidden features.

[0054] In one example, the server sends the training nodes into the first graph neural network for training. The graph neural network's detection result for the training nodes is the probability of classifying the training nodes into various types. Set the parameter γ as the ratio between the largest probability and the second-largest probability. The larger γ is set, the higher the requirement for the detection result, resulting in more training nodes that cannot be correctly classified. The smaller γ is set, the lower the requirement for the detection result, resulting in more training nodes that can be correctly classified. After a batch of training is completed, the correctly classified training nodes are removed, and the remaining uncorrectly classified nodes are sent into the next graph neural network for training until all training nodes are correctly classified, obtaining a trained model group.

[0055] Step 104: Input the data to be detected in the detection set into the trained model group, obtain the detection results of all the nodes to be detected in the data to be detected, and store them in the X sequence.

[0056] In a specific implementation, after the server obtains the trained model group, it is necessary to input the data to be detected in the detection set into the trained model group, obtain the detection results of all the nodes to be detected in the data to be detected, and store them in the X sequence.

[0057] In one example, the server inputs the data to be detected in the detection set into the trained model group, and each node to be detected in the data to be detected enters each graph neural network in turn for classification. In the case where all the nodes to be detected are correctly classified in all the graph neural networks, the nodes to be detected are marked as benign nodes. In the case where at least one of the nodes to be detected is not correctly classified in at least one graph neural network, the nodes to be detected are marked as threat nodes.

[0058] Step 105: Extract node information according to the X sequence, and use the graph pruning algorithm and the graph traversal algorithm to perform an attack restoration operation in the Neo4j graph database to obtain the attack flow chart corresponding to the data to be detected.

[0059] In a specific implementation, after obtaining the X sequence, the server will extract node information according to the X sequence, and use the graph pruning algorithm and the graph traversal algorithm to perform an attack restoration operation in the Neo4j graph database to obtain the attack flow chart corresponding to the data to be detected. The graph pruning algorithm and the graph traversal algorithm together constitute the graph optimization algorithm. In a complex graph model data, many data are conventional information, which are redundant information in the attack flow chart. If a simple method of extracting first-order neighbor nodes is adopted, the extracted information cannot well reflect the attack means of the intruder. As a system administrator, it is also impossible to quickly capture effective information from such an attack flow chart and cannot make good defenses. In addition to the blockage of redundant information, simple extraction means can only extract the neighbor information of abnormal nodes. For the attack method of hidden long lines, it is often difficult to see the whole process of the attack, while the graph pruning algorithm and the graph traversal algorithm can solve these problems.

[0060] In one example, the graph optimization process is as Figure 4 shown, the principle of the graph pruning algorithm is as Figure 5 shown, and the principle of the graph traversal algorithm is as Figure 6As shown in the figure. The server selects a node from the X sequence to extract first-order neighbor nodes, extracts first-order neighbor nodes for the threat nodes among them and repeats the graph pruning algorithm. For the benign nodes among them, it detects the first-order neighbor nodes of the repeated types. For the edge nodes, only one node of the same type is retained. For the non-edge nodes, it detects whether there are nodes in the X sequence among their first-order neighbor nodes. If so, it continues to repeat the graph pruning algorithm for the threat node. Among them, the edge node is a node with only one in-degree, and the non-edge node is a node with two or more in-degrees. For the threat nodes, it counts their path types, sorts them according to the proportion, and when it detects that there are nodes in the X sequence and the graph pruning algorithm has ended, starting from a threat node in the X sequence, it first uses the graph pruning algorithm to extract first-order neighbor nodes. When it cannot be connected to the main subgraph, it adds the path with the highest proportion in the threat nodes and the nodes connected to it to the subgraph, and continues to query from this node. If this node is an edge node, it backtracks and queries the path with the second highest proportion, and so on, until the two subgraphs are connected into one subgraph, and finally obtains the attack flow chart corresponding to the data to be detected.

[0061] In one example, the restored attack flow chart can be as Figure 7 shown. The paths and nodes irrelevant to the attack are all trimmed, and only the core network topology related to the attack is retained. From Figure 7 it can be seen that the optimized attack flow chart is intuitive, concise and highly readable.

[0062] An attack restoration method for advanced persistent threats based on graph model data proposed in this embodiment extracts the attack flow chart based on the node-level model group, and can accurately capture the effective information of the attacker's intrusion path. Considering that in the scenario of intrusion detection, the number of nodes is unbalanced, so when performing anomaly detection, if only a single graph neural network is used for detection, it is very likely to be misclassified as a threat node due to incorrect classification, resulting in detection errors. The model group can effectively learn hidden features and improve the detection accuracy. The feature extraction process of the graph neural network makes full use of the neighbor information of the nodes, achieving the purpose of maximizing the detection efficiency. When performing attack restoration, the graph pruning algorithm and the graph traversal algorithm are used to extract attack information, removing redundant nodes and edges, and only retaining the key information in the attack process, effectively improving the efficiency of attack restoration and reducing the cost of attack restoration. The neo4j graph database is used to store the benign data subgraph and the attack data subgraph, and the graph pruning algorithm and the graph traversal algorithm are used to perform attack restoration operations in the neo4j graph database, which can make full use of the high performance of the neo4j graph database.

[0063] All in all, this embodiment combines the anomaly node and the attack flow chart restoration, improving the accuracy and integrity of the attack restoration, and greatly improving the readability of the extracted attack flow chart.

[0064] The step division of the above various methods is only for clear description. When implementing, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, is within the protection scope of this application.

[0065] In one embodiment, to verify the detection performance of the model group, we conducted a comparative experiment between the model group (hereinafter referred to as OURS) and the flash model. The results of the comparative experiment are as Figure 8 、 Figure 9 shown. In the theia dataset, cadets dataset, trace dataset, fivedirections dataset, SC-1 dataset, and SC-2 dataset, OURS achieved better performance than the flash model in terms of the prec metric, recall metric, and f1 metric.

[0066] Another embodiment of this application proposes an attack restoration system for advanced persistent threats based on graph model data. The following specifically describes the implementation details of an attack restoration system for advanced persistent threats based on graph model data proposed in this embodiment. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this example. Figure 10 FIG. is a schematic structural diagram of an attack restoration system for advanced persistent threats based on graph model data proposed in this embodiment. The system includes: a data extraction and preprocessing module 201, a data set model group construction module 202, a model group training module 203, a detection module 204, and an attack restoration module 205.

[0067] The data extraction and preprocessing module 201 is used to extract data from the system audit logs of the network system to construct a data traceability graph, perform data cleaning and data reduction on the data traceability graph, extract a benign data subgraph and an attack data subgraph containing advanced persistent threat intrusions, and save them to the neo4j graph database.

[0068] The data set model group construction module 202 is used to construct a training set based on the benign data subgraph, construct a detection set based on the attack data subgraph, and build a model group composed of multiple graph neural networks.

[0069] The model group training module 203 is configured to input the training data in the training set into the model group. Some nodes in the training data serve as training nodes and enter the first graph neural network for training. After the training of one batch is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, and the trained model group is obtained.

[0070] The detection module 204 is configured to input the data to be detected in the detection set into the trained model group, obtain the detection results of all nodes to be detected in the data to be detected, and store them in the X sequence.

[0071] The attack restoration module 205 is configured to extract node information according to the X sequence, and perform attack restoration operations in the neo4j graph database using the graph pruning algorithm and the graph traversal algorithm to obtain the attack flow chart corresponding to the data to be detected.

[0072] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0073] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above method embodiment are still valid in this embodiment, and will not be repeated here to avoid redundancy. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.

[0074] Another embodiment of this application proposes an electronic device, and its specific structure is as Figure 11 shown, including: at least one processor 301; and a memory 302 communicatively connected to the at least one processor 301; wherein, the memory 302 stores instructions executable by the at least one processor 301, and the instructions are executed by the at least one processor 301 to enable the at least one processor 301 to execute an attack restoration method for advanced persistent threats based on graph model data as described in each of the above method embodiments.

[0075] Among them, the memory and the processor can be connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be further described herein. The bus interface is responsible for providing an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices over a transmission medium. The data processed by the processor is transmitted over a wireless medium via an antenna. Further, the antenna also receives data and transmits the data to the processor.

[0076] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.

[0077] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement an attack restoration method for advanced persistent threats based on graph model data as described in the above method embodiments.

[0078] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (such as a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, external hard drives, ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0079] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.

Claims

1. A method for attack restoration of advanced persistent threats based on graph model data, characterized in that: include: Extract data from the system audit log of the network system to build a data traceability graph, perform data cleaning and data reduction on the data traceability graph, extract benign data subgraphs and attack data subgraphs containing advanced persistent threat intrusions, and save them in the neo4j graph database; Construct a training set based on the benign data subgraph, and construct a detection set based on the attack data subgraph; Build a model group consisting of multiple graph neural networks, input the training data in the training set into the model group, and some nodes in the training data will enter the first graph neural network as training nodes for training. When a batch of training is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, and a trained model group is obtained; Input the data to be tested in the test set into the trained model group, obtain the test results of all nodes to be tested in the data to be tested, and store them in the X sequence; Node information is extracted according to the X sequence, and the graph pruning algorithm and graph traversal algorithm are used to perform attack restoration operations in the neo4j graph database to obtain the attack flow chart corresponding to the data to be detected.

2. The method for attack restoration of advanced persistent threats based on graph model data as claimed in claim 1, characterized in that: The training set is constructed based on the benign data subgraph, and the detection set is constructed based on the attack data subgraph, including: The benign data subgraph is used as training data, all node data is obtained from the benign data subgraph and processed into a training format, some nodes are selected from all nodes and marked as training nodes, and each training node is labeled with a type label used to characterize the node type, and a training set is formed based on all labeled training data; The attack data subgraph is used as the data to be detected, and each node in the data to be detected is labeled with a type label used to characterize the node type. A detection set is formed based on all the labeled data to be detected.

3. The method for attack restoration of advanced persistent threats based on graph model data as claimed in claim 2, characterized in that: Multiple graph neural networks in the model group are all graph attention networks. In the attention mechanism of the graph attention network, there is an attention coefficient between each node and its neighboring nodes. Let the current node be node i, and one of the neighboring nodes of the current node be node j. The attention coefficient between node i and node j is expressed by the formula: in, represents the projection matrix of node i, represents the projection matrix of node j, a(·) represents the function of calculating the correlation between two nodes, e ij Represents the attention coefficient between node i and node j; In order to simplify the calculation without affecting the accuracy of detection, the graph attention network calculates the second-order neighbor nodes of the node to be detected. The second-order neighbor nodes of the node to be detected refer to the nodes that are two edges away from the node to be detected. There is a node between the second-order neighbor nodes and the node to be detected. The graph attention network first takes the neighboring nodes of the node to be detected as the first-order neighboring nodes of the node to be detected, calculates the attention coefficient between the node to be detected and the first-order neighboring nodes to update the feature value of the node to be detected, and then takes the neighboring nodes of the first-order neighboring nodes except the node to be detected as the second-order neighboring nodes of the node to be detected, calculates the attention coefficient between the first-order neighboring nodes and the second-order neighboring nodes to update the feature value of the node to be detected, and finally calculates the attention coefficient between the updated node to be detected and the updated first-order neighboring nodes to update the feature value of the node to be detected again; For a set of input node features, the graph attention network first performs a self-attention mechanism on each node, then uses the attention mechanism to calculate within the range of the second-order neighbor nodes of each node to update the feature value of each node, and finally uses a softmax function with an output dimension of 1 to calculate the predicted type of each node; In the model group training phase, the loss function is calculated using the predicted type and type label for training; In the model group detection phase, the accuracy of the model group is determined by comparing the predicted types and type labels; In the actual application stage, the predicted type is output as the type judgment result of the node.

4. A method for attack restoration of advanced persistent threats based on graph model data as claimed in any one of claims 2 to 3, characterized in that: The training data in the training set is input into the model group. Some nodes in the training data are used as training nodes to enter the first graph neural network for training. When a batch of training is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, and a trained model group is obtained, including: The training nodes are sent to the first graph neural network for training. The detection result of the graph neural network on the training nodes is the probability of classifying the training nodes into various types. The parameter γ is set to the ratio between the largest probability and the second largest probability. The larger the γ is set, the higher the requirement for the detection result, resulting in more training nodes that cannot be correctly classified. The smaller the γ is set, the lower the requirement for the detection result, resulting in more training nodes that can be correctly classified. When a batch of training is completed, the correctly classified training nodes are removed, and the remaining incorrectly classified nodes are sent to the next graph neural network for training until all training nodes are correctly classified and a trained model group is obtained.

5. The method for attack restoration of advanced persistent threats based on graph model data as claimed in claim 1, characterized in that: Input the data to be tested in the test set into the trained model group to obtain the test results of all nodes to be tested in the data to be tested, including: The data to be tested in the test set is input into the trained model group, and each node to be tested in the data to be tested is sequentially entered into each graph neural network for classification: When the node to be detected is correctly classified in all graph neural networks, the node to be detected is marked as a benign node. When the node to be detected is not correctly classified in at least one graph neural network, the node to be detected is marked as a threat node.

6. The method for attack restoration of advanced persistent threats based on graph model data as claimed in claim 5, characterized in that: Extract node information according to the X sequence, use the graph pruning algorithm and graph traversal algorithm to perform attack restoration operations in the neo4j graph database, and obtain the attack flow chart corresponding to the data to be detected, including: Select a node from the X sequence to extract the first-order neighbor nodes. For the threat nodes, extract the first-order neighbor nodes and repeat the graph pruning algorithm. For the benign nodes, detect the repeated first-order neighbor nodes. For the edge nodes, only keep one node of the same type. For the non-edge nodes, detect whether there is a node in the X sequence among their first-order neighbor nodes. If so, continue to repeat the graph pruning algorithm for the threat node. Among them, the edge node is a node with only one in-degree, and the non-edge node is a node with two or more in-degrees. For threat nodes, their path types are counted and sorted by proportion. When it is detected that there are nodes in the X sequence and the graph pruning algorithm has ended, starting from a threat node in the X sequence, the graph pruning algorithm is first used to extract the first-order neighbor nodes. When it is impossible to connect to the main subgraph, the path with the highest proportion in the threat node and the nodes connected to it are added to the subgraph, and the query continues from this node. If the node is an edge node, it will fall back and re-query the path with the second highest proportion, and so on, until the two subgraphs are connected into one subgraph, and finally the attack flow chart corresponding to the data to be detected is obtained.

7. A method for attack restoration of advanced persistent threats based on graph model data according to any one of claims 1 to 6, characterized in that: The neo4j graph database is a high-performance NOSQL graph database. It is an embedded, disk-based, fully transactional Java persistence engine. It can be regarded as a high-performance graph engine with all the features of a mature database. The neo4j graph database stores structured data on the network instead of in a table. The py2neo library is a third-party library based on Python. The py2neo library can be used to operate the neo4j database conveniently and simply. When the benign data subgraph and the attack data subgraph are saved in the neo4j graph database, the operation is implemented through the py2neo library.

8. An attack recovery system for advanced persistent threats based on graph model data, characterized in that: include: The data extraction and preprocessing module is used to extract data from the system audit log of the network system to build a data traceability graph, perform data cleaning and data reduction on the data traceability graph, extract benign data subgraphs and attack data subgraphs containing advanced persistent threat intrusions, and save them in the neo4j graph database; The data set model group construction module is used to build a training set based on the benign data subgraph, build a detection set based on the attack data subgraph, and build a model group composed of multiple graph neural networks; The model group training module is used to input the training data in the training set into the model group. Some nodes in the training data are used as training nodes to enter the first graph neural network for training. When the training of a batch is completed, the nodes that are not correctly classified will enter the next graph neural network for training until all training nodes are correctly classified, and a trained model group is obtained; The detection module is used to input the data to be detected in the detection set into the trained model group, obtain the detection results of all the nodes to be detected in the data to be detected, and store them in the X sequence; The attack restoration module is used to extract node information according to the X sequence, use the graph pruning algorithm and the graph traversal algorithm to perform attack restoration operations in the neo4j graph database, and obtain the attack flow chart corresponding to the data to be detected.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an attack restoration method for advanced persistent threats based on graph model data as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement an attack restoration method for advanced persistent threats based on graph model data as described in any one of claims 1 to 7.