Network attack tracing method based on interpretable graph neural network and related device

By constructing and analyzing the network attack traceability map based on interpretable graph neural networks, the shortcomings of the existing technology in real-time, accuracy and reliability are solved, and more efficient and reliable network attack traceability are achieved.

CN120034384APending Publication Date: 2025-05-23CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +3
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510197320.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing cyber attack traceability technology has flaws in real-time, accuracy and reliability, and it is difficult to effectively deal with long-chain attacks and multi-stage attacks in complex network environments.

Method used

A method based on interpretable graph neural network is adopted to construct a traceability map by obtaining network attack traceability log data, obtain feature vectors, and filter the abnormal traceability map through an exception detection model. Combined with interpretable methods, key nodes and key edges are obtained, and the minimum abnormal traceability sub-map is generated. Finally, a traceability analysis map is obtained to locate the attack propagation path and attack source.

Benefits of technology

It improves the accuracy and completeness of tracing the source of cyber attacks, enhances the reliability and transparency of tracing, and helps security teams quickly locate attack sources and attack paths, and then adopts targeted defense measures to minimize losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034384A_ABST
    Figure CN120034384A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of network space security, and discloses a network attack traceability method based on an interpretable graph neural network and a related device, and the method comprises the steps: obtaining the attack traceability log data of a network in a preset time period, and constructing a traceability graph at each moment in the preset time period according to the attack traceability log data; obtaining the feature vector of each traceability graph, and through a pre-trained graph neural network-based anomaly detection model, selecting the traceability graphs whose anomaly detection results are abnormal to obtain a plurality of abnormal traceability graphs; obtaining key nodes and key edges of each abnormal traceability graph through an interpretable method, and obtaining a minimum abnormal traceability sub-graph of each abnormal traceability graph according to the key nodes and the key edges; and combining the minimum anomaly traceability sub-graphs of the anomaly traceability graphs to obtain a traceability analysis graph, and obtaining an attack propagation path and an attack source of the network attack according to the traceability analysis graph. The problems that the model output result cannot be effectively explained and the traceability precision is low at present are solved, and the attack traceability accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of cyberspace security and relates to a network attack tracing method based on an interpretable graph neural network and a related device. Background Art

[0002] In today's era of rapid development of informatization, the network environment is becoming increasingly complex, and cyber attacks are frequent, posing serious threats to individuals, enterprises and national security. The means of cyber attacks are constantly evolving, from simple virus propagation to complex hacker intrusions, ransomware and advanced persistent threats (APT), etc. The attack chain is getting longer and longer, and the concealment is getting stronger. In order to effectively deal with these cyber attacks, accurately tracing the source of the attack has become an important part of network security protection. Attack tracing can not only help victims locate attackers in time and take corresponding defensive measures, but also provide strong evidence for law enforcement agencies to combat cybercrime. Therefore, the development of efficient and accurate attack tracing technology is of great significance to maintaining network security.

[0003] At present, network attack source tracing technology mainly relies on traditional means such as log analysis and traffic monitoring. Log analysis collects and analyzes log information generated by network devices, servers, and application systems, and attempts to restore the attack process and track the source of the attack. Traffic monitoring monitors and analyzes network traffic in real time to identify abnormal traffic patterns and thus discover potential network attack behaviors. These technologies can effectively assist in attack source tracing to a certain extent and provide strong support for network security protection. However, with the increasing complexity of the network environment and the continuous development of attack methods, these traditional technologies are also facing many challenges.

[0004] Existing log analysis and traffic monitoring technologies have many defects in attack tracing. First, poor real-time performance is a prominent problem. Log information often needs to be analyzed some time after the attack occurs, and it is impossible to track the source of the attack in real time, resulting in delayed defense measures. Secondly, the high false alarm rate is also an important factor restricting the application effect of these technologies. Due to the complex and changeable network environment, there is a lot of noise and interference in log information and traffic data, which can easily lead to false alarms and missed reports, affecting the accuracy of attack tracing. In addition, long chain attacks and multi-stage attacks in complex network environments are often difficult to handle effectively and cannot be accurately traced. Therefore, improving the accuracy, real-time and reliability of attack tracing has become a key problem that needs to be solved urgently. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a network attack tracing method and related devices based on an interpretable graph neural network.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a network attack tracing method based on an interpretable graph neural network, comprising: obtaining attack tracing log data of the network within a preset time period, and constructing a tracing graph at each moment within the preset time period based on the attack tracing log data; obtaining feature vectors of each tracing graph and selecting a tracing graph with an abnormal detection result as an abnormality through a pre-trained anomaly detection model based on a graph neural network to obtain a plurality of abnormal tracing graphs; obtaining key nodes and key edges of each abnormal tracing graph through an interpretable method, and obtaining a minimum abnormal tracing subgraph of each abnormal tracing graph based on the key nodes and key edges; combining the minimum abnormal tracing subgraphs of each abnormal tracing graph to obtain a tracing analysis graph, and obtaining an attack propagation path and attack source of the network attack based on the tracing analysis graph.

[0008] Optionally, constructing a tracing graph at each moment within a preset time period based on the attack tracing log data includes: constructing a basic graph with entities in the network as nodes and relationships between entities as edges; setting a moment at regular intervals within the preset time period, and obtaining node features of the nodes and edge features of the edges at each moment based on the attack tracing log data; and adding the node features of the nodes and edge features of the edges at each moment to the basic graph, respectively, to obtain the tracing graph at each moment within the preset time period.

[0009] Optionally, the preset time period is: the time when the attack event is discovered is the starting time, the time two weeks before the starting time is the ending time, and the time period between the starting time and the ending time is the preset time period; the certain time is half an hour; the node features include one or more of the following: IP address, user name and device attributes; the edge features include one or more of the following: timestamp, login and access frequency, call type, file type and traffic size.

[0010] Optionally, obtaining the feature vector of each provenance graph includes: obtaining the feature vector of the provenance graph through a graph convolutional neural network.

[0011] Optionally, the obtaining of key nodes and key edges of each anomaly traceability graph through an interpretable method includes: obtaining the feature weights of each node and each edge in each anomaly traceability graph by using the GNNExplainer interpreter; combining a preset node feature weight threshold, taking nodes whose feature weights are not less than the node feature weight threshold as key nodes, and obtaining the key nodes of each anomaly traceability graph; combining a preset edge feature weight threshold, taking edges whose feature weights are not less than the edge feature weight threshold as key edges, and obtaining the key edges of each anomaly traceability graph.

[0012] Optionally, obtaining the attack propagation path and attack source of the network attack based on the source tracing analysis graph includes: obtaining the attack phase, time series and propagation characteristics of each node and each edge in the source tracing analysis graph; obtaining the attack propagation path and attack source of the network attack based on the attack phase, time series and propagation characteristics of each node and each edge in the source tracing analysis graph.

[0013] Optionally, obtaining the attack stage of each node and edge in the source tracing analysis graph includes: obtaining the techniques and tactics of each node and edge in the source tracing analysis graph; according to the techniques and tactics of each node and edge in the source tracing analysis graph, combined with a preset attack stage framework built based on ATT&CK, obtaining the attack stage of each node and edge in the source tracing analysis graph.

[0014] In a second aspect, the present invention provides a network attack tracing system based on an interpretable graph neural network, comprising: a graph construction module, used to obtain attack tracing log data of the network within a preset time period, and construct a tracing graph at each moment within the preset time period based on the attack tracing log data; a graph detection module, used to obtain feature vectors of each tracing graph and select a tracing graph with an abnormal detection result as an abnormality through a pre-trained anomaly detection model based on a graph neural network, to obtain a number of abnormal tracing graphs; a graph interpretation module, used to obtain key nodes and key edges of each abnormal tracing graph through an interpretable method, and obtain the minimum abnormal tracing subgraph of each abnormal tracing graph based on the key nodes and key edges; a tracing module, used to combine the minimum abnormal tracing subgraph of each abnormal tracing graph to obtain a tracing analysis graph, and obtain the attack propagation path and attack source of the network attack based on the tracing analysis graph.

[0015] Optionally, the obtaining of key nodes and key edges of each anomaly traceability graph through an interpretable method includes: obtaining the feature weights of each node and each edge in each anomaly traceability graph by using the GNNExplainer interpreter; combining a preset node feature weight threshold, taking nodes whose feature weights are not less than the node feature weight threshold as key nodes, and obtaining the key nodes of each anomaly traceability graph; combining a preset edge feature weight threshold, taking edges whose feature weights are not less than the edge feature weight threshold as key edges, and obtaining the key edges of each anomaly traceability graph.

[0016] Optionally, obtaining the attack propagation path and attack source of the network attack based on the source tracing analysis graph includes: obtaining the attack phase, time series and propagation characteristics of each node and each edge in the source tracing analysis graph; obtaining the attack propagation path and attack source of the network attack based on the attack phase, time series and propagation characteristics of each node and each edge in the source tracing analysis graph.

[0017] According to a third aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned network attack tracing method based on an interpretable graph neural network when executing the computer program.

[0018] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the network attack tracing method based on an interpretable graph neural network are implemented.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] The network attack tracing method based on the interpretable graph neural network of the present invention is based on the current situation that network attacks are usually complex and nonlinear. A tracing graph is constructed according to the attack tracing log data, and then the feature vectors of each tracing graph are obtained and several abnormal tracing graphs are obtained through the pre-trained abnormal detection model based on the graph neural network. The relationship between different nodes is captured by the graph neural network, which can better model the attack path and improve the accuracy and completeness of tracing. In addition, the key nodes and key edges of each abnormal tracing graph are obtained by an interpretable method, and the minimum abnormal tracing subgraph of each abnormal tracing graph is obtained according to the key nodes and key edges. The key nodes and key edges are presented by generating the minimum abnormal tracing subgraph, which is conducive to improving the reliability of tracing and fully understanding the tracing details. Finally, by accurately tracing the attack path, it can help the security team quickly locate the attack source and attack path, and then take targeted defense measures to minimize the loss, improve the tracing accuracy and security decision-making efficiency. This method combines interpretable methods with graph neural networks, which can improve the transparency of attack tracing, help researchers better understand attack paths, identify attack sources, and analyze attack propagation processes, and provide a basis for attack tracing optimization. This combination can make graph neural networks not just a black box model, but a tool that can explain its decision-making process, thereby providing more insights in applications in the field of network security. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a flow chart of a network attack tracing method based on an interpretable graph neural network according to an embodiment of the present invention.

[0022] Figure 2 This is a structural block diagram of a network attack tracing system based on an interpretable graph neural network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0026] See also Figure 1 In one embodiment of the present invention, a network attack tracing method based on an interpretable graph neural network is provided, which can accurately trace the attack path under the current situation of increasingly complex network attacks, and solve the lack of interpretability caused by the black box problem of traditional graph neural network models, and effectively reduce the losses caused by network attacks.

[0027] Specifically, the network attack tracing method based on the interpretable graph neural network of the present invention includes the following steps:

[0028] S1: Obtain attack source tracing log data of the network within a preset time period, and construct a tracing graph at each moment within the preset time period based on the attack source tracing log data.

[0029] S2: Obtain the feature vector of each traceability graph and select the traceability graph whose anomaly detection result is abnormal through the pre-trained anomaly detection model based on graph neural network to obtain several abnormal traceability graphs.

[0030] S3: Obtain key nodes and key edges of each anomaly traceability graph through an interpretable method, and obtain the minimum anomaly traceability subgraph of each anomaly traceability graph based on the key nodes and key edges.

[0031] S4: Combine the minimum anomaly tracing subgraphs of each anomaly tracing graph to obtain a tracing analysis graph, and obtain the attack propagation path and attack source of the network attack based on the tracing analysis graph.

[0032] The network attack tracing method based on the interpretable graph neural network of the present invention is based on the current situation that network attacks are usually complex and nonlinear. A tracing graph is constructed according to the attack tracing log data, and then the feature vectors of each tracing graph are obtained and several abnormal tracing graphs are obtained through the pre-trained abnormal detection model based on the graph neural network. The relationship between different nodes is captured by the graph neural network, which can better model the attack path and improve the accuracy and completeness of tracing. In addition, the key nodes and key edges of each abnormal tracing graph are obtained by an interpretable method, and the minimum abnormal tracing subgraph of each abnormal tracing graph is obtained according to the key nodes and key edges. The key nodes and key edges are presented by generating the minimum abnormal tracing subgraph, which is conducive to improving the reliability of tracing and fully understanding the tracing details. Finally, by accurately tracing the attack path, it can help the security team quickly locate the attack source and attack path, and then take targeted defense measures to minimize the loss, improve the tracing accuracy and security decision-making efficiency. This method combines interpretable methods with graph neural networks, which can improve the transparency of attack tracing, help researchers better understand attack paths, identify attack sources, and analyze attack propagation processes, and provide a basis for attack tracing optimization. This combination can make graph neural networks not just a black box model, but a tool that can explain its decision-making process, thereby providing more insights in applications in the field of network security.

[0033] In a possible implementation, constructing a tracing graph at each moment within a preset time period based on the attack tracing log data includes: constructing a basic graph with entities in the network as nodes and relationships between entities as edges; setting a moment at regular intervals within the preset time period, and obtaining node features of the nodes and edge features of the edges at each moment based on the attack tracing log data; and adding the node features of the nodes and edge features of the edges at each moment to the basic graph, respectively, to obtain the tracing graph at each moment within the preset time period.

[0034] Explanatory, the entities and relationships in the attack tracing log data are modeled as a basic graph. Entities generally include users, devices, hosts, and servers, and each entity is a node. If there is an entity relationship between two entities, such as network traffic, inter-device communication, file transfer, system call, and access records, then an edge is connected between the two entities. Then, based on the basic graph, the node features and edge features are combined to form the final tracing graph. Among them, node features generally include IP addresses, user names, and device attributes, and edge features generally include timestamps, login and access frequencies, call types, file types, and traffic sizes.

[0035] Specifically, in this implementation, the distributed environment traceability diagram construction tool (SPADE) is used to realize the construction of the traceability diagram. SPADE is a mature traceability diagram construction tool that supports multiple types of data and can collect audit logs to automatically generate traceability diagrams. Using SPADE to construct the traceability diagram G based on the system log<V,E> Among them, V = {user, device, host, server}, E = {network traffic, inter-device communication, file transfer, system call, access record}, ​​node feature X v ={IP address, user name, device attribute}, edge feature E vv = {timestamp, login and access frequency, call type, file type, traffic size}. The attacker's activities will flow in a certain pattern in the traceability graph, which is convenient for subsequent traceability. For example, an attacker may enter a system by exploiting a vulnerability on a device and interact with other devices through the system. These behaviors and paths are represented as node relationships in the traceability graph.

[0036] Optionally, the preset time period is: the time when the attack event is discovered is the starting time, the time two weeks before the starting time is the ending time, and the time period between the starting time and the ending time is the preset time period; the certain time is half an hour.

[0037] Explanatory, the time point when the attack event is discovered is used as the starting time, and the time window is set to two weeks. Log data is intercepted from two weeks before the attack event is discovered. The time range is from 0:00 to 24:00 every day, and each day is divided into 48 half-hour time periods. Each half-hour period represents a timestamp, forming a continuous time series. The attack source tracing log data is divided into multiple timestamp traceability graphs, and the divided traceability graphs are stored to facilitate subsequent feature extraction and traceability reasoning.

[0038] It should be noted that the above-mentioned preset time period and certain time are not strictly limited, and can be adaptively designed based on experience in specific applications. Moreover, in special scenarios, only one traceability graph can be constructed.

[0039] In a possible implementation, obtaining the feature vector of each provenance graph includes: obtaining the feature vector of the provenance graph through a graph convolutional neural network.

[0040] Explanatory, obtaining the feature vectors of each traceability graph can be understood as graph representation learning. There are generally two methods for graph representation learning: graph embedding technology and graph neural network. Graph embedding technology is a method of mapping graph data structure to low-dimensional vector space, such as random walks, etc. Graph neural network not only learns the low-dimensional representation of the graph, but also learns the local and global structure of the graph, and updates the feature representation of the node through information propagation and aggregation. Therefore, in this embodiment, a graph convolutional neural network is used to perform graph representation learning and extract graph features.

[0041] Specifically, we use graph convolutional neural networks to learn graph representations of traceability graphs and encode node and edge information into low-dimensional vector representations. First, we extract node features:

[0042]

[0043] Among them, H (l+1) is the feature expression of the l+1th layer node, H (l) is the feature expression of the l-th layer node, W (l) is the weight matrix of the l-th layer nodes, is the adjacency matrix of the normalized nodes.

[0044] Convert the traceability graph G in the form of a point graph into a traceability graph G in the form of a line graph link , regard the edges of the traceability graph G as the traceability graph G link If two edges in the traceability graph G have common nodes, then convert them to the traceability graph G link The two nodes after that have a connecting edge, for the traceability graph G link Performing point feature extraction means extracting edge features from the traceability graph G:

[0045]

[0046] Among them, H (l+1) is the feature expression of the l+1th layer edge, is the characteristic expression of the l-th layer edge, is the weight matrix of the l-th layer edge, is the normalized adjacency matrix.

[0047] In a possible implementation, during the pre-training process of the anomaly detection model based on the graph neural network, the traceability graph is first labeled. If it is a normal traceability graph, the label is set to 0; if it is an abnormal traceability graph, the label is set to 1, and then it is input into the anomaly detection model for training. In the pre-training stage, a large amount of data is input for training to help the anomaly detection model capture and learn complex attack patterns and relationships.

[0048] In the actual application stage, the feature vector of the traceability graph of the attack traceability log data is input into the pre-trained anomaly detection model for classification prediction. Specifically, first, through preprocessing, the real-time collected attack traceability log data is converted into a traceability graph according to the set time window and rules, and the feature vector is extracted; then, the feature vector is input into the trained anomaly detection model for reasoning, and the anomaly detection model classifies the traceability graph based on the feature vector and outputs the detection result of whether the traceability graph is abnormal.

[0049] In a possible implementation, the method of obtaining the key nodes and key edges of each anomaly traceability graph through an interpretable method includes: obtaining the feature weights of each node and each edge in each anomaly traceability graph by using the GNNExplainer interpreter; combining a preset node feature weight threshold, taking nodes whose feature weights are not less than the node feature weight threshold as key nodes, and obtaining the key nodes of each anomaly traceability graph; combining a preset edge feature weight threshold, taking edges whose feature weights are not less than the edge feature weight threshold as key edges, and obtaining the key edges of each anomaly traceability graph.

[0050] Explanatory, graph neural network, as a deep learning method widely used in graph data analysis in recent years, has achieved remarkable results in tasks such as node classification and link prediction. Graph neural network can learn and automatically extract attack patterns from a large amount of historical data, adapt to complex attack behaviors, and provide efficient traceability capabilities. However, most existing graph neural network models are "black box" models, lack sufficient interpretability, and have low trust in their prediction results, making it difficult for security personnel to understand attack paths and traceability decisions.

[0051] The interpretability of the model focuses on revealing how it learns the importance of nodes in the graph through interactions and information transfer between nodes, thereby identifying the nodes and edges that have the greatest impact on the prediction results. At present, the mainstream interpretability methods mainly include gradient-based and back-propagation methods and random perturbation methods represented by the GNNExplainer interpreter. In order to be closely integrated with the prediction model, in this implementation, the GNNExplainer interpreter is used to interpret the model prediction results.

[0052] GNNExplainer is a method for explaining graph neural networks. It learns continuous masks of edges or features through an optimization process to explain the model's predictive behavior. For the trained graph neural network model and its prediction results, GNNExplainer uses mask learning to identify key subgraph structures and feature subsets in the graph to maximize the mutual information between the model prediction and the input graph. This method can simultaneously reveal the network structure and node and edge features that play a key role in prediction, thereby providing a deep understanding of the model's decision-making process.

[0053] Specifically, in this implementation, the GNNExplainer interpreter is used to interpret the anomaly traceability graph, revealing the key security events or attack steps behind the model prediction results, and providing support for accurate attack tracing. The input of the GNNExplainer interpreter is the anomaly traceability graph, and the output is a binary mask that highlights the key nodes and edges, that is, the feature weights that represent the importance of nodes and edges. The GNNExplainer interpreter finds the optimal subgraph G by maximizing the mutual information (MI). s and the corresponding feature X s , while retaining the original prediction information, making the model prediction more interpretable, where:

[0054] MI(Y,(G S ,X S ))=H(Y)-H(Y|G=G S ,X=X S )

[0055] Among them, H(Y) is the total entropy of the predicted label Y, H(Y|G=G S ,X=X S ) is the conditional entropy under given subgraphs and features. The trained anomaly detection model H(Y) is fixed, so max[MI(Y,(G S ,X S ))] is transformed into the problem of finding min[H(Y|G=G S ,X=X S )] problem.

[0056] After interpreting each anomaly traceability graph using the GNNExplainer interpreter, the feature weights of each node and each edge in each anomaly traceability graph are obtained, and then the key nodes and key edges related to the prediction results are obtained according to the pre-selected node feature weight threshold. θ Set the edge feature weight threshold E to 0.8 to screen the key nodes that have a significant impact on the abnormal traceability graph identification. θ The edge weight indicates the importance of the edge connecting the nodes in the information transmission process, reflecting the contribution of the interaction between nodes to the model output. If the feature weight of an edge is greater than or equal to 0.7, then the edge will be regarded as a key edge that has a greater impact on the model prediction results.

[0057] Exemplarily, in order to further analyze and visualize the key edges and key nodes, the key edges and key nodes may be highlighted in the anomaly tracing graph.

[0058] Finally, based on the selected key nodes and key edges, the minimum anomaly tracing subgraph G is generated.trace , visually displays the most influential graph structure part of the prediction result, helps to locate the attack path and the key interaction chain, and provides strong data support for subsequent security response and defense measures.

[0059] In a possible implementation manner, obtaining the attack propagation path and the attack source of the cyber attack according to the traceability analysis graph includes: obtaining the attack stage, time series, and propagation characteristics of each node and each edge in the traceability analysis graph; and obtaining the attack propagation path and the attack source of the cyber attack according to the attack stage, time series, and propagation characteristics of each node and each edge in the traceability analysis graph.

[0060] Explanatory, based on the minimum abnormal traceability sub-graph of the foregoing abnormal traceability graphs, combined with multi-dimensional information such as attack stage, time series, and propagation characteristics, to assist in accurately identifying the attack propagation path and the attack source.

[0061] Specifically, during the processing of several traceability graphs, multiple abnormal traceability graphs may be identified. After subsequent feature extraction, prediction, explanation, and visualization display steps, multiple minimum abnormal traceability sub-graphs G trace1 , G trace2 , G trace3 , …, G traceN are obtained. The multiple minimum abnormal traceability sub-graphs are spliced into an overall graph to obtain the traceability analysis graph G trace-f .

[0062] Explanatory, by determining the attack stage of each edge and node in the traceability analysis graph, conduct attack traceability from the dimension of attack technical means. Generally, attackers gradually advance to higher-level attack stages over time. The time series can determine the time sequence of the nodes and edges in the traceability analysis graph being attacked, clarify the time span and key behaviors of each attack stage, and conduct attack traceability from the dimension of attack time. Combining with the equipment professional knowledge of specific industries, there is a fixed information flow direction between different devices, and conduct attack traceability from the level of the attack path propagation direction. Through the mutual verification of these three dimensions of attack stage, time series, and propagation characteristics, not only the accuracy of traceability is improved, but also complex situations such as multiple attack chains can be processed and analyzed.

[0063] Exemplarily, in a graph traversal manner, compare the attributes of two nodes A and B in the traceability analysis graph in sequence. If the attack stage of A, A. phase < the attack stage of B, B. phase and the attack time of A, A. time < the attack time of B, B. time , then A is the parent node of B.

[0064] In a possible implementation, obtaining the attack stage of each node and each edge in the source tracing analysis graph includes: obtaining the techniques and tactics of each node and each edge in the source tracing analysis graph; according to the techniques and tactics of each node and each edge in the source tracing analysis graph, combined with a preset attack stage framework built based on ATT&CK, obtaining the attack stage of each node and each edge in the source tracing analysis graph.

[0065] Explanatory, ATT&CK is a knowledge-based public library created by MITRE to describe attacker tactics and techniques. It is a framework that describes the techniques used in each stage of an attack from the attacker's perspective. It provides a variety of different "techniques" and "tactics" to make the expression of attack methods have a consistent standard. "Tactics" are behaviors that support strategic goals, and "techniques" are possible methods of executing tactics. The attack phase framework built based on ATT&CK can determine which stage of the attack chain the attacker's attack behavior belongs to, so as to determine the order of the attack behavior in the attack chain and provide data support for the attack stage for tracing.

[0066] Specifically, the ATT&CK framework includes 11 attack steps, namely: initial penetration, code execution, permission maintenance, detection escape, asset discovery, lateral movement, environmental search, command control, response suppression, compromise process and impact. The attack phase framework built based on ATT&CK has a clear list of techniques and tactics for each phase. For example, initial penetration includes open application vulnerability exploitation, external remote service exploitation, email attack, etc., and code execution includes script execution, file injection, user execution, etc. Then, the techniques and tactics based on the attacker's behavior are matched with the attack phase in the framework to determine the position of the attack behavior in the attack chain.

[0067] In general, the present invention is based on a network attack tracing method based on an interpretable graph neural network. Aiming at the problem of poor reliability of network security event model tracing, based on the interpretability of graph neural networks, an attack tracing method based on graph neural networks and GNNExplainer interpreter is proposed. The interpretable graph neural network attack tracing consists of two parts: an anomaly detection model based on graph neural networks and a GNNExplainer interpreter based on mask learning. The former predicts whether it is an abnormal tracing graph, and the latter generates network structure-based interpretations and feature-based interpretations through training. Combine multi-dimensional information such as attack stages, time series, and propagation characteristics to accurately identify attack propagation paths and attack sources.

[0068] The following are device embodiments of the present invention, which can be used to implement the method embodiments of the present invention. For details not disclosed in the device embodiments, please refer to the method embodiments of the present invention.

[0069] See also Figure 2In another embodiment of the present invention, a network attack tracing system based on an interpretable graph neural network is provided, which can be used to implement the above-mentioned network attack tracing method based on an interpretable graph neural network. Specifically, the network attack tracing system based on an interpretable graph neural network includes a graph construction module, a graph detection module, a graph interpretation module and a tracing module.

[0070] Among them, the graph construction module is used to obtain the attack tracing log data of the network within a preset time period, and construct the tracing graph at each moment within the preset time period based on the attack tracing log data; the graph detection module is used to obtain the feature vector of each tracing graph and select the tracing graph with abnormal detection results to obtain several abnormal tracing graphs through a pre-trained anomaly detection model based on a graph neural network; the graph interpretation module is used to obtain the key nodes and key edges of each abnormal tracing graph through an interpretable method, and obtain the minimum abnormal tracing subgraph of each abnormal tracing graph based on the key nodes and key edges; the tracing module is used to combine the minimum abnormal tracing subgraph of each abnormal tracing graph to obtain a tracing analysis graph, and obtain the attack propagation path and attack source of the network attack based on the tracing analysis graph.

[0071] In a possible implementation, constructing a tracing graph at each moment within a preset time period based on the attack tracing log data includes: constructing a basic graph with entities in the network as nodes and relationships between entities as edges; setting a moment at regular intervals within the preset time period, and obtaining node features of the nodes and edge features of the edges at each moment based on the attack tracing log data; and adding the node features of the nodes and edge features of the edges at each moment to the basic graph, respectively, to obtain the tracing graph at each moment within the preset time period.

[0072] In a possible implementation, the preset time period is: the time when the attack event is discovered is the starting time, the time two weeks before the starting time is the ending time, and the time period between the starting time and the ending time is the preset time period; the certain time is half an hour; the node features include one or more of the following: IP address, user name and device attributes; the edge features include one or more of the following: timestamp, login and access frequency, call type, file type and traffic size.

[0073] In a possible implementation, obtaining the feature vector of each provenance graph includes: obtaining the feature vector of the provenance graph through a graph convolutional neural network.

[0074] In a possible implementation, the method of obtaining the key nodes and key edges of each anomaly traceability graph through an interpretable method includes: obtaining the feature weights of each node and each edge in each anomaly traceability graph by using the GNNExplainer interpreter; combining a preset node feature weight threshold, taking nodes whose feature weights are not less than the node feature weight threshold as key nodes, and obtaining the key nodes of each anomaly traceability graph; combining a preset edge feature weight threshold, taking edges whose feature weights are not less than the edge feature weight threshold as key edges, and obtaining the key edges of each anomaly traceability graph.

[0075] In a possible implementation, obtaining the attack propagation path and attack source of the network attack based on the source tracing analysis graph includes: obtaining the attack phase, time series and propagation characteristics of each node and each edge in the source tracing analysis graph; obtaining the attack propagation path and attack source of the network attack based on the attack phase, time series and propagation characteristics of each node and each edge in the source tracing analysis graph.

[0076] In a possible implementation, obtaining the attack stage of each node and each edge in the source tracing analysis graph includes: obtaining the techniques and tactics of each node and each edge in the source tracing analysis graph; according to the techniques and tactics of each node and each edge in the source tracing analysis graph, combined with a preset attack stage framework built based on ATT&CK, obtaining the attack stage of each node and each edge in the source tracing analysis graph.

[0077] All relevant contents of each step involved in the aforementioned embodiment of the network attack tracing method based on interpretable graph neural network can be referred to the functional description of the functional modules corresponding to the network attack tracing system based on interpretable graph neural network in the embodiment of the present invention, and will not be repeated here.

[0078] The division of modules in the embodiments of the present invention is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present invention may be integrated into one processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0079] In another embodiment of the present invention, a computer device is provided, the computer device including a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in a computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the network attack tracing method based on an interpretable graph neural network.

[0080] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in a computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by a processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the network attack tracing method based on an interpretable graph neural network in the above embodiment.

[0081] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0082] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0083] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0084] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A network attack tracing method based on an interpretable graph neural network, characterized in that: include: Obtain attack source tracing log data of the network within a preset time period, and construct a tracing graph at each moment within the preset time period based on the attack source tracing log data; Obtain the feature vector of each traceability graph and select the traceability graph with abnormal detection results through the pre-trained graph neural network-based anomaly detection model to obtain several abnormal traceability graphs; Obtain key nodes and key edges of each anomaly traceability graph through an interpretable method, and obtain the minimum anomaly traceability subgraph of each anomaly traceability graph based on the key nodes and key edges; The minimum anomaly tracing subgraphs of each anomaly tracing graph are combined to obtain a tracing analysis graph, and the attack propagation path and attack source of the network attack are obtained based on the tracing analysis graph.

2. The network attack tracing method based on interpretable graph neural network according to claim 1 is characterized in that: The constructing of a tracing graph at each moment in a preset time period according to the attack tracing log data includes: The basic graph is constructed with entities in the network as nodes and relationships between entities as edges; A time is set at regular intervals within a preset time period, and node features of nodes and edge features of edges at each time are obtained according to the attack tracing log data; The node features of the nodes and the edge features of the edges at each moment are added to the basic graph respectively to obtain the traceability graph at each moment in the preset time period.

3. The network attack tracing method based on interpretable graph neural network according to claim 2 is characterized in that: The preset time period is: the time when the attack event is discovered is the starting time, the time two weeks before the starting time is the ending time, and the time period between the starting time and the ending time is the preset time period; The certain time is half an hour; The node characteristics include one or more of the following: IP address, user name and device attributes; The edge features include one or more of the following: timestamp, login and access frequency, call type, file type and traffic size.

4. The network attack tracing method based on interpretable graph neural network according to claim 1 is characterized in that: The step of obtaining the characteristic vectors of each traceability graph comprises: The feature vector of the traceability graph is obtained through the graph convolutional neural network.

5. The network attack tracing method based on interpretable graph neural network according to claim 1 is characterized in that: The key nodes and key edges of each anomaly traceability graph obtained by an interpretable method include: The feature weights of each node and each edge in each anomaly traceability graph are obtained by using the GNNExplainer interpreter; Combined with the preset node feature weight threshold, nodes whose feature weights are not less than the node feature weight threshold are taken as key nodes to obtain the key nodes of each anomaly traceability graph; Combined with the preset edge feature weight threshold, the edges whose feature weights are not less than the edge feature weight threshold are taken as key edges, and the key edges of each anomaly traceability graph are obtained.

6. The network attack tracing method based on interpretable graph neural network according to claim 1 is characterized in that: The attack propagation path and attack source of the network attack obtained according to the source tracing analysis diagram include: Obtain the attack phase, time series, and propagation characteristics of each node and edge in the traceability analysis graph; According to the attack phase, time series and propagation characteristics of each node and edge in the source tracing analysis graph, the attack propagation path and attack source of the network attack are obtained.

7. The network attack tracing method based on interpretable graph neural network according to claim 6 is characterized in that: The attack phase of obtaining each node and each edge in the traceability analysis graph includes: Obtain the techniques and tactics of each node and edge in the traceability analysis diagram; According to the techniques and tactics of each node and edge in the source tracing analysis diagram, combined with the preset attack stage framework built based on ATT&CK, the attack stage of each node and edge in the source tracing analysis diagram is obtained.

8. A network attack tracing system based on an interpretable graph neural network is characterized by: include: A graph construction module is used to obtain attack source tracing log data of the network within a preset time period, and to construct a tracing graph at each moment within the preset time period based on the attack source tracing log data; The graph detection module is used to obtain the feature vector of each traceability graph and select the traceability graph with abnormal detection results through the pre-trained graph neural network-based anomaly detection model to obtain several abnormal traceability graphs; A graph interpretation module, used to obtain key nodes and key edges of each anomaly traceability graph through an interpretable method, and obtain a minimum anomaly traceability subgraph of each anomaly traceability graph based on the key nodes and key edges; The tracing module is used to combine the minimum anomaly tracing subgraphs of each anomaly tracing graph to obtain a tracing analysis graph, and obtain the attack propagation path and attack source of the network attack based on the tracing analysis graph.

9. The network attack tracing system based on interpretable graph neural network according to claim 8 is characterized in that: The key nodes and key edges of each anomaly traceability graph obtained by an interpretable method include: The feature weights of each node and each edge in each anomaly traceability graph are obtained by using the GNNExplainer interpreter; Combined with the preset node feature weight threshold, nodes whose feature weights are not less than the node feature weight threshold are taken as key nodes to obtain the key nodes of each anomaly traceability graph; Combined with the preset edge feature weight threshold, the edges whose feature weights are not less than the edge feature weight threshold are taken as key edges, and the key edges of each anomaly traceability graph are obtained.

10. The network attack tracing system based on interpretable graph neural network according to claim 8 is characterized in that: The attack propagation path and attack source of the network attack obtained according to the source tracing analysis diagram include: Obtain the attack phase, time series, and propagation characteristics of each node and edge in the traceability analysis graph; According to the attack phase, time series and propagation characteristics of each node and edge in the source tracing analysis graph, the attack propagation path and attack source of the network attack are obtained.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the network attack tracing method based on an interpretable graph neural network as described in any one of claims 1 to 7 are implemented.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the network attack tracing method based on an interpretable graph neural network as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Internet of Things data security collaborative defense system based on edge computing

    CN120263548A

  • A method for detecting abnormalities in agricultural product quality traceability

    CN120410342B

  • Redundant fragment calculation verification and edge loss recovery error correction method for attack traceability graph construction

    CN121509099A