Negative sample enhanced APT attack detection method based on graph structure learning

Through the negative sample enhancement method based on graph structure learning, the problem of insufficient data set compatibility and edge information utilization in APT attack detection is solved, and efficient APT attack detection is realized, which is suitable for unified formatting processing of multi-source heterogeneous logs and dynamic anti-negative sampling, improving the accuracy and efficiency of detection.

CN120301664AActive Publication Date: 2025-07-11XIDIAN UNIV

Patent Information

Application Number
CN202510528585.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-11
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

There are problems in the existing APT attack detection technology that insufficient data set compatibility, insufficient edge information utilization and scarce negative samples lead to poor detection results. Traditional graph neural network models are difficult to effectively capture the key causal relationships in the attack chain, and the compatibility of multi-source heterogeneous logs leads to insufficient model generalization capabilities.

Method used

A negative sample enhancement method based on graph structure learning was designed. By constructing a traceability graph and dividing it into snapshots by timestamp, the graph neural network cyclic RNN of the redesigned encoder and decoder is used for iterative training, combining dynamic adversarial negative sampling and adaptive loss functions, dynamically adjusting the generator parameters, merging edge information and node information, and building a time-sensitive dynamic metric modeling mechanism, solving the compatibility problem of multi-source logs.

Benefits of technology

It significantly improves the accuracy and efficiency of APT attack detection, can adapt to multi-scenario deployment, reduces false positive rate, improves the generalization ability of the model and its sensitivity to attack patterns, and can identify time-related attack patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301664A_ABST
    Figure CN120301664A_ABST
Patent Text Reader

Abstract

The invention provides a negative sample enhanced APT attack detection method based on graph structure learning, and mainly solves the problem of poor detection effect caused by insufficient data set compatibility, side information utilization and data set negative samples in the prior art. According to the scheme, the method comprises the following steps: 1) acquiring a heterogeneous data set, and constructing a visual traceability graph; 2) preprocessing the traceability graph, dividing the traceability graph into snapshots according to timestamps, and constructing a training set and a test set; 3) designing an encoder and a decoder of the drawing neural network, taking the time window, the node features, the adjacency matrix and the edge features as encoder input, and generating a reconstructed adjacency matrix through the decoder; 4) designing a loss function, and performing RNN training on the graph neural network; and 5) inputting to-be-detected data into the trained model, and identifying abnormal attack traffic to complete detection. According to the method, the accuracy and robustness of APT attack detection can be effectively improved, and the method can be used for development and deployment of APT attack detection and defense systems in the field of network security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer networks, and further relates to network attack detection technology. Specifically, it is a negative sample enhanced APT attack detection method based on graph structure learning, which can be used for security detection of the internal networks of enterprises and institutions. Background Art

[0002] In the information age, information security has become increasingly important, and the cyberspace is regarded as the fifth battlefield. With the development of the Internet and artificial intelligence technologies, while people enjoy the convenience brought by the Internet, they also face the threats posed by network attacks. In 2023, the 404 Advanced Threat Intelligence Team of Knownsec found that more than 400 assets in China were successfully attacked and controlled by APT groups, and China remains one of the main targets of advanced persistent threats. APT attacks are network attack and intrusion behaviors launched by hackers to steal core data, and they are a kind of malicious latent threat with long-term planning. When deeply analyzing the current network security situation, it is not difficult to find that many advanced persistent threat APT attack events often closely follow current affairs hotspots, highlighting their high sensitivity and timeliness. The APT groups behind have the ability to independently discover and exploit 0day vulnerabilities, which makes their attack means more cunning and difficult to prevent. The types of security vulnerabilities exploited by these APT groups are extremely extensive, including but not limited to key areas such as browsers, operating systems, processors, and mail services, showing the comprehensiveness and complexity of their attack means.

[0003] Existing APT attack detections face multiple bottlenecks. When traditional graph neural network models process the traceability graphs in security scenarios, there are generally defects in the insufficient utilization of edge information. For example, edge features such as process call relationships and network communication types are only used as auxiliary information of the topological structure and are not fully integrated into the node representation learning, resulting in the difficulty of effectively capturing the key causal relationships in the attack chain. At the same time, the serious shortage of negative samples unique to the security field severely restricts the model performance. The proportion of APT attack samples in real scenarios is less than one in ten thousand. Due to the excessive dependence on the negative sample distribution, the traditional cross-entropy loss function is prone to cause the model to overfit to normal behaviors and be less sensitive to attack patterns. In addition, the compatibility problem of multi-source heterogeneous logs has long existed. The structural differences in mainstream data sets at levels such as entity definition, time granularity, and relationship dimension make it difficult for existing detection systems to achieve cross-platform generalization, and often need to customize models for a single data source, severely weakening the actual deployment efficiency. Therefore, we urgently need a highly automated and responsive APT attack detection method.

[0004] In recent years, significant achievements have been made in the development of graph neural networks in key fields such as drug research and development, protein structure determination, recommendation systems, and traffic flow prediction. By graphically modeling data in real-world scenarios, graph neural networks can learn a highly abstract representation for each node in the graph. This representation is further applied to downstream tasks such as node classification, link prediction, and graph classification, showing extremely high accuracy. These advancements fully demonstrate the broad application prospects and great potential of graph neural networks in multiple fields. However, there are also some problems with graph neural network models in terms of secure applications. First, there is an obvious gap between the attack link distributions predicted by most systems (links that may be attacked predicted through existing datasets) and the actual attack links. Second, conventional encoders such as GCN and GraphSage cannot incorporate edge features (temporal attributes other than source and destination addresses) into the framework. Third, most systems abstract static graphs from logs, thus ignoring the important temporal dynamics between different times.

[0005] For example, the patent application with the publication number CN119232465A and the title "An APT Attack Detection Method Based on Traceability Graph Behavior Information" proposes to detect attacks by constructing a traceability graph and extracting behavioral attribute structure information, uses GraphSAGE and Residual Gated Graph ConvNets models to classify edge features, and combines ER random graphs to initialize the behavioral structure graph BSG and pooling strategies to generate edge features. This method realizes fine-grained anomaly detection through incremental training of multiple sub-models and outputs the behavioral structure graph of abnormal edges to enhance interpretability. However, the defect is that this patent does not design an adaptation mechanism for heterogeneous datasets and does not solve the semantic difference problem of multi-source logs (such as Windows event logs and Linux auditd logs). Although this patent extracts edge features through the behavioral structure graph BSG, its ER random graph initialization strategy does not introduce security domain knowledge (such as MITRE ATT&CK tactical tags), resulting in the weakening of the semantic association between edge features and real APT attack patterns. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies, and proposes a negative sample enhanced APT attack detection method based on graph structure learning to solve the problems of insufficient dataset compatibility, poor detection effect caused by insufficient utilization of edge information and insufficient negative samples in the dataset in the existing technology.

[0007] The technical idea of the present invention is as follows: First, for heterogeneous data sets such as LANL and DARPA TC, an independent parser is designed. Then, a PGTS (Provenance Graph to Snapshot) modeling method is designed, that is, the provenance graph is divided into snapshots according to timestamps. Then, a graph neural network recurrent RNN with a redesigned encoder and decoder is used for iterative training, and the trained network model is used to detect abnormal APT attack traffic to achieve attack detection.

[0008] To achieve the above object, the technical solution of the present invention includes the following:

[0009] (1) Obtain a security log data set from the public network, and use the network topology graph tool NetworkX to construct a visual provenance graph;

[0010] (2) Preprocess the provenance graph to obtain a training set, a validation set, and a test set;

[0011] (3) Construct an encoder of the graph neural network, take the time window, node features, adjacency matrix, and edge features as inputs, and obtain an encoded result vector:

[0012] (3a) Use the message passing layer to merge the edge features of neighboring nodes into the node embedding, and take the edge features and edge weights as inputs respectively;

[0013] (3b) Construct an edge feature processing layer EF and an edge weight processing layer EW, which are respectively used to obtain the aggregated feature value EF of the edge features containing the time window t t and the aggregated weight feature value EW t ;

[0014] (3c) Based on the information passing neural network, use three EWs and one EF to construct an encoder, and its structure includes the following: the first EW layer - the second EW layer - the first activation function layer - the weighted EW layer - the second activation function layer - the EF layer - the multi-layer perceptron MLP - the third activation function layer; the encoded result vector is output through the third activation function layer;

[0015] (4) Construct a decoder of the graph neural network, and generate a reconstructed adjacency matrix according to the following steps

[0016] (4a) Capture the community structure that the attacker may form by randomly sampling and aggregating the neighbor information of the nodes;

[0017] (4b) Obtain the aggregated node embedding according to the following formula

[0018]

[0019] Among them, S represents the identity matrix, and each row of this matrix is s neighborhoods extracted from the adjacency matrix A t ; Z t is the embedding representation matrix of the time window t;

[0020] (4c) Process through the multi-layer perceptron MLP and the non-linear activation function Softmax to generate a reconstructed adjacency matrix This matrix represents the probability that the corresponding edge is a malicious edge:

[0021]

[0022] (5) Design the loss function of the graph neural network, which is expressed as follows:

[0023]

[0024] Among them, is the average precision AP loss; β is a parameter specifically adjusted according to the dataset; is the decoder loss of the sampled edges calculated using the cross-entropy function CE;

[0025] (6) Use the loss function to train the encoder and decoder of the graph neural network with the recurrent neural network RNN to obtain an optimized graph neural network model;

[0026] (7) Input the data to be tested into the graph neural network model, and use the model to identify abnormal attack traffic to complete attack detection.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] First, since the present invention combines the traceability graph with the graph neural network, captures the complex associations in the APT attack through the traceability graph, and uses the graph neural network to process and analyze it, the APT attack can be treated as a graph, effectively extracting the complex patterns and features hidden in the graph structure, thereby significantly improving the APT attack detection effect.

[0029] Second, for datasets with different structures such as LANL and DARPA TC, the present invention designs a unified formatting and preprocessing method for multi-source heterogeneous logs. By constructing an entity relationship mapping rule library and a time series alignment algorithm, the logs of heterogeneous data sources are uniformly converted into a standardized traceability graph structure that can be recognized by the model, and divided into traceability graph slices according to appropriate time windows. At the same time, a hybrid pruning strategy based on node degree threshold and time decay factor is adopted to compress the dataset, remove redundant data, reduce the impact of noise, and effectively improve the detection efficiency. At the same time, it solves the problems of poor cross-data source compatibility and high dependence on manual rules in traditional solutions, can adapt to the deployment requirements of multiple scenarios such as finance, cloud platforms, and industrial Internet, and significantly improves the generalization ability of the detection model.

[0030] Third, the present invention specially designs a graph neural network architecture with edge-node joint encoding for APT attacks. Through a mechanism similar to the message passing neural network, edge information and node information are merged. The influence weights of edge features such as the number of process calls and file read / write permissions on adjacent nodes are dynamically calculated through the multi-head attention mechanism, and residual connections are used to achieve cross-layer propagation of edge features. Thus, the utilization degree of edge information in the graph is improved, and the problem of poor learning effect caused by the lack of negative sets in security problems in the current graph neural network learning is improved, and the accuracy of graph neural network learning is significantly improved.

[0031] Fourth, since the present invention proposes dynamic adversarial negative sampling and an adaptive loss function; based on the generative adversarial network (GAN), difficult negative examples with high confusion are synthesized, and the parameters of the generator are dynamically adjusted in combination with the meta-learning framework to make the negative sample distribution approximate the real attack pattern. Compared with the prior art, in the extremely imbalanced scenario with a positive-negative sample ratio of 1:10 5 the false alarm rate is reduced, and the problems of model overfitting or information loss caused by traditional oversampling / undersampling methods are solved.

[0032] Fifth, the present invention constructs a time series sensitive dynamic metric modeling mechanism, uses LSTM to predict the attack phase transition point, dynamically adjusts the time window length of the graph slice, and embeds relative timestamp encoding in the node features to achieve three-dimensional tensor fusion of topological features and spatio-temporal features. Compared with the existing static graph construction scheme, the detection accuracy of time-related attack patterns (such as zero-day attacks in the latency period and periodic data exfiltration) is improved, and the problem of loss of stage evolution features caused by fixed time windows in traditional methods is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart for implementing the method of the present invention;

[0034] Figure 2 is a schematic diagram of the overall architecture of the graph neural network model in the present invention;

[0035] Figure 3 This is a schematic diagram of the encoder structure of the graph neural network model in the present invention. Specific embodiments

[0036] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0037] Example 1: Refer to Figure 1 A negative sample enhanced APT attack detection method based on graph structure learning proposed by the present invention specifically includes the following steps:

[0038] Step 1) Obtain a security log dataset from the public network and use the network topology graph tool NetworkX to construct a visual traceability graph; in this embodiment, the obtained security log dataset preferably adopts the LANL dataset of the Los Alamos National Laboratory and the OpTC dataset of the Operationally Transparent Network.

[0039] Step 2) Preprocess the traceability graph to obtain a training set, a validation set, and a test set; in this embodiment, in this step, the traceability graph constructed in step 1) is preprocessed. Specifically, the LANL dataset is data-labeled through the content of its corresponding red team attack information to identify malicious data flows; then it is segmented by 30 - 60 minute snapshots, and all events sharing the same source node and target node are merged into one edge; all snapshots before the start of the red team attack are intercepted; each snapshot has an average of 4000 - 7000 edges, and 5% - 10% of the edges are discarded as the validation set, and the remaining snapshots are used for testing; for the OpTC dataset, node features X t and edge weights Wt are constructed, the snapshot size is set to 180 - 360 seconds, and 1,680 - 3360 snapshots in one week are obtained; 20 - 40% of the snapshots are used for testing, and the rest are used for training.

[0040] Step 3) Construct an encoder of the graph neural network, take the time window, node features, adjacency matrix, and edge features as inputs, and obtain an encoded result vector:

[0041] (3a) Use the message passing layer to merge the edge features of neighboring nodes into the node embedding, and take the edge features and edge weights as inputs respectively;

[0042] (3b) Construct an edge feature processing layer EF and an edge weight processing layer EW, which are respectively used to obtain the aggregated feature value EF t of the edge features in the t time window and the aggregated weight feature value EW t .

[0043] In this embodiment, the above-mentioned aggregated feature value EF t is specifically obtained according to the following formula:

[0044]

[0045] Among them, let \(V\) represent the set of nodes, \(u\in V\); \(v\) represents the neighbor node of \(u\); \(EF(X u ,F u,v ) represents the aggregated feature of node \(u\) in the time window \(t\); \(X u represents the node feature of \(u\); \(F u,v represents the edge feature between node \(u\) and \(v\); \(N(u)\) represents the neighborhood of \(u\); \(M θ represents the eigenvalue of the features transmitted by the aggregated neighborhood nodes, and \(\sigma\) represents the activation function.

[0046] The above aggregation weight value \(EW t , is obtained according to the following formula:

[0047]

[0048] Among them, \(EW(X, W)\) represents the aggregation weight value \(EW\) of the node feature \(X\) in the time window \(t\) t ; is the pair angle matrix, which is used to normalize the edge weight matrix \(W\); \(\Theta\) represents the trainable parameter; \(\sigma\) represents the activation function.

[0049] (3c) Based on the information transfer neural network, an encoder is constructed using three \(EW\)'s and one \(EF\), and its structure includes the following: the first \(EW\) layer - the second \(EW\) layer - the first activation function layer - the weighted \(EW\) layer - the second activation function layer - the \(EF\) layer - the multi-layer perceptron \(MLP\) - the third activation function layer; the encoded result vector is output through the third activation function layer;

[0050] Step 4) Construct the decoder of the graph neural network, and generate the reconstructed adjacency matrix according to the following steps

[0051] (4a) Capture the community structure that the attacker may form by randomly sampling and aggregating the neighbor information of the nodes;

[0052] (4b) Obtain the aggregated node embedding according to the following formula

[0053]

[0054] Among them, \(S\) represents the identity matrix, and each row of this matrix is \(s\) neighborhoods extracted from the adjacency matrix \(A t ; \(Z t is the embedding representation matrix of the time window \(t\);

[0055] (4c) Process through the multi-layer perceptron \(MLP\) and the non-linear activation function Softmax to generate the reconstructed adjacency matrix This matrix represents the probability that the corresponding edge is a malicious edge:

[0056]

[0057] Step 5) Design the loss function of the graph neural network, which is expressed as follows:

[0058]

[0059] Among them, is the average precision AP loss; β is a parameter specifically adjusted according to the dataset; is the decoder loss of the sampled edges calculated using the cross-entropy function CE.

[0060] The loss function designed in this embodiment in this step is specifically obtained according to the following steps:

[0061] (5a) Calculate the decoder loss L of the sampled edges using the cross-entropy function CE DEC :

[0062]

[0063] (5b) Introduce the average precision AP loss

[0064]

[0065] Among them, P t is the positive edge in G t nP t is the number of positive edges, e i is a positive edge, n is the number of all edges in G t e s is an arbitrary edge in G t y s is the edge label, I is the identification function, which outputs 1 / 0 when the parameter is true / false; l(e s ; e i ) = (m - (f(e i ) - f(e s ))) 2 is a smooth square loss approximating I(f(e s ) ≥ f(e i ), where f is the prediction model and m is the margin value;

[0066] (5c) Combine L ap and L DEC and assign different weights according to the characteristics of the dataset to obtain the final loss function:

[0067]

[0068] Step 6) Use the loss function to perform recurrent neural network (RNN) training on the encoder and decoder of the graph neural network to obtain an optimized graph neural network model. In this embodiment, the above-mentioned RNN training is implemented as follows:

[0069] (6a) The encoder receives the graph structure data at the current moment, including node features and adjacency relationships, and generates node embedding vectors;

[0070] (6b) Integrate the node embedding vectors generated at all historical moments through RNN, capture the dynamic evolution law through time series modeling, and output a temporally enhanced node embedding sequence;

[0071] (6c) Based on the node embedding sequence, the decoder reconstructs the corresponding adjacency matrix according to the node embedding at each moment, and calculates the decoder loss by comparing the reconstructed matrix with the real graph structure;

[0072] (6d) Use the loss function to calculate the gradients of the encoder, decoder, and RNN, and update the parameters through backpropagation to achieve optimization.

[0073] Step 7) Input the data to be tested into the graph neural network model, and use the model to identify abnormal attack traffic to complete attack detection.

[0074] Embodiment 2: The overall implementation process of the detection method proposed in this embodiment is the same as that in Embodiment 1. Now, with reference to the attached Figure 2 and 3 specific examples are given to further describe the implementation process of the present invention in detail:

[0075] Step 1. Construct a traceability graph

[0076] Obtain the LANL dataset, which is derived from five different sources within the enterprise internal computer network of Los Alamos National Laboratory. After 58 consecutive days of collection, ensure the integrity and continuity of the data; obtain the OpTC dataset, which is supported by the large-scale network hunting program CHASE of the Defense Advanced Research Projects Agency of the United States. The data contains approximately 1TB of compressed JSON-compatible format data from this evaluation.

[0077] Step 2. Preprocess the traceability graph, including the processing processes of the two datasets as follows:

[0078] For the LANL dataset: Data label the auth file through the content of the redteam file to identify malicious data flows. Then, in this embodiment, it is preferable to split the LANL-auth data by 1-hour (3,600 seconds) snapshots, and merge all events sharing the same source node and target node into one edge. Use the snapshots of the first 42 hours for model training. Each snapshot has an average of 7,957 edges, and 5% of the edges are discarded as the validation set. The remaining snapshots are used for testing;

[0079] For the OpTC dataset: By constructing node features Xt and edge weights Wt. Set the snapshot size to 360 seconds, which is much smaller than the default value of LANL because the number of nodes in OpTC (814) is much less than that in LANL (15,610), and the event frequency between a pair of hosts is much higher. Setting a larger snapshot size will merge too many events. Finally, in this embodiment, it is preferable to obtain 1,680 snapshots in one week. Use the snapshots of the first 5 days for training and the snapshots of the remaining 3 days for testing;

[0080] Step 3. Design a graph neural network encoder, with the input being the time window t, node features X t , adjacency matrix A t and edge features F t ; In the encoder designed in this embodiment, three EW layers and one EF layer are adopted, and the structure is as Figure 2 shown;

[0081] (3.1) Use the message passing layer to merge the edge features of neighboring nodes into the node embeddings;

[0082] (3.2) Respectively use the edge weights and edge features as the inputs of two types of layers EF and EW:

[0083] (3.2.1) EF layer design:

[0084] The design of EF is similar to that of GCN and can handle node features and edge weights. Given a node u ∈ V,

[0085]

[0086] where, X u represents the node features of u, F u ,v represents the edge features between the edge (u, v), N(u) represents the neighborhood of u, M θ aggregates the eigenvalue of the features transmitted by the neighboring nodes and can be simply called a multi-layer perceptron (MLP), and σ represents the activation function. The function output value is Ht, which is the node features at the time window t.

[0087] (3.2.3) EW layer design:

[0088] The input is the node feature X = (N, F), the weight matrix W = (N, N), and the scientific system parameter matrix Θ = (F, F'):

[0089]

[0090] Among them, is the pair angle matrix, used to normalize the edge weight matrix W; Θ represents the trainable parameter; σ represents the activation function, set to Tanh. The function outputs the updated node feature matrix (N, F').

[0091] Step 4. Design the decoder and loss function of the graph neural network:

[0092] (4.1) Design an aggregation operation that randomly samples from the neighborhood of each node as the decoder, used to update the node embedding:

[0093]

[0094] Among them, the input S is an identity matrix, and s neighborhoods are sampled from the adjacency matrix A for each row t ; the output is the node embedding

[0095] (4.2) This system uses the cross-entropy function (CE) to calculate the decoder loss (L DEC ) of the sampled edges:

[0096]

[0097] (4.3) Average Precision (AP) loss:

[0098] Under the condition of CE, L DEC can still optimize the graph neural network model under the precision metric. Precision is a metric more suitable for imbalanced security datasets, so the AP loss is introduced. Specifically, the loss of AP can be approximately written as the following formula:

[0099]

[0100] Among them, P t is the positive edge in G t , nP t is the number of positive edges, e i is a positive edge, n is the number of all edges in G t , e s is any edge in G t , y s is the edge label (y s= 1 indicates the positive side), I is the identification function, which outputs 1 / 0 when the parameter is true / false. l(e s ; e i ) = (m - (f(e i ) - f(e s ))) 2 is the smoothed squared loss approximating I(f(e s ) ≥ f(e i ))), where f is the prediction model (i.e., the graph neural network system) and m is the margin value. By rewriting Equation 13 as a finite sum of combination functions and performing SGD-style or Adam-style stochastic optimization, L ap can be minimized.

[0101] (4.4) Combines L ap and L DEC to construct a joint loss function and assigns them different weights according to the characteristics of the dataset (e.g., a highly imbalanced dataset requires a larger weight for L ap ); the final loss function is:

[0102]

[0103] where β is a parameter adjusted specifically according to the dataset. The output loss value loss is continuously reduced through RNN training to update the model.

[0104] Step 5. Graph Neural Network Recurrent (RNN) Training:

[0105] (5.1) Using X t and A t as inputs, uses the encoder ENC constructed in 3 to generate the node embedding Z′ t .

[0106] (5.2) Assuming T is the number of all observed snapshots, the RNN layer receives [Z0′, Z1′,..., Z T ′] and updates it to [Z0, Z1,..., Z T using the observed time dynamics.

[0107] [Z0, Z1,…, Z T = RNN([Z′0, Z′1,…, Z′ T )

[0108] (5.3) In the snapshot G t the decoder reconstructs the adjacency matrix using Z t generated by the encoder The decoding function may be different from σ(ZZ T ), and the decoder DEC is defined as:

[0109]

[0110] (5.4) During forward propagation, the encoder generates node embeddings [Z0, Z1, ..., Z T , and the decoder reconstructs from them Calculate the loss between the ground truth values [A0, A1, ..., A T .

[0111] (5.5) During backpropagation, the loss is used to calculate the gradients and [A0, A1, ..., A T . It is used to calculate the gradients and update the RNN, encoder, and decoder. During the testing process, after collecting the logs of the snapshot duration T + k (k > 0), Xt, At, and Ft, as well as the prior embeddings Z0, Z1, ......, Z T+k-1 are used to generate its embedding Z T+k .

[0112] Step 6. Testing of the graph neural network model:

[0113] Bring the trained model into the test sets of the LANL and OpTC datasets, calculate the false positives (FP) and false negatives (FN), F1 score, and AUC value, and identify abnormal traffic by predicting the values of the reconstructed adjacency matrix to achieve attack detection.

[0114] The present invention aims to systematically break through the technical barriers of existing APT attack detection. Aiming at the problem of insufficient utilization of side information, by designing an edge-node joint encoder, multimodal fusion of edge features such as process call frequencies and file read / write permissions with node semantics is performed, enabling the graph neural network to simultaneously analyze entity attributes and relationship strengths, thereby accurately modeling the causal transmission law in the attack stage. Facing the challenge of scarce negative samples, a dynamic adversarial negative sampling strategy is innovatively proposed, combined with a meta-learning framework to dynamically generate hard negative examples, and the loss function weight allocation mechanism is reconstructed, enabling the model to still maintain strong discriminative power for attack features under an extremely low positive sample ratio. To solve the problem of multi-source log compatibility, a unified formatting preprocessing framework is developed. Through an entity relationship mapping rule library, a time series alignment algorithm, and an adaptive pruning strategy, heterogeneous data is transformed into standardized traceability graph slices, which not only preserves the semantic integrity of the original logs but also realizes model generality across data sources, and finally constructs a high-precision detection scheme that can cover multi-dimensional attack surfaces such as hosts, networks, and users.

[0115] The application prospect of the present invention is broad, and it has significant value especially in the field of advanced persistent threat (APT) defense. With the accelerating digital transformation of enterprises, APT attacks have penetrated into key fields such as finance, energy, and government affairs. Traditional rule-based or single-point detection technologies are difficult to cope with multi-stage and long-cycle stealth attacks. By integrating the traceability graph constructed from multi-source logs and an optimized graph neural network model, this technology can be widely applied to the network security operation center (SOC) of large enterprises to analyze host process, network traffic, and user behavior data in real time, and accurately identify attack stages such as lateral movement and data leakage. For example, in the financial industry, the system can correlate abnormal logs of transaction systems with internal personnel operation records to quickly locate supply chain attacks or internal penetration behaviors; in the cloud service scenario, it can dynamically track the attack path across virtual nodes to solve the detection problems of new threats such as container escape and covert communication between microservices. Its multi-source data compatibility feature can also empower the industrial Internet by integrating heterogeneous logs of OT and IT systems to cope with APT attacks against critical infrastructure.

[0116] The parts not described in detail in the present invention belong to the common general knowledge of those skilled in the art.

[0117] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.

Claims

1. A negative sample enhanced APT attack detection method based on graph structure learning, characterized in that, It includes the following steps: (1) Obtain a security log dataset from the public network and construct a visual traceability graph using the network topology graph tool NetworkX; (2) Preprocess the traceability graph to obtain a training set, a validation set, and a test set; (3) Construct an encoder of the graph neural network, taking the time window, node features, adjacency matrix, and edge features as inputs to obtain an encoded result vector: (3a) Use a message passing layer to merge the edge features of neighboring nodes into the node embeddings, taking the edge features and edge weights as inputs respectively; (3b) Construct an edge feature processing layer EF and an edge weight processing layer EW, which are used to obtain the aggregated feature value EF containing edge features in the t time window t and the aggregated weight feature value EW t ; (3c) Based on the information passing neural network, use three EWs and one EF to construct an encoder, whose structure is as follows: the first EW layer - the second EW layer - the first activation function layer - the weighted EW layer - the second activation function layer - the EF layer - the multi-layer perceptron MLP - the third activation function layer; output the encoded result vector through the third activation function layer; (4) Construct the decoder of the graph neural network to generate the reconstructed adjacency matrix according to the following steps (4a) Aggregate the neighbor information of nodes through random sampling to capture the community structure that the attacker may form; (4b) Obtain the aggregated node embeddings according to the following formula Among them, S represents the identity matrix, and each row of this matrix is s neighborhoods extracted from the adjacency matrix A t ; Z t is the embedding representation matrix of the time window t; (4c) Processed by a multi-layer perceptron MLP and a non-linear activation function Softmax Generate a reconstructed adjacency matrix This matrix represents the probability that the corresponding edge is a malicious edge: (5) Design a loss function of the graph neural network, which is expressed as follows: Among them, is the average precision AP loss; β is a parameter specifically adjusted according to the dataset; is the decoder loss of the sampled edges calculated using the cross-entropy function CE; (6) Use the loss function to perform recurrent neural network (RNN) training on the encoder and decoder of the graph neural network to obtain an optimized graph neural network model; (7) Input the data to be tested into the graph neural network model, and use the model to identify abnormal attack traffic to complete attack detection.

2. The method according to claim 1, wherein: The security log dataset described in step (1) adopts the Los Alamos National Laboratory (LANL) dataset and the Operationally Transparent Network (OpTC) dataset.

3. The method according to claim 2, wherein: In step (2), preprocess the traceability graph, construct a sample set and divide it into a training set, a validation set, and a test set, and the implementation is as follows: For the LANL dataset, perform data annotation on the dataset according to the content of its corresponding red team attack information to identify malicious data flows; then segment it by 30 - 60 - minute snapshots, and merge all events sharing the same source node and target node into one edge; intercept all snapshots before the start of the red team attack; each snapshot has an average of 4000 - 7000 edges, discard 5% - 10% of the edges as the validation set, and use the remaining snapshots for testing; Construct node features X for the OpTC dataset t and edge weights Wt, set the snapshot size to 180 - 360 seconds to obtain 1,680 - 3,360 snapshots in a week; use 20 - 40% of these snapshots for testing and the rest for training.

4. The method according to claim 1, characterized in that: The aggregated eigenvalue EF described in step (3b) t , which is obtained according to the following formula: Among them, let V denote the set of nodes, u ∈ V; v represents the neighbor node of u; EF(X u ,F u,v ) represents the aggregated feature of node u under the time window t; X u represents the node feature of u; F u,v represents the edge feature between node u and v; N(u) represents the neighborhood of u; M θ represents the eigenvalue of the aggregated features transmitted by the neighborhood nodes, and σ represents the activation function.

5. The method according to claim 1, wherein: The aggregation weight value EW described in step (3b) t , is obtained according to the following formula: Among them, EW(X, W) represents the aggregated weight value EW of node feature X under the time window t t ; is the angular matrix for normalizing the edge weight matrix W; Θ represents the trainable parameter; σ represents the activation function.

6. The method according to claim 1, wherein: The loss function of the graph neural network described in step (5) is obtained according to the following steps: (5a) Calculate the decoder loss L of the sampled edges using the cross-entropy function CE DEC : (5b) Introduce the average precision AP loss Among them, P t is the positive edge in G t , nP t is the number of positive edges, e i is a positive edge, n is the number of all edges in G t , e s is any edge in G t , y s is the edge label, I is the identification function that outputs 1 / 0 when the parameter is true / false; l(e s ; e i ) = (m - (f(e i ) - f(e s ))) 2 is the smoothed square loss approximating I(f(e s ) ≥ f(e i ))), where f is the prediction model and m is the margin value; (5c) Combine L ap and L DEC , and assign different weights according to the characteristics of the dataset to obtain the final loss function:

7. The method according to claim 1, characterized in that: The recurrent neural network (RNN) training described in step (6) is implemented as follows: (6a) The encoder receives the graph structure data at the current moment, including node features and adjacency relationships, and generates node embedding vectors; (6b) Integrate the node embedding vectors generated at all historical moments through the RNN, capture the dynamic evolution law through time series modeling, and output a time - series enhanced node embedding sequence; (6c) Based on the node embedding sequence, the decoder reconstructs the corresponding adjacency matrix according to the node embeddings at each moment, and calculates the decoder loss by comparing the reconstructed matrix with the real graph structure; (6d) Use the loss function to calculate the gradients of the encoder, decoder, and RNN, and update the parameters through backpropagation to achieve optimization.

Citation Information

Patent Citations

  • Attack path tracing and attack source detection method based on machine learning

    CN115412328A

  • APT attack detection method and device based on mask graph auto-encoder

    CN116192477A

  • APT attack investigation method and device, computer equipment and storage medium

    CN116776102A

  • Intrusion detection method based on gating time convolutional network and graph

    CN117579324A

  • APT attack detection and tracing method based on graph attention sequential network

    CN117749437A

Cited By

  • Industrial control network risk monitoring method and device based on immune recognition

    CN120979802A

  • Cross-device lateral movement attack dynamic detection method and device

    CN121864387A