APT abnormal behavior detection method based on multi-dimensional feature fusion

By fusing the multi-dimensional features of the traceability diagram and the function call tree in exception detection and combining the attention mechanism for feature fusion, the problem of abnormal detection that is difficult to achieve high accuracy and high generalization capabilities in the existing technology is solved, and stronger adaptability and detection accuracy are achieved.

CN119995952AActive Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510064615.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Existing anomaly detection technologies are difficult to meet the needs of complex information systems for high-precision and high generalization capabilities, especially when facing changing network conditions and new attacks.

Method used

APT abnormal behavior detection method based on multi-dimensional feature fusion is adopted. By building a traceability diagram and function call tree, the semantic attribute embedding and structural information embedding of nodes are extracted, and multi-dimensional feature fusion is combined with attention mechanism, and an exception behavior detection is finally performed using a classifier.

Benefits of technology

It significantly improves the generalization capability and detection accuracy of the anomaly detection model, can more comprehensively evaluate the system security status, and adapt to complex and changeable network security threats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995952A_ABST
    Figure CN119995952A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network security, and discloses a multi-dimensional feature fusion-based APT abnormal behavior detection method, which comprises the following steps of: constructing a traceability graph by using an acquired system operation log, separating system call sequence information from acquired data, and respectively calculating point embedding of the traceability graph and sequence embedding of function call; and the two features are embedded and fused, and abnormal behavior discovery is carried out through a detection model based on the fused features, so that the generalization ability of the abnormal detection model is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security technology, and specifically relates to an APT abnormal behavior detection method based on multi-dimensional feature fusion, which effectively improves the generalization ability of the anomaly detection model. Background Art

[0002] At present, anomaly detection technology has become a key research area to ensure the security of information systems. Traditional anomaly detection methods, such as rule-based and statistical methods, have obvious shortcomings. Rule-based detection relies on predefined rules or signature libraries, can only identify known attack patterns, and is difficult to be effective against new attacks and cannot adapt to changing network conditions and attack methods. Although statistical-based detection can detect unknown anomalies, it has limited ability to model complex system behaviors and is prone to high false alarm rates when the system fluctuates normally.

[0003] Machine learning technology is increasingly used in anomaly detection, and data feature representation is particularly critical. Provenance graphs have emerged in network security analysis. They use nodes and edges to represent system entities and causal relationships, record data sources and propagation paths at a macro level, and are used for anomaly detection to explore potential anomaly propagation paths. However, detection based only on provenance graphs focuses on node connections and topological structures, and does not adequately explore details such as the execution logic of node internal processes.

[0004] The function call tree depicts the process behavior from a microscopic perspective. When the process is executed, functions are called in sequence and hierarchy to form a call tree, which can reflect the logical structure of the process, the calling relationship of functional modules, and the data transfer process. By extracting its features, abnormal function call patterns can be found, such as sequence anomalies or illegal calls. However, relying solely on the function call tree features to detect anomalies will ignore the data interaction relationship at the system level, making it difficult to fully evaluate the system security status.

[0005] In summary, in the field of anomaly detection, whether it is based on traditional methods or machine learning methods based on a single type of feature (such as only traceability graphs or only function call trees), it is difficult to meet the needs of complex information systems for high-precision and high-generalization anomaly detection. Summary of the invention

[0006] The purpose of the present invention is to provide an APT abnormal behavior detection method based on multi-dimensional feature fusion, which uses the collected system operation logs to build a traceability graph, separates the system call sequence information from the collected data, calculates the point embedding of the traceability graph and the sequence embedding of the function call respectively, and fuses the feature embeddings of the two. Abnormal behavior is discovered through a detection model based on fused features, so as to significantly improve the generalization ability of the anomaly detection model and improve the accuracy of abnormal behavior detection.

[0007] To achieve the above object, the technical solution adopted by the present invention is:

[0008] A method for detecting APT abnormal behavior based on multi-dimensional feature fusion, comprising:

[0009] Collect system kernel log data, build a traceability graph based on the system kernel log data, and extract function call sequence information from the system kernel log data;

[0010] Adopting adaptive neighborhood sampling method to sample the neighborhood subgraph of process nodes from the provenance graph;

[0011] Construct a semantic description statement based on the triples in the neighborhood subgraph, and extract features from the description statement to get the semantic attribute embedding of the process node;

[0012] The neighborhood subgraph and semantic attributes of the process node are embedded into the input graph neural network to obtain the structural information embedding of the process node output by the graph neural network. The semantic attribute embedding and the structural information embedding are combined as the semantic structure feature vector of the process node.

[0013] Based on the extracted function call sequence information, a call tree of a process node is constructed in units of processes. The vertex labels of the call tree are iterated to generate all subtrees of the call tree. The execution time of the root node of the subtree is set as the weight of the subtree to obtain a subtree set. The subtree set includes all subtrees and the weights of the subtrees. Feature extraction is performed on the subtree set to obtain a call tree feature vector of the process node.

[0014] Using the attention-based multi-dimensional feature fusion method, the multi-dimensional fusion features are obtained according to the semantic structure feature vector and call tree feature vector of the same process node, and the classifier is used to output the APT abnormal behavior detection results based on the multi-dimensional fusion features.

[0015] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution, but are merely further supplements or preferences. Under the premise that there are no technical or logical contradictions, each optional method can be combined with the above-mentioned overall solution separately, and multiple optional methods can also be combined.

[0016] Preferably, the method of sampling the neighborhood subgraph of the process node from the provenance graph using an adaptive neighborhood sampling method comprises:

[0017] Calculate the behavior density index of the node:

[0018]

[0019] Where D(v) represents the behavior density index of node v, E v is the number of outgoing edges of node v, represents the average number of outgoing edges of the adjacent nodes of node v;

[0020] If D(v)≥θ, node v is considered to be a behavior-dense point, and a breadth-first search strategy is used to sample the neighborhood of node v; if D(v)<θ, node v is considered to be a behavior-sparse point, and a depth-first search strategy is used to sample the neighborhood of node v, where θ is the density threshold.

[0021] Preferably, the step of constructing a description statement with semantics according to the triples in the neighborhood subgraph, and extracting features from the description statement to obtain the semantic attribute embedding of the process node comprises:

[0022] Take the triple in the neighborhood subgraph as <source node, event, target node>;

[0023] Combine the source node attributes, event type, and target node attributes in a triple into a semantic sentence;

[0024] Sort the sentences corresponding to the triples of all process nodes containing the features to be extracted in the neighborhood subgraph according to the timestamps of the events to obtain semantically descriptive sentences;

[0025] The word2vec model trained on benign system kernel log data is used to encode the description sentences and obtain the semantic attribute embedding of the process node.

[0026] Preferably, the iterative call tree node labels generate all subtrees of the call tree, including:

[0027] Initial label assignment: Assign an initial label to each vertex in the call tree. The initial label is the function call name represented by the vertex.

[0028] Label expansion: In each iteration, for each vertex, the label of each vertex is appended with the label of its child vertices to form a signature structure;

[0029] Label compression: After label expansion, a virtual new label is used to represent the signature structure;

[0030] Termination condition: if the current number of iterations reaches the set maximum number of iterations, the iteration ends and a label is used as a subtree; otherwise, the iteration continues to perform label expansion and label compression, and the maximum number of iterations is set to d-1 times according to the maximum depth d of the call tree.

[0031] Preferably, setting the execution time of the subtree root node as the weight of the subtree includes:

[0032] Determine the timestamp of the system call corresponding to the leaf node of the call tree;

[0033] Calculate the relative time interval between two adjacent timestamps and set the relative time interval as the weight of the leaf node;

[0034] For non-leaf nodes, the sum of the weights of all child nodes of the non-leaf node is used as the weight of the non-leaf node;

[0035] Therefore, the weights of each node in all subtrees in the call tree are obtained, and the weight of the root node of the subtree is set as the weight of the subtree.

[0036] Preferably, the multi-dimensional feature fusion method based on attention is used to obtain the multi-dimensional fusion feature according to the semantic structure feature vector and the call tree feature vector of the same process node, including:

[0037] Multi-dimensional feature extraction: semantic structure feature vector E of process node v and the call tree feature vector E of the process node t , convolution operation is performed through n convolution kernels of different sizes to obtain the semantic structure feature vector E v and the call tree feature vector E t Feature information F under different convolution kernels vj and F tj , j = 1, 2, ..., n;

[0038] Attention mechanism calculation: Use average pooling operation to process feature information F vj and F tj , get the fixed length feature vector G vj and G tj , and then the attention score A is obtained through the full connection layer mapping vj and A tj ;

[0039] Multi-dimensional feature fusion: using attention score A vj For feature information F vj Perform weighted fusion to obtain semantic structure fusion features Using attention score A tj For feature information F tj Perform weighted fusion to obtain semantic structure fusion features Finally, we get the multi-dimensional fusion features

[0040] The present invention provides an APT abnormal behavior detection method based on multi-dimensional feature fusion, which has the following beneficial effects compared with the prior art: 1. Combining the traceability graph node features with the process function call tree features, it not only covers the traceability association information between entities in the system, but also includes the logical details of the function calls within the process. 2. The fused features can dig out the complex interactive behavior patterns between system processes based on the traceability relationship and function call logic, breaking through the limitations of the prior art that only analyzes from a single behavior level (such as only tracing or only function calls). 3. The model constructed by multi-dimensional feature fusion and deep behavior analysis has stronger adaptability and generalization capabilities, as well as higher detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of an APT abnormal behavior detection method based on multi-dimensional feature fusion according to the present invention;

[0042] Figure 2 A schematic diagram of a description statement construction process of the present invention;

[0043] Figure 3 A schematic diagram of obtaining a subtree by iterative call tree of the present invention;

[0044] Figure 4 This is a schematic diagram of subtree weight calculation according to the present invention. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0047] In order to effectively improve the performance of the anomaly detection model, the present invention organically combines the traceability graph node features with the process function call tree features. Through feature fusion technology, the advantages of both in describing system behaviors at different levels are fully utilized, so that the model can learn the normal and abnormal behavior patterns of the system more comprehensively and accurately, thereby improving the detection and generalization capabilities of unknown anomalies and better responding to complex and changeable network security threats.

[0048] like Figure 1As shown, this embodiment provides an APT abnormal behavior detection method based on multi-dimensional feature fusion, comprising the following steps:

[0049] (1) Collect system kernel log data, build a traceability graph based on the system kernel log data, and extract function call sequence information from the system kernel log data.

[0050] (1-1) System operation data collection. Use the published log collection tool to collect system kernel log data (such as the log collection tool disclosed in Patent No. 2022110610516) to fully capture the operations during the operation process, including function call sequences, process creation and termination, file read and write access, network connection establishment and disconnection, and other information.

[0051] (1-2) Construct a traceability graph. Parse and compress the collected system kernel log data, identify entities such as processes, files, and network connections as nodes of the traceability graph, and construct directed edges based on the operational causal relationship between them. At the same time, add rich attribute information to each node and edge, such as the node type identifier, the timestamp of the operation, the event type of the edge, etc., and construct a complete and information-rich streamlined traceability graph through compression rules. The compression operations are as follows: 1. Merge duplicate events: Only one edge of the same type of event between nodes is calculated; 2. Delete low-priority events: Delete low-priority events that do not affect behavioral analysis, such as deleting file registry opening, closing file cache cleaning, and other events (low-priority events can be preset through black and white lists).

[0052] (1-3) Extracting function call sequence information. Filter out the function call sequence information from the collected system kernel log data. For example, in Windows systems, the stackinfo log data of the ETW (Event Tracing for Windows) tool can be used to restore the callstack. The extracted function call sequence is stored in a special sequence database, using the process id as the identifier of the function call sequence of the process.

[0053] (2) Node feature extraction: For the process nodes in the traceability graph, the node subgraph is first sampled using the set neighborhood sampling strategy, and then the node features are calculated to form a feature vector.

[0054] (2-1) Adaptively select the subgraph sampling strategy according to the sparsity of the neighborhood behavior of the node in the traceability graph, and use the adaptive neighborhood sampling method to sample the neighborhood subgraph of the process node from the traceability graph: Design adaptive neighborhood sampling, and sample the neighborhood subgraph of the node to be detected from the traceability graph. Adaptation specifically refers to selecting different subgraph sampling strategies based on whether the node neighborhood behavior is sparse. In order to determine whether the behavior of the node in the traceability graph is dense or sparse, a behavior density index D(v) is introduced for evaluation. The behavior density index formula of node v is as follows:

[0055]

[0056] Where D(v) represents the behavior density index of node v, E v is the number of outgoing edges of node v, Indicates the average number of outgoing edges of the neighboring nodes of node v. If D(v)≥θ, it is considered to be a behavior-dense point, and the breadth-first search strategy is used for neighborhood sampling (the number of sampling layers is set, for example, it is set to 3 based on experience); if D(v)<θ, it is considered to be a behavior-sparse node, and the depth-first search strategy is used for neighborhood sampling (the sampling depth is set, for example, it is set to 5 based on experience), where θ is the preset density threshold.

[0057] For example, in the traceability graph of a file server, a core process node that frequently processes a large number of file read and write requests has a large number of read and write operation edges with multiple file nodes, while the average number of outgoing edges of other adjacent auxiliary process nodes is relatively small. The D(v) of the core process node is greater than the threshold, and it is determined to be a behavior-intensive node. On the other hand, a background process node that occasionally performs system update checks has only a small number of interaction edges with network connection nodes, system configuration file nodes, etc., and the calculated D(v) is less than the threshold, so it is considered to be a behavior-sparse node.

[0058] (2-2) Embed the node attributes through the learning model to obtain the feature vector of the node attribute dimension: Based on the neighborhood subgraph of the node v to be calculated obtained by sampling in step (2-1), encode the attributes of each node, and finally form a feature matrix composed of feature vectors. The representation is obtained through the learning model, and finally the semantic attribute features and graph structure features of the node are combined as the semantic attribute embedding E of the node v The main process is as follows:

[0059] (2-2-1) Construct a semantic description sentence through the attributes of the node and generate an embedded representation vector: Different types of nodes have different attributes, including the process name and command line parameters of the process node, the file path of the file node, the network IP address and port of the socket node, etc. The semantic attributes of the node, such as process name, command line parameters, etc., are combined with the event type between nodes and their one-hop neighbors to construct a sentence describing the node, and the events are sorted according to the timestamps of the events to maintain the time order, thus obtaining a semantic description sentence.

[0060] For ease of understanding, this embodiment explains the formation of a description statement from the perspective of triples as follows: take the triple in the neighborhood subgraph as <source node, event, target node>; combine the attributes of the source node, the type of event, and the attributes of the target node in a triple into a sentence with semantics; sort the sentences corresponding to the triples of all process nodes in the neighborhood subgraph that contain the features to be extracted according to the timestamp of the event to obtain a description statement with semantics.

[0061] Then, the word2vec model trained on the kernel log data of the benign system is used to encode the description sentences and encode the sentences into fixed-length vectors, thereby converting the semantic attributes of the nodes into low-dimensional dense vector representations and obtaining the semantic attribute embedding of the process nodes.

[0062] For a process of describing the construction of a sentence, Figure 2 As shown, Figure 2 (a) is part of a traceability graph. Rectangles are process nodes, ovals are file nodes, and diamonds are networks. The edges are <event, timestamp>, and the timestamp sequence is t1, t2, …, t7. Taking process node P1 as an example, we can get a sentence with semantics such as Figure 2 As shown in (b) of Figure 2, all sentences are sorted according to the timestamps as shown in Figure 2 (c), and the description statement of process node P1 is finally obtained as follows Figure 2 As shown in (d), P0 starts P1, P1 writes F1, P1 starts P2, P1 sends S1, and P1 writes F2.

[0063] (2-2-2) Use the trained machine learning model to obtain the structural feature vector of the node. Specifically, the neighborhood subgraph and semantic attributes of the process node are embedded into the input graph neural network to obtain the structural information embedding of the process node output by the graph neural network. This embodiment uses a semi-supervised node classification method to train a GNN (Graph Neural Network) model (such as a GraphSage model), and uses the semantic attribute embedding obtained in (2-2-1) and the node v sampling domain subgraph in step (2-1) as structural information input to obtain a multi-dimensional feature vector of the node as structural information embedding.

[0064] Then the semantic attribute embedding of step (2-2-1) and the structural information embedding of step (2-2-2) are combined (such as splicing) as the semantic structural feature vector E of the process node v .

[0065] (3) Function call embedding calculation. Based on the function call sequence information extracted in step (1), a call tree of the corresponding process node is constructed in units of processes. The vertex labels of the call tree are iterated to generate all subtrees of the call tree. The execution time of the root node of the subtree is set as the weight of this subtree, thereby obtaining a subtree set of the call tree. The subtree set contains all subtrees and the weights of the subtrees. The model is used to learn the embedding of the subtree set to generate the call tree feature vector of the process node.

[0066] (3-1) Constructing a function call tree: Based on the function call sequence information extracted in step (1), for each process node, construct a call tree as an independent unit. With the process start as the root node, when the system executes a function call, add the corresponding node to the call tree, and construct the parent-child node relationship based on the execution order and nested relationship.

[0067] (3-2) Iterate vertex labels: Use the improved graph core algorithm to generate all subtrees of the tree called in step (3-1). The main process of the improved graph core algorithm is as follows:

[0068] (3-2-1) Initial label assignment: Assign an initial label to each vertex in the call tree, that is, the function call name represented by the vertex. For example, if the vertex represents system call A, then set its initial label to A.

[0069] (3-2-2) Label expansion: In each iteration, for each vertex, its label is appended with the labels of its child vertices to form a signature structure. Assuming that vertex v has child vertices v1 (labeled B) and v2 (labeled C), the signature structure of vertex v is (A, BC), which is used to explore the relationship structure between the vertex and its child nodes.

[0070] (3-2-3) Label compression: After label expansion, compress the newly generated virtual labels and use the new virtual labels to represent the specific signature structure. If multiple vertices have the same signature structure, assign them a unified new label to reduce data redundancy and facilitate subsequent subtree structure identification and processing.

[0071] (3-2-4) Termination condition: If the current number of iterations reaches the set maximum number of iterations, the iteration is terminated and a label is used as a subtree; otherwise, the iteration continues to perform label expansion and label compression, and the maximum number of iterations is set to d-1 times according to the maximum depth d of the call tree.

[0072] This embodiment provides an iterative specific process such as Figure 3 As shown, Figure 3 (a) in the figure is the one that has been assigned an initial label. For example, if the root function is A, its label is A. Figure 3 The call tree depth is 3, and the maximum number of iterations is 2; take the node labeled A as an example to iterate the vertex label for the first time. Label A is first expanded, and the labels of the child nodes are B, C and D respectively. Then the signature structure of label A after expansion exploration is (A, BCD). In order to facilitate subsequent processing, label compression is set to set label H to be equivalent to the signature structure (A, BCD), that is, H = (A, BCD), H is also Figure 3 The rectangle represents Figure 3 (b) in the figure is the result of the first iteration; similarly, the result of the second iteration can be obtained, that is, Figure 3 As shown in (c), the new labels H, I, J, K, L and M are finally expanded, and their equivalent signature structures are H = (A, BCD), I = (B, D), J = (C, E), K = (D, FG), L = (H, IJK) and M = (I, K), respectively.

[0073] (3-3) The extracted subtree is used as text, and the subtree execution time is used as the weight to obtain the subtree set of calls: In order to encode the time information, the process running time represented by the root node of the subtree is used as the weight of the subtree. Due to the limitation that the exact value of the function running time cannot be obtained, the relative running time of the timestamp is used to represent the running time of the process represented by the root node of the subtree. The leaf node of the call tree T is also the system call. A fixed basic time unit is set (through preliminary analysis of system behavior, it can be assumed that the equivalent time unit is 1 millisecond). The timestamp of each system call is recorded. By calculating the relative time interval between two adjacent timestamps, if the last system call has appeared before, it is set to the existing one. If it appears for the first time, the relative time interval is the basic time unit. Therefore, the weight of the leaf node is the relative time interval. For the non-leaf node, the weight is the sum of the weights of its child nodes. Therefore, the weight of each node in all subtrees in the call tree is obtained, and the weight of the root node of the subtree is set as the weight of the subtree.

[0074] All subtrees generated by the iterative marking of the call tree T in step (3-2) are added to the set S. The weight of the subtree is the weight of its root node. Finally, the same subtrees are aggregated and the weights of the same subtrees are added. Different subtrees are saved to form the final subtree set.

[0075] This embodiment provides a specific process of generating a subtree set. Figure 4 As shown, after step (3-2), the call tree obtains the subtree set S = {A, B, C, D, D, F, G, E, F, G, H, I, J, K, L, M}. After merging the same subtrees, the updated subtree set is S = {A, B, C, D, E, F, G, H, I, J, K, L, M}, where F, G, E, F, G are system calls, and their occurrence timestamps are t1, t2, t3, t4, and t5 respectively. The weights obtained by the timestamp difference are 10, 5, 15, 10, and 5 respectively, which can be obtained as follows Figure 4 (a) shows the initialization node weight. For the system call corresponding to the occurrence timestamp t1, since there is no relative time interval, its execution time is first set as the basic time unit, and then after the relative time interval is calculated for the system call at time t4, it is synchronously updated to the corresponding relative time interval.

[0076] like Figure 4 As shown in (b), after step (3-2) and merging the same subtrees, the set of subtrees obtained can obtain their corresponding weights. The weights of the same subtrees are added and merged, and finally the subtree set {A, B, C, D, F, G, E, H, I, J, K, L, M} is obtained, and its corresponding weights are {15, 10, 15, 20, 20, 10, 15, 45, 15, 15, 15, 45, 15}.

[0077] (3-4) Generation of call tree feature vector: Using the call subtree set obtained in step (3-3) as input data, a general model is trained using a natural language processing model (such as doc2vec) to obtain the call tree feature vector E. t .

[0078] (4) Anomaly detection based on feature fusion: By combining the semantic structure feature vector E obtained in step (2) v And the call tree feature vector E for the process node obtained in step (3) t , the fusion feature E is obtained by using the multi-dimensional feature fusion method based on attention guidance f , and use common patterns to build anomaly detection models for detection.

[0079] (4-1) Multi-dimensional feature extraction: For the semantic structure feature vector E of the process node vand the call tree feature vector E of the process node t , convolution operation is performed through n convolution kernels of different sizes to obtain the semantic structure feature vector E v and the call tree feature vector E t Feature information F under different convolution kernels vj and F tj , j = 1, 2, ..., n. Assuming that this embodiment uses 2 convolution kernels, that is, n = 2, the specific formula is as follows:

[0080]

[0081] Among them, s vj Yes E v The kernel size of the convolution, n vj is the number of convolution kernels, and the feature F can be obtained v1 and F v2 Similarly, the call tree is convolved to obtain the feature F t1 and F t2 For example, the convolution kernel dimensions are set to 3 and 5, and the number is set to 16 and 32 respectively.

[0082] (4-2) Attention mechanism calculation: Use average pooling operation to process feature information F vj and F tj , get the fixed length feature vector G vj and G tj , and then the attention score A is obtained through the full connection layer mapping vj and A tj ; The specific formula of this embodiment is as follows:

[0083] G vj = GAP(F vj )(j=1,2)

[0084] G tj = GAP(F vj )(j=1,2)

[0085] A vj =softmax(W v2 (ReLU(W v1 G vj +b v1 ))+b v2 )(j=1,2)

[0086] A tj =softmax(W t2 (ReLU(W t1 G tj +b t1 ))+b t2 )(j=1,2)

[0087] Where GAP represents the global average pooling operation, ReLU is the activation function, and softmax is the softmax function. v The features of the fully connected layer are W v1 and W v2 , the bias term is b v1 and b v2 , the sum of the attention weights of each scale feature is 1. t The features of the fully connected layer are W t1 and W t2 , the bias term is b t1 and b t2 , the sum of the attention weights of each scale feature is 1.

[0088] (4-3) Multi-dimensional feature fusion: using attention score A vj For feature information F vj Perform weighted fusion to obtain semantic structure fusion features Using attention score A tj For feature information F tj Perform weighted fusion to obtain semantic structure fusion features Finally, we get the multi-dimensional fusion features The specific formula of this embodiment is as follows:

[0089]

[0090] where € means element-wise multiplication.

[0091] (4-4) Use the classifier to build an anomaly detection model: Take the E from step (4-3) f As feature vectors, common classifiers, such as support vector machine (SVM) and K nearest neighbor (KNN) models, are used to train anomaly detection models on benign system kernel log data, and the anomaly detection model is used to target the multi-dimensional fusion feature E f Output APT abnormal behavior detection results.

[0092] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0093] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A method for detecting abnormal behavior of APT based on multi-dimensional feature fusion, characterized in that: The APT abnormal behavior detection method based on multi-dimensional feature fusion includes: Collect system kernel log data, build a traceability graph based on the system kernel log data, and extract function call sequence information from the system kernel log data; Adopting adaptive neighborhood sampling method to sample the neighborhood subgraph of process nodes from the provenance graph; Construct a semantic description statement based on the triples in the neighborhood subgraph, and extract features from the description statement to get the semantic attribute embedding of the process node; The neighborhood subgraph and semantic attributes of the process node are embedded into the input graph neural network to obtain the structural information embedding of the process node output by the graph neural network. The semantic attribute embedding and the structural information embedding are combined as the semantic structure feature vector of the process node. Based on the extracted function call sequence information, a call tree of a process node is constructed in units of processes. The vertex labels of the call tree are iterated to generate all subtrees of the call tree. The execution time of the root node of the subtree is set as the weight of the subtree to obtain a subtree set. The subtree set includes all subtrees and the weights of the subtrees. Feature extraction is performed on the subtree set to obtain a call tree feature vector of the process node. Using the attention-based multi-dimensional feature fusion method, the multi-dimensional fusion features are obtained according to the semantic structure feature vector and call tree feature vector of the same process node, and the classifier is used to output the APT abnormal behavior detection results based on the multi-dimensional fusion features.

2. The APT abnormal behavior detection method based on multi-dimensional feature fusion according to claim 1 is characterized in that: The method of sampling a neighborhood subgraph of a process node from a traceability graph using an adaptive neighborhood sampling method includes: Calculate the behavior density index of the node: Where D(v) represents the behavior density index of node v, E v is the number of outgoing edges of node v, represents the average number of outgoing edges of the adjacent nodes of node v; If D(v)≥θ, node v is considered to be a behavior-dense point, and a breadth-first search strategy is used to sample the neighborhood of node v; if D(v)<θ, node v is considered to be a behavior-sparse point, and a depth-first search strategy is used to sample the neighborhood of node v, where θ is the density threshold.

3. The APT abnormal behavior detection method based on multi-dimensional feature fusion according to claim 1 is characterized in that: The step of constructing a description statement with semantics according to the triples in the neighborhood subgraph and extracting features from the description statement to obtain the semantic attribute embedding of the process node includes: Take the triple in the neighborhood subgraph as <source node, event, target node>; Combine the source node attributes, event type, and target node attributes in a triple into a semantic sentence; Sort the sentences corresponding to the triples of all process nodes containing the features to be extracted in the neighborhood subgraph according to the timestamps of the events to obtain semantically descriptive sentences; The word2vec model trained on benign system kernel log data is used to encode the description sentences and obtain the semantic attribute embedding of the process node.

4. The APT abnormal behavior detection method based on multi-dimensional feature fusion according to claim 1 is characterized in that: The node labels of the iterative call tree generate all subtrees of the call tree, including: Initial label assignment: Assign an initial label to each vertex in the call tree. The initial label is the function call name represented by the vertex. Label expansion: In each iteration, for each vertex, the label of each vertex is appended with the label of its child vertices to form a signature structure; Label compression: After label expansion, a virtual new label is used to represent the signature structure; Termination condition: if the current number of iterations reaches the set maximum number of iterations, the iteration ends and a label is used as a subtree; otherwise, the iteration continues to perform label expansion and label compression, and the maximum number of iterations is set to d-1 times according to the maximum depth d of the call tree.

5. The APT abnormal behavior detection method based on multi-dimensional feature fusion according to claim 1 is characterized in that: The step of setting the execution time of the root node of the subtree as the weight of the subtree includes: Determine the timestamp of the system call corresponding to the leaf node of the call tree; Calculate the relative time interval between two adjacent timestamps and set the relative time interval as the weight of the leaf node; For non-leaf nodes, the sum of the weights of all child nodes of the non-leaf node is used as the weight of the non-leaf node; Therefore, the weights of each node in all subtrees in the call tree are obtained, and the weight of the root node of the subtree is set as the weight of the subtree.

6. The APT abnormal behavior detection method based on multi-dimensional feature fusion according to claim 1 is characterized in that: The multi-dimensional feature fusion method based on attention is used to obtain multi-dimensional fusion features according to the semantic structure feature vector and call tree feature vector of the same process node, including: Multi-dimensional feature extraction: semantic structure feature vector E of process node v and the call tree feature vector E of the process node t , convolution operation is performed through n convolution kernels of different sizes to obtain the semantic structure feature vector E v and the call tree feature vector E t Feature information F under different convolution kernels vj and F tj , j = 1, 2, ..., n; Attention mechanism calculation: Use average pooling operation to process feature information F vj and F tj , get the fixed length feature vector G vj and G tj , and then the attention score A is obtained through the full connection layer mapping vj and A tj ; Multi-dimensional feature fusion: using attention score A vj For feature information F vj Perform weighted fusion to obtain semantic structure fusion features Using attention score A tj For feature information F tj Perform weighted fusion to obtain semantic structure fusion features Finally, we get the multi-dimensional fusion features

Citation Information

Patent Citations

  • Kernel log joint compression and query method fusing semantics and deep neural network

    CN117453646A

  • APT detection method based on semantic enhancement and attention mechanism

    CN119272277A

  • Complex network attack detection method based on cross-host abnormal behavior recognition

    WO2024216729A1

Cited By

  • Student behavior anomaly detection system based on clustering processing

    CN120524399A

  • Student behavior anomaly detection system based on clustering processing

    CN120524399B

  • Complex network anomaly detection method and system based on dynamic graph neural network

    CN121644163A