A Real-time Detection and Analysis Method for APT Based on Heterogeneous Graph

By constructing heterogeneous graphs and applying graph embedding and profile technology, combined with pre-trained detection models, real-time and efficient detection of APT attacks is achieved, and the problem that traditional methods are difficult to capture long-term behavior patterns and dynamic modeling is easily poisoned, with high precision and scalability.

CN114896591BActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210593319.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-05-30
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

Traditional APT detection methods are difficult to capture long-term behavior patterns, ignore the timing, logic and interaction between log entries, and dynamic modeling is easily poisoned by attackers, resulting in reduced accuracy of identifying attacks.

Method used

APT real-time detection and analysis method based on heterogeneous graphs is adopted, and real-time and efficient APT detection is achieved by obtaining operating system log data, extracting five meta attributes, and constructing heterogeneous graphs. Graph embedding technology and graph summary technology are used to generate low-dimensional vector representations and compact sketches, and compared with pre-trained detection models to achieve real-time and efficient APT detection.

Benefits of technology

Real-time and efficient detection of APT attacks is realized, and the causality, timing and logical relationships between logs can be detected granularly, with scalability and low computational storage overhead, preventing attackers from poisoning the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114896591B_ABST
    Figure CN114896591B_ABST
Patent Text Reader

Abstract

The present invention discloses an APT real-time detection and analysis method based on a heterogeneous graph, which includes extracting five meta-attributes from log data, and converting the log data into a heterogeneous graph based on the preset edge generation rules according to the five meta-attributes; using graph embedding technology to extract the context relationships of each log entry in the heterogeneous graph to form a log sequence, and converting each log entry in the log sequence into a low-dimensional vector representation; generating a sketch of a fixed size based on the low-dimensional vectors of all nodes in the heterogeneous graph by adopting graph summary technology; comparing the sketch with a pre-trained detection model, if the sketch fits into a cluster in the evolved model under the detection model, the sketch is normal, that is, the obtained log data is log data in a scenario without APT attacks; otherwise, the sketch is abnormal, that is, the obtained log data is log data in an APT attack scenario. The present invention realizes the high efficiency, accuracy and real-time performance of APT detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information security, and particularly relates to a method for real-time detection and analysis of APT based on heterogeneous graphs. Background Art

[0002] At present, Advanced Persistent Threat (APT) has become increasingly common. In APT, the goal of APT participants is to gain control of a specific system (network), while remaining undetected for a long time, persistently harming multiple hosts and stealing confidential information. APT has a long-term and persistent attack pattern, and the frequent use of zero-day attacks makes it difficult to detect.

[0003] However, current traditional detection systems are difficult to detect APT: anomaly-based systems usually analyze a series of system calls and log adjacent system events, but most of them are difficult to model long-term behavior patterns; a provenance-based APT detection method uses simple edge matching rules based on prior attack knowledge and cannot detect new categories of APT; a prototype-based anomaly detection system relies on graph neighborhood exploration to understand normal behavior through static or dynamic models, and actual computational constraints limit the scope of feasible context analysis; in addition, existing detection algorithms are mostly based on behavior detection, and most methods consider the sequence relationship of logs and the behavior sequence of users, but ignore other relationships, resulting in inapplicability to diverse attack scenarios.

[0004] Therefore, the defects of traditional detection methods are mainly manifested in: First, only the single relationship between logs is considered, without simultaneously considering the temporal relationship, logical relationship, and interaction relationship between log entries; Second, static models cannot capture long-term system behavior. The attack scenario of APT is very long, and the inability to capture long-term system behavior will lead to a high false alarm rate; Third, the computational method in memory storage has poor scalability in the face of long-running attack scenarios; Fourth, there is a risk of being poisoned by attackers in runtime dynamic modeling. The long-term and persistent APT means will gradually poison the dynamic model and reduce its accuracy in identifying attacks. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for real-time detection and analysis of APT based on heterogeneous graphs to achieve real-time and efficient detection of APT attacks.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is:

[0007] A method for real-time detection and analysis of APT based on heterogeneous graphs, the method for real-time detection and analysis of APT based on heterogeneous graphs includes:

[0008] Step 1, obtain log data in the operating system;

[0009] Step 2: Extract five meta-properties from the log data. Based on the five meta-properties, convert the log data into a heterogeneous graph according to the preset edge generation rules. The nodes in the heterogeneous graph are log entries;

[0010] Step 3: Use graph embedding technology to extract the context relationships of each log entry in the heterogeneous graph, form a log sequence, and regard the log sequence as a sentence. Use the word2vec method to convert each log entry in the log sequence into a low-dimensional vector representation;

[0011] Step 4: Based on the low-dimensional vectors of all nodes in the heterogeneous graph, use graph summarization technology to generate a sketch of a fixed size;

[0012] Step 5: Compare the sketch with the pre-trained detection model. If the sketch fits into one of the clusters in the evolutionary model under the detection model, the sketch is normal, that is, the obtained log data is log data in a scenario without APT attacks; otherwise, the sketch is abnormal, that is, the obtained log data is log data in an APT attack scenario;

[0013] Among them, the generation process of the detection model includes:

[0014] Regularly obtain the sketches for this training from the log data in a scenario without APT attacks, and arrange the sketches for this training and the previously trained sketches in chronological order of creation time to form a sketch sequence;

[0015] Use a clustering algorithm to cluster the sketch sequence to obtain multiple clusters;

[0016] Generate the evolutionary model for this training based on the chronological order of the sketches in each cluster and the statistical data of each cluster, and aggregate the evolutionary models obtained from multiple trainings as the detection model.

[0017] The following also provides several optional methods, which are not additional limitations to the above overall solution, but are only further supplements or optimizations. Without technical or logical contradictions, each optional method can be combined with the above overall solution separately, or multiple optional methods can be combined with each other.

[0018] Preferably, the five meta-properties are source object, operation type, target object, time, and host.

[0019] Preferably, the preset edge generation rules include:

[0020] Rule 1: Connect the log entries of the same source object on the same day in chronological order to form a sequence of daily log entries. This rule has the highest weight;

[0021] Rule 2: Connect the log entries of the same target object on the same day in chronological order to form a sequence of daily log entries, and this rule has the highest weight;

[0022] Rule 3: Connect the sequences of daily log entries with the same source object and a similarity higher than the threshold;

[0023] Rule 4: Connect the sequences of daily log entries with the same target object and a similarity higher than the threshold.

[0024] Preferably, the similarity between the sequences of daily log entries is positively correlated with the similarity of the number of log entries they contain.

[0025] Preferably, the method of using graph embedding technology to extract the context relationship of each log entry in the heterogeneous graph to form a log sequence includes:

[0026] The graph embedding technology is the random walk method, and the definitions of edge types and edge weights in the random walk method are as follows:

[0027] The edge types: Divide the edge types of the edges generated based on Rule 1 and Rule 2 into type 1, and divide the edge types of the edges generated based on Rule 3 and Rule 4 into type 2, generating two sets of edge type sets as {Rule 1, Rule 3} and {Rule 2, Rule 4};

[0028] The edge weights: Each time the random walk is executed, one of the two sets of edge type sets is taken, so the edge weight function is expressed as:

[0029]

[0030] w(T,V) = e -s(m,n)

[0031] In the formula, w(T,V) represents the weight of the sequence T of log entries and the sequence V of log entries on the corresponding edge type set. The sequence T is the sequence of log entries generated by connecting according to the rule with edge type 1 in the edge type set, and the sequence V is the sequence of log entries generated by connecting according to the rule with edge type 2 in the edge type set. m is the number of log entries contained in the sequence T, n is the number of log entries contained in the sequence V, and s(m,n) is the similarity of the number of log entries in the sequence T and the sequence V.

[0032] Preferably, the calculation of the transition probability at node v in the random walk method is as follows:

[0033]

[0034] Wherein, P(t|v) represents the transition probability from node v to node t, E is the set of all edges in the heterogeneous graph, N(v) is the neighbor nodes of node v, and W N(v) is the sum of the edge weights between node v and all its neighbor nodes N(v), w(t,v) = w(T,V), is the edge type between node v and node t, and S n is the set of all edge types, and R w(t,v) is the sorting value based on the edge type of the edge weights of the nodes connected to node v in descending order among the edge weights of all nodes connected to node v, and neigh is the preset number of neighbor nodes.

[0035] The APT real-time detection and analysis method based on a heterogeneous graph provided by the present invention realizes constructing log data entries into a heterogeneous graph and formulating corresponding rules, so it can take into account the causal relationship, temporal relationship, and logical relationship between logs at the same time, and perform fine-grained detection of APT; by adopting the graph summary technology, all node vectors in the incrementally generated heterogeneous graph can be kept compact and of a fixed size, while also maintaining the similarity between log sequences, realizing scalability, low computational and storage overheads, and high-precision APT detection; this method directly models the evolutionary behavior of the system during data training and does not update the model afterwards, so it can prevent attackers from poisoning the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flowchart of the APT real-time detection and analysis method based on a heterogeneous graph of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0039] In order to overcome the problem of poor recognition of APT attacks in the prior art, this embodiment proposes an APT real-time detection and analysis method based on a heterogeneous graph, as Figure 1 shown, which specifically includes the following steps:

[0040] Step 1: Obtain the log data in the operating system.

[0041] In this embodiment, the SPADE framework is used to obtain the log data in the operating system, and the data only includes the log data collected under benign scenarios (i.e., scenarios without APT attacks).

[0042] Specifically, before model training, the SPADE infrastructure is adopted to collect fine-grained log data. SPADE provides software for inferring, storing, and querying data sources. It is cross-platform and can be used for multiple sources, including Linux, macOS, and Windows. The collection of sources can be completed without any changes to the application or the target platform. SPADE is easy to install and configure, and provides a simple mechanism for users to select from multiple storage formats.

[0043] Step 2: Extract five meta-properties from the log data. According to the five meta-properties, transform the log data into a heterogeneous graph based on the preset edge generation rules. The nodes in the heterogeneous graph are log entries.

[0044] For each log entry, define five meta-properties, including: source object, operation type, target object, time, and host. And according to three relationships between log entries, including: causal relationship, sequential relationship, and logical relationship of log entries, formulate the corresponding edge generation rules, and then construct the log entries into a heterogeneous graph according to the properties of the log entries and the edge generation rules.

[0045] Specifically, perform preliminary processing on the collected log entries to extract five meta-properties of the log. Among them, the source object represents the set of execution objects of the log (such as processes, threads, etc.), and the target object represents the set of objects operated on in the log (such as processes, files, sockets, inter-process communication, memory, network); the operation type represents the set of operation types of the source object on the target object (such as fork, execute, exit, clone operations between processes, read, open, close, write, loadlib, create, unlink operations between processes and files, connect, send operations between processes and networks, mprotect, mmap operations between processes and memory, connect, send, recv, read, close, accept, write operations between processes and sockets, etc.); the time represents the time when the log occurs, and the host is the operating system where the log occurs (such as a computer or a server).

[0046] Then, rules for three relationships between log entries are generated, considering different combinations of these meta-properties to associate fewer log entries and map fine-grained log relationships to a graph. Some of these rules are given in this embodiment and can be supplemented later:

[0047] Rule 1: Connect the log entries of the same source object on the same day in chronological order to form a sequence of daily log entries. This rule has the highest weight.

[0048] Rule 2: Connect the log entries of the same target object on the same day in chronological order to form a sequence of daily log entries. This rule has the highest weight.

[0049] Rule 3: Connect the sequences of daily log entries with the same source object and a similarity higher than the threshold. The similarity between sequences of daily log entries is positively correlated with the similarity of the number of log entries they contain.

[0050] Rule 4: Connect the sequences of daily log entries with the same target object and a similarity higher than the threshold. The similarity between sequences of daily log entries is positively correlated with the similarity of the number of log entries they contain.

[0051] Understanding that having the highest weight in Rule 1 and Rule 2 means that log entries within the same day must be connected, so the weight value is the highest, and the sequences between days need to be selected according to similarity.

[0052] Rule 1 connects the log sequences within a day to obtain a sequence of log entries for a day. Rule 3 calculates the similarity (the number of log entries) between these sequences connected by Rule 1 between days to determine which sequences of days are more similar, and the similar sequences can be connected. And the connection between days does not need to be in chronological order. The connection between days requires the head and tail nodes to connect to the log sequences (head and tail nodes) of other days according to similarity.

[0053] And it can be set that the similarity between sequences of daily log entries is the same as the similarity of the number of log entries they contain. For example, if the number of log entries contained in two sequences is the same, the similarity of the two sequences is 100%. For example, if the number of log entries contained in two sequences differs by 5, the similarity of the two sequences is 95%. It should be noted that the calculation method of the similarity of the number of log entries in two sequences can be selected according to actual needs.

[0054] Step 3: Use graph embedding technology to extract the context relationship of each log entry in the heterogeneous graph to form a log sequence, and regard the log sequence as a sentence. Use the word2vec method to convert each log entry in the log sequence into a low-dimensional vector representation.

[0055] In this embodiment, the graph embedding technology specifically uses a random walk method. Random walk is a popular graph traversal algorithm. Suppose a walker stays at a node in the graph. The next node to visit is determined according to the weight and type of each edge. The path (node sequence) generated by it is regarded as the context of these nodes. The word2vec model is used to calculate the vector of each node and its path (context). Specifically, these paths are regarded as sentences in natural language and processed by the model to learn the vector of each word (node). This method maintains the proximity between the node and its context, which means that the node (log entry) and its neighbors (log entries closely related to it) share similar embeddings (vectors).

[0056] The random walk method adopted in this embodiment can extract the context of each log entry from the heterogeneous graph. This method solves the anomaly detection problem in different attack scenarios by controlling the number of neighbor nodes and extracting the context with different sets of edge types in different proportions. Random walk generates paths according to the edge type and weight. The weight of the edge is proportional to the similarity of the number of logs contained in the two sets of log sequences.

[0057] In this embodiment, the definitions of edge type and edge weight in the random walk method are as follows:

[0058] Edge type: The edge types of the edges generated based on Rule 1 and Rule 2 are classified as Type 1, and the edge types of the edges generated based on Rule 3 and Rule 4 are classified as Type 2, generating two sets of edge types {Rule 1, Rule 3} and {Rule 2, Rule 4}.

[0059] This embodiment improves the conventional random walk method. When the random walk generates paths, it will generate different numbers of paths according to the proportion of the set of edge types (Rule 1 + Rule 3 or Rule 2 + Rule 4) we set, so that different attack methods can be focused on according to different scenarios.

[0060] Edge weight: Each time the random walk is executed, one of the two sets of edge types is taken. Therefore, the edge weight function is expressed as:

[0061]

[0062] w(T,V)=e -s(m,n)

[0063] Wherein, w(T, V) represents the weight of the sequence T of log entries and the sequence V of log entries on the corresponding set of edge types. The sequence T is a sequence of log entries generated by connecting rules where the edge type in the set of edge types is of type 1, and the sequence V is a sequence of log entries generated by connecting rules where the edge type in the set of edge types is of type 2. m is the number of log entries contained in the sequence T, n is the number of log entries contained in the sequence V, and s(m, n) is the similarity of the number of log entries between the sequence T and the sequence V.

[0064] And the transition probability calculation for the node v in the random walk method is as follows:

[0065]

[0066] Wherein, P(t|v) represents the transition probability from node v to node t, E is the set of all edges in the heterogeneous graph, N(v) is the neighbor nodes of node v, and W N(v) is the sum of the edge weights between node v and all its neighbor nodes N(v). is the edge type between node v and node t, and S n is the set of all edge types, and R w(t,v) is the sorting value (essentially limiting the number of neighbor nodes) of the edge weights of the nodes connected to node v based on the edge type in descending order among the edge weights of all nodes connected to node v, and neigh is the preset number of neighbor nodes.

[0067] When calculating the transition probability between nodes, if the nodes are sequences between days, then take the transition of the first / last node of the sequence to represent the transition between sequences. At this time, take w(t, v) as w(T, V), that is, node t is the first / last node of sequence T, and node v is the first / last node of sequence V.

[0068] N(v) as the neighbor nodes of node v needs to meet the following three conditions:

[0069] 1) Node t needs to have at least one edge with node v.

[0070] 2) There is at least one edge type between node t and node v that belongs to S n .

[0071] 3) w(t, v) represents the weight of node t and node v on the edge type , and it meets the condition if and only if its sorting value is not greater than neigh.

[0072] The node t that meets the above three conditions is used as the neighbor node N(v) of node v.

[0073] Note that for each random walk, only a set of edge type collections is considered. This definition can determine the importance of each set of edge types based on different scenarios, allowing for the differential extraction of the context of each node. Different attack scenarios will leave different traces for detection. That is, in each attack, only one or two meta - attributes of the log entries become abnormal. In each set, there is one type that focuses on within - group (e.g., Rule 1, Rule 2) and another type that focuses on between - group (e.g., Rule 3, Rule 4).

[0074] As mentioned above, each group includes two edge types, focusing on within - group (same day) and between - group (day - to - day) respectively. The edge types within the group construct a sequence of log entries, and each node in this sequence has two neighbor nodes. The random walk will not visit the last node that has already been visited. So, in fact, there is only one node to visit in the sequence. For the first and last nodes of this sequence, the log sequences to be associated need to be selected from the head and tail nodes of multiple other sequences. Therefore, the number of neighbor nodes (neigh) needs to be controlled. By tuning the value of neigh, a small number of malicious log sequences can be isolated from the vector space. At the same time, the optimal value of neigh needs to vary based on different users. So, the value of neigh is a hyperparameter, set by experiments. Generally, a lower value is set to ensure that the most similar sequences are connected to a given sequence.

[0075] Next, the word2vec model skip - gram is used to map log entries from the paths generated in the heterogeneous graph to low - dimensional vectors. It is an objective function that maximizes the probability of its neighbors conditioned on a node.

[0076] Step 4: Based on the low - dimensional vectors of all nodes in the heterogeneous graph, use graph sketching techniques to generate a compact and fixed - size sketch.

[0077] Using graph sketching techniques can support real - time stream analysis of growing heterogeneous graphs while maintaining the Jaccard similarity before and after the vector representation of log entries, which is beneficial for the analysis and clustering of heterogeneous graphs.

[0078] Specifically, since similarity calculation can be very expensive, scalable methods are needed to estimate the similarity of large - scale data sets. Graph sketching techniques construct a compact and fixed - size sketch for the vector representation of the log entry set in the heterogeneous graph. Graph sketching techniques are a constant - time algorithm, so it is fast enough to support real - time stream analysis of rapidly growing heterogeneous graphs. It also maintains the Jaccard similarity before and after the vector representation of log entries, allowing us to effectively calculate the similarity between sketches.

[0079] Step 5: Compare the sketch with the pre-trained detection model. If the sketch fits into one of the clusters in the evolutionary model under the detection model, the sketch is normal, that is, the obtained log data is the log data in the scenario without APT attacks; otherwise, the sketch is abnormal, that is, the obtained log data is the log data in the APT attack scenario.

[0080] In the actual detection scenario, the pre-trained detection model will not be modified, which can prevent the detection model from being poisoned by attack data. When deploying detection, for the increasing log entries, the heterogeneous graph needs to be incrementally updated regularly, and its corresponding sketch is calculated and generated, and compared with all the evolutionary models learned during the training process to make it fit into one of the clusters in an evolutionary model.

[0081] If the sketch obtained in real-time detection cannot be clustered into one of the clusters of the pre-trained evolutionary model, this sketch is considered abnormal. When an abnormal behavior occurs, the system will generate an alarm. The two abnormal behaviors detected in this embodiment include: the sketch that does not fit into the cluster corresponding to the final state and the sketch that does not fit into the cluster corresponding to the transition state.

[0082] Since the sketches generated by similar operations will be clustered into the same cluster after clustering, after a new operation appears, there may be a state of transition from one operation to another, and this state is also regarded as normal behavior. It is very difficult to ensure whether the state to be detected currently is a final state or a transition state during real-time detection, so these states all require sketch maintenance.

[0083] In this embodiment, when pre-training the detection model, the training process is as follows:

[0084] Regularly obtain the log data in the scenario without APT attacks, and obtain the sketch of this training according to Steps 1-4. Arrange the sketch of this training and the sketches of previous trainings in chronological order of creation to form a sketch sequence.

[0085] Use the clustering algorithm to cluster the sketch sequence to obtain multiple clusters. In this embodiment, the K-medoids algorithm is used to cluster the sketch sequence, and the silhouette coefficient is used to determine the optimal value of K.

[0086] Generate the evolutionary model of this training based on the chronological order of the sketches in each cluster and the statistical data of each cluster (such as diameter, medoid), and aggregate the evolutionary models obtained from multiple trainings as the detection model.

[0087] During data training, only the log entries in the benign scenario are trained. A series of sketches with temporal relationships are created during the training process. For each sketch, an evolutionary model is created, and this evolutionary model captures the update of the execution state during system operation. The final detection model consists of multiple evolutionary models in the training data.

[0088] Given a sketch and a similarity metric, clustering is a common data mining method for identifying outliers. The duration of an APT attack scenario is long enough that it cannot capture the evolutionary behavior of the system, which can lead to too many false positives. This embodiment utilizes its stream processing ability to create an evolutionary model to capture the normal changes in system behavior. Importantly, these evolutionary models are built during training rather than during deployment because there is a risk that a dynamically evolving model can be poisoned during the extended attack phase of an APT, thereby reducing the accuracy of detection.

[0089] The APT real-time detection and analysis method based on a heterogeneous graph provided in this application constructs a heterogeneous graph from log entries, represents the log entries as low-dimensional vectors using a graph embedding method, utilizes graph summarization techniques to keep the size of the graph sketch fixed and compact, models the graph sketch, and detects abnormal behavior through a clustering algorithm, achieving the efficiency, accuracy, and real-time performance of APT detection.

[0090] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0091] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A real-time detection and analysis method for APT based on heterogeneous graphs, characterized in that, the real-time detection and analysis method for APT based on heterogeneous graphs includes: Step 1, obtain log data in the operating system; Step 2, extract five meta-attributes from the log data, and based on the five meta-attributes, convert the log data into a heterogeneous graph according to a preset edge generation rule, where the nodes in the heterogeneous graph are log entries; among them, the five meta-attributes are source object, operation type, target object, time, and host, and the preset edge generation rules include: Rule 1, connect the log entries of the same source object on the same day in chronological order to form a sequence of daily log entries, and this rule has the highest weight; Rule 2, connect the log entries of the same target object on the same day in chronological order to form a sequence of daily log entries, and this rule has the highest weight; Rule 3, connect the sequences of daily log entries with the same source object and a similarity higher than the threshold; Rule 4, connect the sequences of daily log entries with the same target object and a similarity higher than the threshold; Step 3, use graph embedding technology to extract the context relationship of each log entry in the heterogeneous graph to form a log sequence, and regard the log sequence as a sentence and use the word2vec method to convert each log entry in the log sequence into a low-dimensional vector representation; among them, the use of graph embedding technology to extract the context relationship of each log entry in the heterogeneous graph to form a log sequence includes: The graph embedding technology is the random walk method, and the definitions of edge types and edge weights in the random walk method are as follows: The edge type: divide the edge types of the edges generated based on Rule 1 and Rule 2 into type 1, and divide the edge types of the edges generated based on Rule 3 and Rule 4 into type 2, and generate two sets of edge type sets as {Rule 1, Rule 3} and {Rule 2, Rule 4}; The edge weight: each time the random walk is executed, take one of the two sets of edge type sets, so the edge weight function is expressed as: w(T, V) = e -s(m,n) In the formula, w(T, V) represents the weight of the sequence T of log entries and the sequence V of log entries on the corresponding edge type set. The sequence T is the sequence of log entries generated by connecting according to the rule with edge type 1 in the edge type set, the sequence V is the sequence of log entries generated by connecting according to the rule with edge type 2 in the edge type set, m is the number of log entries included in the sequence T, n is the number of log entries included in the sequence V, and s(m, n) is the similarity of the number of log entries in the sequence T and the sequence V; Step 4, generate a sketch of a fixed size based on the low-dimensional vectors of all nodes in the heterogeneous graph; Step 5, compare the sketch with a pre-trained detection model. If the sketch fits into a cluster in the evolved model under the detection model, the sketch is normal, that is, the obtained log data is the log data in the scenario without APT attacks; otherwise, the sketch is abnormal, that is, the obtained log data is the log data in the APT attack scenario; Among them, the generation process of the detection model includes: Timely obtain the log data in the scenario without APT attacks to get the sketch of this training, and arrange the sketch of this training and the sketches of previous trainings in chronological order of creation time to form a sketch sequence; Use a clustering algorithm to cluster the sketch sequence to obtain multiple clusters; Generate the evolutionary model of this training based on the chronological order of the sketches in each cluster and the statistical data of each cluster, and aggregate the evolutionary models obtained from multiple trainings as the detection model.

2. The APT real-time detection and analysis method based on a heterogeneous graph according to claim 1, characterized in that, The similarity between the sequences of the daily log entries is positively correlated with the similarity of the number of log entries they contain.

3. The APT real-time detection and analysis method based on a heterogeneous graph according to claim 1, characterized in that, The transition probability calculation at node v in the random walk method is as follows: Where P(t|v) represents the transition probability from node v to node t, E is the set of all edges in the heterogeneous graph, N(v) is the neighbor node of node v, and W N(v) is the sum of the edge weights between node v and all its neighboring nodes N(v), w(t, v) = w(T, V), is the edge type between node v and node t, S n is the set of all edge types, R w(t,v) Based on edge type The edge weights of the nodes connected to node v are sorted in descending order among the edge weights of all nodes connected to node v, and neigh is the number of preset neighbor nodes.