Power monitoring system APT attack detection method based on graph neural network

By adopting a graph neural network-based detection method in the power monitoring system, APT detection is transformed into the classification problem of benign and malignant data traceability graph, the problem of high false alarm rate and reliance on prior knowledge in the prior art is solved, and the APT detection effect with high accuracy and low false alarm rate is achieved.

CN119989037APending Publication Date: 2025-05-13YANCHENG POWER SUPPLY CO STATE GRID JIANGSU ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411341233.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect APT attacks in power systems, especially when detection methods relying on prior knowledge cannot cope with new APTs, the false alarm rate is high and it is difficult to ensure the integrity of graph features.

Method used

The APT attack detection method of the power monitoring system based on graph neural network is adopted to transform the APT detection problem into the classification problem of benign and malignant data traceability maps. Through the combination of time-sequence segmentation of the data traceability map, InfoGraph unsupervised graph representation learning and LSTM model, APT detection with high accuracy and low false alarm rate is achieved.

Benefits of technology

It realizes a high APT detection accuracy and a low false alarm rate. It does not rely on prior knowledge and can effectively detect new APTs. It also solves the algorithm memory overhead problem through a timing segmentation strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989037A_ABST
    Figure CN119989037A_ABST
Patent Text Reader

Abstract

The invention provides an APT attack detection method for a power monitoring system based on a graph neural network. The APT attack detection method comprises the following steps: step 1, collecting a data traceability graph of a running system; step 2, storing and querying data by adopting a graph database; and step 3, performing benign and malicious classification on the feature sequence by using an LSTM model to realize APT detection. According to the APT attack detection method for the power monitoring system based on the graph neural network, the APT detection problem is converted into the classification problem of benign and malignant data traceability graphs, and the APT detection accuracy is high and the false alarm rate is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of cyberspace security technology, and in particular relates to an APT attack detection method for an electric power monitoring system based on a graph neural network. Background Art

[0002] With the access of a high proportion of power electronic equipment, the probability of new power system communication networks being attacked by network attacks has increased significantly. Among them, APT attacks with strong concealment, great destructive power and long duration have become one of the main threats to the security of new power system communication networks. Once the power grid is attacked by APT, it will not only cause large-scale power outages and affect production and life, but also cause the leakage of sensitive information, such as power operation data and user information, resulting in huge economic losses. Therefore, the power system should focus on improving the ability to identify APT attacks during network security operation monitoring.

[0003] Since APT attacks are slow and last for a long time, it is difficult for time series models to perceive the causal relationship between two events that are separated by a long time. Research in recent years has shown that data provenance graphs are a better data source for APT detection. Data provenance graphs describe the information flow between system entities (processes, etc.) and objects (files, sockets, etc.), and express the causal relationship between system events. Compared with traditional anomaly detection methods, detection models based on data provenance graphs can perceive the causal relationship between system events generated by APT attacks, even if the two events occur far apart in time, thereby better coping with the slow and persistent characteristics of APT attacks. Most of the current research uses prior knowledge to a certain extent, summarizes the TTP attack rules of APT or defines the attack features in the provenance graph. Although this detection method has a low false alarm rate, they all rely on prior expert knowledge and are difficult to detect new APTs that use new technologies. Although the StreamSpot and UNICORN systems do not rely on prior expert knowledge, both methods are essentially based on anomaly detection, that is, the model only learns the behavior of benign samples, distinguishes and detects abnormal behavior by the difference from the behavior of benign samples. Once the training samples cannot cover all normal behaviors, false alarms are likely to occur during actual detection. In addition, both systems use traditional graph algorithms to extract traceability graph features, which makes it difficult to ensure the integrity of the extracted graph features.

[0004] APT attacks are advanced, slow, meticulous, and highly concealed, which makes it difficult for traditional threat intrusion detection methods to deal with them. Detectors based on malicious code signatures cannot detect attacks based on 0-day vulnerabilities; typical anomaly detection systems record and analyze system call sequences and log sequences, but most of them have difficulty modeling the long-term behavior patterns of the system and perform poorly in APT detection. In recent years, APT detection systems based on data traceability graphs have achieved initial results, but there are still problems such as reliance on expert knowledge and high false alarm rates. Therefore, it is of great significance to design a new APT detection system. Summary of the invention

[0005] In order to solve the above problems, the present invention provides an APT attack detection method for an electric power monitoring system based on a graph neural network, which converts the APT detection problem into a classification problem of benign and malignant data traceability graphs, and has a higher APT detection accuracy and a lower false alarm rate.

[0006] The present invention is specifically a method for detecting APT attacks in a power monitoring system based on a graph neural network, which includes three parts: data provenance graph collection, data provenance graph management, and APT detection algorithm, and includes the following steps:

[0007] Step 1: Collect a data traceability diagram of a running system;

[0008] Step 2: Use graph database to store and query data;

[0009] Step 3: Use the LSTM model to classify the feature sequences into benign and malicious to achieve APT detection.

[0010] Furthermore, the step 1 specifically includes the following sub-steps:

[0011] Step 1.1: Process the generated log files and convert them into a time-series data traceability diagram that can reflect the causal relationship of events;

[0012] Step 1.2: Store the generated data provenance graph in the graph database by default.

[0013] Furthermore, the step 2 specifically includes the following sub-steps:

[0014] Step 2.1, the present invention uses a graph database to save the data traceability graph;

[0015] Step 2.2: Based on the time series data provenance graph segmentation strategy, the data provenance graph is segmented into multiple subgraphs according to the time sequence of the edges to form a subgraph sequence.

[0016] Furthermore, the step 3 specifically includes the following sub-steps:

[0017] Step 3.1, use one-hot encoding to vectorize the subgraph according to the subtypes of nodes and edges to obtain a vectorized subgraph sequence;

[0018] Step 3.2: Use the InfoGraph unsupervised algorithm to train the GINE graph neural network, so that GINE can extract feature vectors that can represent the subgraph structure information and generate a subgraph feature sequence.

[0019] Step 3.3: Use the LSTM model to classify the feature sequences into benign and malicious to achieve APT detection.

[0020] Compared with the prior art, the beneficial effects are:

[0021] 1. The algorithm memory overhead problem is solved by introducing the time-series segmentation strategy of the data traceability graph.

[0022] 2. The InfoGraph unsupervised graph representation learning algorithm is introduced to train the GINE graph neural network to extract subgraph features, and the subgraph sequence is converted into a subgraph feature sequence, which can still complete efficient detection without prior knowledge.

[0023] 3. Convert the APT detection problem into a classification problem of benign and malicious data traceability graphs, which has a higher APT detection accuracy and lower false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a framework diagram of an APT attack detection method for an electric power monitoring system based on a graph neural network according to the present invention;

[0025] Figure 2 It is a schematic diagram of the APT detection algorithm flow of the present invention. DETAILED DESCRIPTION

[0026] The following is a detailed description of the specific implementation of a method for detecting APT attacks in a power monitoring system based on a graph neural network according to the present invention in conjunction with the accompanying drawings.

[0027] The present invention proposes a method for detecting APT attacks in a power monitoring system based on a graph neural network, comprising the following steps:

[0028] Step 1: Collect a data traceability diagram of a running system;

[0029] Step 1.1: Process the generated log files and convert them into a time-series data traceability diagram that can reflect the causal relationship of events;

[0030] The data provenance graph contains rich semantic information, such as process name, file name, file path, etc. The present invention does not consider this semantic information and only focuses on the types of nodes and edges.

[0031] Step 1.2: Store the generated data provenance graph in the graph database by default.

[0032] Step 2: Use graph database to store and query data;

[0033] Step 2.1, the present invention uses a graph database to save the data traceability graph;

[0034] Step 2.2: Based on the time series data provenance graph segmentation strategy, the data provenance graph is segmented into multiple subgraphs according to the time sequence of the edges to form a subgraph sequence;

[0035] The time-series-based data provenance graph segmentation strategy is divided into edge file reading and subgraph generation. In the edge file reading part, the edges are pre-read from the graph database in chronological order and stored in the file. The algorithm cyclically reads the edge files arranged in chronological order. The subgraph generation part generates a subgraph based on the threshold and inputs the subgraph into the subsequent APT detection model.

[0036] Step 3: Use the LSTM model to classify the feature sequences into benign and malicious to achieve APT detection;

[0037] Step 3.1, use one-hot encoding to vectorize the subgraph according to the subtypes of nodes and edges to obtain a vectorized subgraph sequence;

[0038] The present invention uses unique hot coding to vectorize nodes and edges for subtypes. If a certain node (edge) does not contain any subtype, then this node (edge) is a class of its own. Considering that the timestamp of the edge reflects the sequence and interval of events in the data traceability graph, and contains rich time series information, the present invention uses universal time coding to vectorize the timestamp in the edge;

[0039] Step 3.2: Use the InfoGraph unsupervised algorithm to train the GINE graph neural network, so that GINE can extract feature vectors that can represent the subgraph structure information and generate a subgraph feature sequence;

[0040] All subgraphs obtained after time series segmentation are used as sample sets of the InfoGraph algorithm for unsupervised graph representation training; the GINE graph encoder obtained after the InfoGraph algorithm converges can accurately extract the global representation of the subgraph, and use the global representation as the feature of the subgraph, thereby converting the subgraph sequence into a sequence of subgraph feature vectors, which is passed to the LSTM classifier for subsequent processing;

[0041] Step 3.3: Use the LSTM model to classify the feature sequence into benign and malicious to achieve APT detection

[0042] The GINE model obtained by InfoGraph training is regarded as a subgraph feature encoder with fixed parameters. The subgraph feature vector is extracted, and the subgraph sequence is converted into a feature sequence. The subgraph sequence is input into a multi-layer LSTM structure. The average of all outputs of the last layer of LSTM is taken and input into a fully connected layer binary classifier to realize the classification of subgraph feature sequences, thereby completing the classification of malicious graphs and benign graphs. Specific embodiments

[0044] This paper proposes a method for detecting APT attacks in power monitoring systems based on graph neural networks. Figure 1 As shown in the figure, ProDetector includes three parts: data traceability graph collection, data traceability graph management, and APT detection algorithm. First, SPADE is used to collect system data traceability graphs; in the data management module, Neo4j graph database is used to store and query data. When reading, the data traceability graph is segmented based on time series to form a subgraph sequence, which solves the problem of excessive memory usage of the detection algorithm; then, the GINE graph neural network model obtained by unsupervised training of the InfoGraph algorithm is used to extract features from each subgraph, and the subgraph sequence is converted into a feature sequence; finally, the LSTM model is used to classify the feature sequence into benign and malicious to achieve APT detection.

[0045] Specifically, the present invention comprises the following steps:

[0046] Step 1: Use SPADE to collect a data traceability diagram of a running system. The specific process is as follows:

[0047] Step 1.1: Use SPADE to process the log files generated by the system and convert them into a time-series data provenance graph that can reflect the causal relationship of events. The data provenance graph generated by SPADE also contains rich semantic information, such as process name, file name, file path, etc. ProDetector does not consider this semantic information and only focuses on the types of nodes and edges. SPADE supports the Open Provenance Model (OPM) data provenance graph format standard. The subtypes and detailed descriptions of nodes and edges in the OPM format are shown in Tables 1 and 2.

[0048] Table 1

[0049]

[0050] Table 2

[0051]

[0052] Step 1.2: SPADE stores the generated data provenance graph in the Neo4j graph database by default.

[0053] Step 2: Use Neo4j graph database to store and query data. When reading, the data traceability graph is segmented into subgraph sequences based on time series, which solves the problem of excessive memory usage of the detection algorithm. The specific process is as follows:

[0054] Step 2.1, use the graph database Neo4j that SPADE relies on to save the data traceability graph. Neo4j is an efficient cross-platform graph database that allows nodes and edges to exist in multiple types and contain multiple attributes. It supports basic query functions such as node and edge queries, as well as graph algorithms such as depth and breadth-first search, label propagation, and connected component query;

[0055] Step 2.2, propose a time-series-based data provenance graph segmentation strategy. ProDetector divides the data provenance graph into multiple subgraphs according to the time sequence of the edges to form a subgraph sequence. The data provenance graph segmentation scale (the maximum number of edges contained in the subgraph) is determined by the pre-set parameter threshold, which can be adjusted according to the model training results.

[0056] Step 3: Use the GINE graph neural network model obtained through unsupervised training of the InfoGraph algorithm to extract features from each subgraph, convert the subgraph sequence into a feature sequence, and finally use the LSTM model to classify the feature sequence into benign and malicious to achieve APT detection. The APT detection algorithm process is as follows: Figure 2 The specific process is as follows:

[0057] Step 3.1, use one-hot encoding to vectorize nodes and edges. If a certain node (edge) does not contain any subtypes, then this node (edge) is a class in itself. Considering that the timestamp of the edge reflects the sequence and interval of events in the data traceability graph, and contains rich time series information, the present invention uses universal time encoding to vectorize the timestamp in the edge;

[0058] Step 3.2: Use the InfoGraph unsupervised graph representation learning algorithm to learn subgraph features. InfoGraph regards node features as local representations (patch representation). GNN uses the READOUT operator to aggregate the features of all nodes to obtain the features of a fixed-length graph. InfoGraph regards graph features as global representations. The core idea of ​​InfoGraph is to maximize the mutual information between the global representation and local representation of the same graph, and minimize the mutual information between the global representation and local representation of different graphs. The GINE graph encoder obtained after the InfoGraph algorithm converges can accurately extract the global representation of the subgraph, and use the global representation as the feature of the subgraph, thereby converting the subgraph sequence into a sequence of subgraph feature vectors, which is passed to the LSTM classifier for subsequent processing;

[0059] Step 3.3: Use the LSTM model to classify the obtained subgraph feature sequence into malicious and benign. The GINE model obtained by InfoGraph training is regarded as a subgraph feature encoder with fixed parameters. The subgraph feature vector is extracted, and the subgraph sequence is converted into a feature sequence, which is input into the multi-layer LSTM structure. The outputs of the last layer of LSTM are averaged and input into a fully connected layer binary classifier to realize the classification of subgraph feature sequences, thereby completing the classification of malicious and benign graphs.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit the present invention. A person skilled in the art should understand that the specific implementation of the present invention can be modified or replaced by equivalents, but these modifications or changes are within the scope of protection of the pending claims.

Claims

1. A method for detecting APT attacks in power monitoring systems based on graph neural networks, characterized in that: The power monitoring system APT attack detection method comprises the following steps: Step 1: Collect a data traceability diagram of a running system; Step 2: Use graph database to store and query data; Step 3: Use the LSTM model to classify the feature sequences into benign and malicious to achieve APT detection.

2. According to claim 1, a method for detecting APT attacks in a power monitoring system based on graph neural network is characterized in that: The step 1 specifically includes the following sub-steps: Step 1.1: Process the generated log files and convert them into a time-series data traceability diagram that can reflect the causal relationship of events; Step 1.2: Store the generated data provenance graph in the graph database by default.

3. According to claim 2, a method for detecting APT attacks in a power monitoring system based on a graph neural network is characterized in that: The data provenance graph in step 1.1 contains rich semantic information, and the present invention uses the types of nodes and edges.

4. According to claim 3, a method for detecting APT attacks in a power monitoring system based on graph neural network is characterized in that: The step 2 specifically includes the following sub-steps: Step 2.1, the present invention uses a graph database to save the data traceability graph; Step 2.2: Based on the time series data provenance graph segmentation strategy, the data provenance graph is segmented into multiple subgraphs according to the time sequence of the edges to form a subgraph sequence.

5. According to claim 4, a method for detecting APT attacks in a power monitoring system based on graph neural network is characterized in that: The time-series-based data provenance graph segmentation strategy in step 2.2 is divided into edge file reading and subgraph generation. In the edge file reading part, the edges are read from the graph database in chronological order in advance and stored in the file, and the algorithm cyclically reads the edge files arranged in chronological order; the subgraph generation part generates a subgraph based on a threshold value and inputs the subgraph into the subsequent APT detection model.

6. The method for detecting APT attacks in a power monitoring system based on graph neural network according to claim 5 is characterized in that: The step 3 specifically includes the following sub-steps: Step 3.1, use one-hot encoding to vectorize the subgraph according to the subtypes of nodes and edges to obtain a vectorized subgraph sequence; Step 3.2: Use the InfoGraph unsupervised algorithm to train the GINE graph neural network, so that GINE can extract feature vectors that can represent the subgraph structure information and generate a subgraph feature sequence; Step 3.3: Use the LSTM model to classify the feature sequences into benign and malicious to achieve APT detection.

7. The method for detecting APT attacks in a power monitoring system based on graph neural network according to claim 6 is characterized in that: In the step 3.1, the present invention uses one-hot encoding for subtypes to vectorize nodes and edges. If a certain node does not contain any subtypes, then this node is a class in itself; the present invention uses universal time encoding to vectorize the timestamps in the edges.

8. The method for detecting APT attacks in a power monitoring system based on graph neural network according to claim 7 is characterized in that: In the step 3.2, all subgraphs obtained after time series segmentation are used as sample sets of the InfoGraph algorithm for unsupervised graph representation training; the GINE graph encoder obtained after the InfoGraph algorithm converges can accurately extract the global representation of the subgraph, and use the global representation as the feature of the subgraph, thereby converting the subgraph sequence into a sequence of subgraph feature vectors, which is passed to the LSTM classifier for subsequent processing.

9. The method for detecting APT attacks in a power monitoring system based on graph neural network according to claim 8 is characterized in that: In step 3.3, the GINE model obtained by InfoGraph training is regarded as a subgraph feature encoder with fixed parameters, and the subgraph feature vector is extracted. The subgraph sequence is converted into a feature sequence and input into a multi-layer LSTM structure. The average of all outputs of the last layer of LSTM is taken and input into a fully connected layer binary classifier to realize subgraph feature sequence classification, thereby completing the classification of malicious graphs and benign graphs.