Internal Threat Detection Method Based on Knowledge Graph and Residual Graph Convolutional Network
By building an internal threat knowledge graph and using residual graph convolution network, combining user communication relationships and behavioral characteristics, the problem of poor detection accuracy and isolated nodes caused by ignoring user communication relationships in the prior art is solved, and more accurate and comprehensive internal threat detection is achieved.
Patent Information
- Application Number
- CN202410248725.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-03-05
AI Technical Summary
Existing internal threat detection methods ignore the communication relationship between users, resulting in poor accuracy of detection results and the existence of isolated nodes in users.
The detection method based on knowledge graph and residual graph convolution network is adopted. By collecting user feature information and preprocessing, an internal threat knowledge graph is constructed, and an adjacency matrix weighted by user communication relationships and behavior similarity is used to train the feature matrix and adjacency matrix in combination with the residual graph convolution network to obtain the detection results.
By conducting comprehensive analysis from the perspectives of user behavior and relationships, the accuracy and comprehensiveness of the internal threat detection model are improved, and the problem of neglected communication relationships between users is solved and the existence of isolated nodes is reduced.
Smart Images

Figure CN118018304B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network and information security technology, and particularly to an internal threat detection method based on a knowledge graph and a residual graph convolutional network. Background Art
[0002] Internal threats are destructive and stealthy, making them a challenging task in network security. Although previous internal threat detection work has achieved excellent results by detecting through user behavior characteristics, these methods still have many drawbacks. Existing methods are insufficient to handle high-dimensional, complex, heterogeneous, and sparse data, with poor model generalization and few applicable scenarios. In addition, existing methods are based on the assumption that the daily behavior of users in the time series is relatively regular and stable, and identify anomalies by comparing subsequent behaviors with daily behaviors, which does not represent the real situation. Communication relationships can provide valuable and necessary information, similar to the social network in our daily life. However, existing methods ignore the communication relationships between users in the internal threat detection task, thus reducing the effect of threat detection. Therefore, it is crucial to make the most of limited data. Summary of the Invention
[0003] The purpose of the present invention is to provide an internal threat detection method based on a knowledge graph and a residual graph convolutional network, aiming to solve the problem that existing methods ignore the communication relationships between users in the internal threat detection task, resulting in isolated nodes among users.
[0004] To achieve the above purpose, in the first aspect, the present invention provides an internal threat detection method based on a knowledge graph and a residual graph convolutional network, including the following steps:
[0005] Collect user feature information and perform preprocessing to obtain internal user data;
[0006] Construct an internal threat knowledge graph ontology based on the internal user data;
[0007] Construct an internal threat knowledge graph using a graph database based on the internal threat knowledge graph ontology;
[0008] Construct a feature matrix based on the internal threat knowledge graph;
[0009] Construct a weighted function based on an adjacency matrix weighted by user communication relationships and user behavior similarity, and construct an adjacency matrix;
[0010] Train the adjacency matrix and the feature matrix using an internal threat detection method based on a residual graph convolutional network to obtain a detection result.
[0011] Wherein, the user feature information includes resume information, personality information, communication records, and behavior records.
[0012] Among them, the preprocessing includes data cleaning and filling of missing values.
[0013] Among them, constructing the internal threat knowledge graph ontology based on the internal user data includes:
[0014] Converting the discrete data in the internal user data into one-hot encoded vectors;
[0015] Constructing the internal threat knowledge graph ontology based on the one-hot encoded vectors.
[0016] Among them, the graph database includes the Neo4j graph database.
[0017] Among them, the residual graph convolutional network is composed of three layers of graph convolutional network layers and a residual module, and three layers of multi-layer perceptron layers are added before the initial residual jumps to connect to the third layer of the graph convolutional network layer.
[0018] In a second aspect, the present invention provides an internal threat system based on a knowledge graph and a residual graph convolutional network, including an ontology modeling module, a data collection and preprocessing module, a knowledge graph construction and storage module, and a knowledge graph application module;
[0019] The ontology modeling module is used to design the ontology structure from the dimension of internal users as the knowledge organization mode of the knowledge graph by analyzing the construction purpose of the knowledge base and the characteristics of the data source, so as to obtain the internal threat knowledge graph ontology;
[0020] The data collection and preprocessing module is used to obtain and collect user feature information, and then perform data cleaning and filling of missing values on the user feature information to obtain internal user data;
[0021] The knowledge graph construction and storage module extracts entities, relationships, and attributes from the internal user data based on rules, creates data instances of the internal threat knowledge graph ontology, and then uses the Py2neo toolkit to store the data of the knowledge graph in the Neo4j graph database to obtain the constructed internal threat knowledge graph;
[0022] The knowledge graph application module uses the visualization interface of the Neo4j database to perform query and export operations, filters the data in the required graph using the Export command in the APOC library, and exports it as a csv file as the input for the subsequent internal threat detection model.
[0023] An internal threat detection method based on a knowledge graph and a residual graph convolutional network of the present invention preprocesses by collecting user feature information to obtain internal user data; constructs an internal threat knowledge graph ontology based on the internal user data; constructs an internal threat knowledge graph using a graph database based on the internal threat knowledge graph ontology; constructs a feature matrix based on the internal threat knowledge graph; constructs an adjacency matrix by constructing a weighting function based on an adjacency matrix weighted by user communication relationships and user behavior similarities; trains the adjacency matrix and the feature matrix using an internal threat detection method based on a residual graph convolutional network to obtain a detection result. By the above method, the present invention addresses the problem that traditional detection models only focus on the behavior of users themselves and ignore the association relationships between users, resulting in poor accuracy of detection results, and realizes comprehensive analysis from two perspectives of user behavior and relationships, improving the accuracy and comprehensiveness of the internal threat detection model, and solving the problem that existing methods ignore the communication relationships between users in the internal threat detection task, resulting in isolated nodes among users. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0025] Figure 1 is a flowchart of an internal threat detection method based on a knowledge graph and a residual graph convolutional network provided by the present invention.
[0026] Figure 2 is an architecture diagram of an internal threat knowledge graph construction method of an internal threat detection system based on a knowledge graph and a residual graph convolutional network provided by the present invention.
[0027] Figure 3 is an internal threat ontology structure diagram.
[0028] Figure 4 is an overall framework diagram of an internal threat detection model based on a residual graph convolutional network.
[0029] 1 - Ontology modeling module, 2 - Data acquisition and preprocessing module, 3 - Knowledge graph construction and storage module, 4 - Knowledge graph application module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0031] Please refer to Figures 1 to 4 , Figure 1 which is a flowchart of an internal threat detection method based on a knowledge graph and a residual graph convolutional network provided by the present invention. Figure 2 which is an architecture diagram of an internal threat knowledge graph construction method for an internal threat detection system based on a knowledge graph and a residual graph convolutional network provided by the present invention. Figure 3 which is an internal threat ontology structure diagram. Figure 4 which is an overall framework diagram of an internal threat detection model based on a residual graph convolutional network.
[0032] In a first aspect, the present invention provides an internal threat detection method based on a knowledge graph and a residual graph convolutional network, including the following steps:
[0033] S1: Collect user feature information and perform preprocessing to obtain internal user data;
[0034] Specifically, the user feature information includes resume information, personality information, communication records, behavior records, etc. The preprocessing includes data cleaning and filling of missing values.
[0035] S2: Construct an internal threat knowledge graph ontology based on the internal user data;
[0036] Specifically, convert the discrete data in the internal user data into one-hot encoded vectors; construct an internal threat knowledge graph ontology based on the one-hot encoded vectors.
[0037] S3: Use a graph database to construct an internal threat knowledge graph based on the internal threat knowledge graph ontology;
[0038] Specifically, the graph database includes a Neo4j graph database.
[0039] S4: Construct a feature matrix based on the internal threat knowledge graph;
[0040] S5: Construct a weighted function based on the adjacency matrix weighted by user communication relationships and user behavior similarities, and construct an adjacency matrix;
[0041] S6: Train the adjacency matrix and the feature matrix using an internal threat detection method based on a residual graph convolutional network to obtain a detection result.
[0042] Specifically, the feature matrix and the adjacency matrix are used as the inputs of the residual graph convolutional network; the feature matrix and the adjacency matrix are passed to the graph convolutional network layer of the residual graph convolutional network. A residual unit is introduced before the first graph convolutional layer. The input of the first graph convolutional layer is used as the initial residual, and the initial residual is skip-connected to the input of the third graph convolutional layer of the residual graph convolutional network. Three layers of MLP are added in the middle to perform flexible transformation on the features, which helps to adapt to different data patterns; cross-entropy is used as the loss function for evaluation during training to obtain the final internal threat detection result.
[0043] The overall research idea framework diagram of the internal threat detection model based on the residual graph convolutional network is as Figure 4 shown.
[0044] The input of the model (residual graph convolutional network) proposed in the present invention can be divided into two parts. One is the feature matrix, which can be represented as a matrix X of N×T. N is the number of nodes on the graph, that is, the number of users, and T is the number of features of each user node. The other is the adjacency matrix A of N×N, which represents the structural information of the graph. As mentioned above, the matrix data derived from the knowledge graph can be used as the input of the detection model. However, since there are many isolated nodes in the network and relying only on communication relationships to establish connections may lead to a decrease in the robustness and anomaly detection ability of the model, the present invention adopts a comprehensive weighting method to establish connections between users before using the graph data. Neo4j provides the GraphData Science library, which contains the NodeSimilarity algorithm, and the similarity between nodes can be calculated according to their attributes.
[0045] Since the cosine similarity function performs better in high-dimensional data and matches the high-dimensional nature of the internal threat information, the present invention chooses to use the cosine similarity method in Node Similarity to calculate S i,j , and the cosine similarity uses the cosine of the angle between two vectors to calculate the similarity. The formula is as follows:
[0046]
[0047] where S i,j is the similarity function for calculating the feature similarity of two nodes; x i and x j represent user node i and user node j respectively.
[0048] The output of the GCN model is represented as a matrix of N×F, where F is the number of classification categories, and each node belongs to one of the F categories. Considering the actual situation that the internal threat detection task belongs to binary classification, the value of F here is 2.
[0049] The designed network model of this module consists of three layers of graph convolutional network layers and residual modules. Among them, the input matrix needs to be transformed through multiple hidden convolutional layers in the GCN. Its forward propagation formula is as follows:
[0050] H (l+1) = f(H (l) , A) = σ(AH (l) W (l) ) (2)
[0051] Among them, H (l) is the node representation of each layer of the GCN. l represents the number of hidden layers. σ is a non-linear function, such as ReLU. W (l) is the learning parameter of the l-th layer. When l = 0, H (0) = X ∈ R N×T , which is the feature matrix of the input user.
[0052] In each hidden layer, the input matrix calculation formula is expressed as follows:
[0053]
[0054] Among them, represents the feature set of user node i in the (l + 1)-th layer, represents the feature set of user node j in the l-th layer. N i represents the set of neighbors (including itself) of user node i, represents the weight parameter used to update the information from user node j, and c ij represents the normalization coefficient.
[0055] In the network model designed by the present invention, the propagation process of the GCN layer is specifically expressed as follows:
[0056]
[0057] Among them, ReLU is the activation function selected by this model. W (n) respectively represent the learning parameters of the n-th layer, is the matrix obtained by normalizing and standardizing the adjacency matrix A. The specific calculation process of this matrix can be expressed by the following formula:
[0058]
[0059]
[0060] Among them, I represents the identity matrix, An adjacency matrix indicating self-connection, abbreviated as a self-connected adjacency matrix. Since there is no connection between nodes in the original adjacency matrix A, the diagonal in A is 0. If it is input into the model for downstream calculations, the model cannot distinguish between self-nodes and unconnected nodes in the adjacency matrix. Therefore, the adjacency matrix is added with the identity matrix to obtain the matrix and then perform a normalization operation. is the degree matrix. The role of the degree matrix is to perform a normalization process.
[0061] The residual module of GCN is mainly reflected in introducing a skip connection in a residual unit. For the (l + 1)-th layer, in addition to the output of the l-th layer as the input, a skip connection before the l-th layer is added to prevent calculation deviation and improve the aggregation efficiency. The basic formula of GCN with a residual module can be expressed as:
[0062]
[0063] where α is a learnable parameter, usually called the residual weight. This parameter controls the weight of the initial residual in the residual connection, that is, the contribution degree of the input feature to the final output.
[0064] The initial motivation for introducing the residual is to prevent the problem of gradient vanishing. In the present invention, the feature matrix H (0) = X ∈ R N×T is selected as the initial residual, and the initial residual is skip-connected to the output of the third layer H (3) . The output of the third layer is expressed by the formula:
[0065]
[0066] The parameter α controlling the initial residual should not be too large. Generally, it is more appropriate to set it at 0.1 or 0.2. If α is too large, it will weaken the role of the upper-layer calculation on the current-layer calculation and seriously affect the learning efficiency. If the residual connection is not applied in the GCN network, the feature homogenization of nodes will occur quickly, resulting in a smoothing phenomenon.
[0067] In addition, in the present invention, a 3-layer MLP is added before the initial residual is skip-connected to the third layer. Such a design allows the network to learn more flexible transformations of the input features, which helps to adapt to different data patterns. It is expressed by the formula as:
[0068] H (MLP) = W 3 (W 2 (W 1 X + b 1 ) + b 2 ) + b 3 (9)
[0069] Among them, W (1-3) corresponds to the weights of three MLP layers respectively, and b (1-3) corresponds to the biases of three MLP layers respectively. To prevent overfitting, the Dropout regularization method is introduced before classification to prevent the model from overfitting and improve the generalization ability of the model.
[0070] Finally, to calculate the classification output of each node, the present invention uses softmax as the activation function in the output layer to output the classification probability distribution, as shown in Formula 10:
[0071] Z = softmax(WH (3) + b) (10)
[0072] The formula of the softmax activation function is as follows:
[0073]
[0074] To better measure the difference between the model output and the true label to encourage the model to learn a better representation, the present invention uses the cross-entropy function shown in Formula 12 as the loss function to evaluate during training, where represents the overall loss, y L represents the set of labeled samples, l represents the node elements in the set, F represents the number of categories, f represents the category, and Y lf represents the actual label in sample l, and Z lf represents the predicted label of the sample.
[0075]
[0076] In a second aspect, the present invention provides an internal threat system based on a knowledge graph and a residual graph convolutional network, including an ontology modeling module 1, a data collection and preprocessing module 2, a knowledge graph construction and storage module 3, and a knowledge graph application module 4;
[0077] The ontology modeling module 1 is used to design the ontology structure from the dimension of internal users as the knowledge organization mode of the knowledge graph by analyzing the construction purpose of the knowledge base and the characteristics of the data source, so as to obtain the internal threat knowledge graph ontology;
[0078] The data collection and preprocessing module 2 is used to obtain and collect user feature information, and then perform data cleaning and filling of missing values on the user feature information to obtain internal user data;
[0079] The knowledge graph construction and storage module 3 extracts entities, relationships, and attributes from the internal user data based on rules, creates data instances of the internal threat knowledge graph ontology, and then uses the Py2neo toolkit to store the knowledge graph data in the Neo4j graph database to obtain the constructed internal threat knowledge graph;
[0080] The knowledge graph application module 4 uses the visualization interface of the Neo4j database for querying and exporting operations. The Export command in the APOC library filters the data in the required graph and exports it as a csv file as the input for the subsequent internal threat detection model.
[0081] Specifically, the idea of the internal threat knowledge graph construction method is to first construct the internal threat knowledge graph ontology, then extract important user features from the original log data file, and aggregate data from different sources according to the user ID with given aggregation conditions. Subsequently, feature extraction is performed on the aggregated data to generate numerical vectors and extract user entities and the relationships between user entities, which are then integrated into the knowledge graph to finally construct the internal threat knowledge graph. By extracting the feature matrix of users and the relationships between all user entities from the knowledge graph, an adjacency matrix is established to provide input for the subsequent internal threat detection model.
[0082] The internal threat knowledge graph can represent information such as user entities, operation behaviors, and user social relationships. Constructing its data can provide more information support for subsequent internal threat detection and achieve regulatory visualization. As Figure 2 shown, the architecture of this method can be divided into four modules: ontology modeling, data collection and preprocessing, knowledge graph construction and storage, and knowledge graph application.
[0083] (1) Ontology modeling module 1. The main task of this module is to model the internal threat domain ontology. By analyzing the construction purpose of this knowledge base and the characteristics of the data source, the ontology structure is designed from the dimension of internal users as the knowledge organization mode of the knowledge graph. By constructing the internal threat behavior knowledge graph, the characteristics of internal users are analyzed and threat detection is performed. The internal threat ontology structure diagram is as Figure 3 shown.
[0084] (2) Data collection and preprocessing module 2. This module first needs to obtain and collect data, and then perform data cleaning and filling of missing values on these data to ensure data quality and consistency.
[0085] (3) Knowledge Graph Construction and Storage Module 3. This module first extracts entities, relationships, and attributes from the data source using a rule-based method to create data instances of the ontology, and then uses the Py2neo toolkit to store the data of the knowledge graph in the Neo4j graph database, obtaining the constructed internal threat knowledge graph.
[0086] (4) Knowledge Graph Application Module 4. After the construction of the knowledge graph is completed, operations such as querying and exporting can be performed using the visualization interface of the Neo4j database. The Export command in the APOC library filters the data in the required graph and exports it as a csv file as the input for the subsequent internal threat detection model.
[0087] Compared with the prior art, the beneficial effects of the present invention are:
[0088] 1. The present invention uses a method for constructing an internal threat knowledge graph. First, construct the ontology of the internal threat knowledge graph, then extract important user features from the original log data file, and aggregate data from different sources according to the user ID under given aggregation conditions. Subsequently, perform feature extraction on the aggregated data to generate numerical vectors and extract user entities and the relationships between user entities, and then integrate them into the knowledge graph to finally construct the internal threat knowledge graph. By extracting the feature matrix of users and the weighted relationships between all user entities from the knowledge graph to establish an adjacency matrix, it provides input for the subsequent internal threat detection model;
[0089] 2. After inputting the feature matrix and adjacency matrix extracted from the knowledge graph into the model, first train the feature matrix and adjacency matrix through three graph convolutional network layers; then, introduce a residual unit before the first graph convolutional layer, take the input of the first graph convolutional layer as the initial residual, and jump-connect the initial residual to the third graph convolutional layer, effectively preventing the problem of gradient disappearance and improving the aggregation efficiency; at the same time, the present invention also adds three layers of MLP before the initial residual jump connection. Such a design allows the network to perform more flexible transformations on the input features, helping to adapt to different data patterns. The present invention constructs an internal threat knowledge graph based on the communication relationships and behavior characteristics of users, realizes the visual supervision of internal users by the organization, and accurately identifies threat users based on this.
[0090] The above-disclosed is only a preferred embodiment of an internal threat detection method based on a knowledge graph and a residual graph convolutional network of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. An internal threat detection method based on knowledge graph and residual graph convolutional network, characterized in that: The following steps are involved: Collect user feature information and pre-process it to obtain internal user data; Building an internal threat knowledge graph ontology based on the internal user data; Based on the internal threat knowledge graph ontology, an internal threat knowledge graph is constructed using a graph database; Constructing a feature matrix based on the internal threat knowledge graph; A weighted function is constructed based on the adjacency matrix weighted by the user communication relationship and the user behavior similarity to build the adjacency matrix; The adjacency matrix and the feature matrix are trained based on the internal threat detection method of the residual graph convolutional network to obtain a detection result; The user characteristic information includes resume information, personality information, communication records and behavior records; The constructing of the internal threat knowledge graph ontology based on the internal user data includes: Converting discrete data in the internal user data into a one-hot encoded vector; Constructing an internal threat knowledge graph ontology based on the one-hot encoded vector; The residual graph convolutional network consists of three layers of graph convolutional network layers and residual modules, and three layers of multi-layer perceptron layers are added before the initial residual jump connection to the third layer of the graph convolutional network layer.
2. The internal threat detection method based on knowledge graph and residual graph convolutional network as claimed in claim 1, characterized in that: The preprocessing includes data cleaning and filling of missing values.
3. The internal threat detection method based on knowledge graph and residual graph convolutional network as claimed in claim 1, characterized in that: The graph database includes a Neo4j graph database.
4. An internal threat detection system based on knowledge graph and residual graph convolutional network, applied to the internal threat detection method based on knowledge graph and residual graph convolutional network according to claim 3, characterized in that: It includes ontology modeling module, data collection and preprocessing module, knowledge graph construction and storage module, and knowledge graph application module; The ontology modeling module is used to design an ontology structure from the dimension of internal users by analyzing the construction purpose of the knowledge base and the characteristics of the data source, as a knowledge organization model of the knowledge graph, to obtain the internal threat knowledge graph ontology; The data collection and preprocessing module is used to obtain and collect user characteristic information, and then perform data cleaning and missing value filling on the user characteristic information to obtain internal user data; The knowledge graph construction and storage module extracts entities, relationships and attributes from the internal user data based on a rule-based method, creates a data instance of the internal threat knowledge graph ontology, and then uses the Py2neo toolkit to store the knowledge graph data in a Neo4j graph database to obtain a constructed internal threat knowledge graph; The knowledge graph application module uses the visual interface of the Neo4j database to perform query and export operations. The Export command in the APOC library filters the required data in the graph and exports it as a csv file as input for the subsequent internal threat detection model.
Citation Information
Patent Citations
Internal threat anomaly detection method based on double-domain graph convolutional neural network
CN116484363A
Internal user attribute portrait construction method
CN117033650A