Cyber threat prediction system and method thereof

WO2025010053A3PCT designated stage Publication Date: 2025-06-26BTS KURUMSAL BİLİŞİM TEKNOLOJİLERİ ANONİM ŞİRKETİ
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/TR2024/051283
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current Intrusion Detection Systems (IDS) lack an advanced, explainable Graph Neural Networks (GNN)-based solution capable of making accurate cyber threat predictions, particularly for sophisticated attacks like BruteForce, DDoS, DoS, and Bot, due to inadequate handling of graph-based relationships and non-transparent decision-making processes, which affects trust and accountability.

Method used

A cyber threat prediction system utilizing Graph Neural Networks (GNNs) with edge embedding representations generated through Self-Supervised Learning, combined with Explainable Artificial Intelligence (XAI) techniques like PGExplainer and CatBoost, to provide high-accuracy predictions and explanations by processing network traffic data as graphs, incorporating both local and global neighborhood information.

Benefits of technology

The system achieves higher accuracy in intrusion detection and provides transparent explanations, effectively detecting sophisticated cyber threats and adapting to dynamic network topologies, enhancing trust and accountability in IDS systems.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to an explainable and graph neural networks-based cyber threat prediction system for IDS environments, which utilizes the structural advantages of Graph Neural Networks (GNNs) to efficiently process network traffic data, and adapts a novel Explainable Artificial Intelligence (XAI) methodology, and to the method thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CYBER THREAT PREDICTION SYSTEM AND METHOD THEREOF

[0002] Technical Field of the Invention

[0003] The invention relates to an explainable and graph neural networks-based cyber threat prediction system for Intrusion Detection System (IDS) environments, which utilizes the structural advantages of Graph Neural Networks (GNN) to efficiently process network traffic data, and adapts a novel Explainable Artificial Intelligence (XAI) methodology, and to the method thereof.

[0004] State of the Art

[0005] The current IDSs literature mainly prioritizes (i) developing advanced Machine Learning (ML) / Deep Learning (DL) models for complex intrusion detection and (ii) using XAI to declassify ML and DL decision-making process. From the first perspective, recent studies show that GNNs are particularly promising for IDS. This is because the natural structure of computer networks is a graph where nodes are network devices (e.g. routers, hosts, etc.) and edges are connections between network devices, packet transfers, and network flows. GNNs reveal the impact of malicious cyber activities on the topology of the network and make use of neighborhood information between network entities. Moreover, existing studies show that it is more appropriate to use network flows in IDS studies rather than individual packet-based monitoring, as combining network features helps to reveal the diverse and heterogeneous characteristics of cyber intrusions. From a second perspective, there is a concerted academic effort to incorporate XAI models into various ML / DL frameworks to provide local and global explanations of their operations. This effort aims to explore the decision-making mechanism behind model predictions, whether by clarifying the importance of specific data points (through local explainability techniques) or shedding light on the overall behavior of the model (through global explainability techniques).

[0006] In the state of the art, taking into account the research efforts on IDS and the security needs of organizations, there is no advanced GNN-based IDS solution methodology capable of making accurate prediction and used in conjunction with a complex XAI framework. Current explanation methods using XAI frameworks use either local or global explainability approaches. This results in poor performance in the IDS field, where both local and global explainability techniques are important at the same time. Existing explanation methods using XAI frameworks are not suitable for explaining graph-based relationships as they ignore relationships between nodes, reducing explainability performance. In addition, the existing prediction and explainability methods developed for IDSs cannot adapt to the dynamic and variable nature of Internet Service Providers (ISPs) that provide Internet services to many customers in various network topologies.

[0007] In the state of the art, the increase in complex cyber threats in the digital and information environment causes the existing security systems to be insufficient. The efficiency of IDSs is of key importance in an era of increasingly sophisticated cyber threats. ML and DL models offer efficient and accurate solutions for detecting intrusions and anomalies in computer networks. However, the use of ML and DL models in IDS leads to a security gap due to non-transparent decision-making processes. This transparency gap in IDS research is a major aspect that affects trust and accountability. Therefore, it is crucial to develop a new explainable and graph neural network-based cyber threat prediction method for IDS environments that solves the aforementioned problems, utilizes the structural advantages of GNNs to efficiently process network traffic data, and also adapts a new XAI methodology.

[0008] Summary and Objects of the Invention

[0009] The invention relates to a new explainable and graph neural networks-based cyber threat prediction system designed for IDS environments that utilizes the structural advantages of GNNs to efficiently process network traffic data, and also adapts a new XAI methodology, and to a method thereof.

[0010] An object of the invention is to enable the generation of edge embedding representations through graph theory with Self-Supervised Learning.

[0011] Another object of the invention is to design explainable artificial intelligence technique for IDS systems and to provide high-accuracy cyber threat prediction. Another object of the invention is to successfully detect sophisticated cyber intrusions, especially the four attacks named BruteForce, DDoS, DoS, and Bot in the current literature.

[0012] Another object of the invention is to use graph structure and edge embedding representations in IDS cyber intrusion predictions with the new ML training model it proposes, and thus, to provide better prediction results by using edge embedding representations within the graph structure instead of directly predicting from the data itself as in existing studies.

[0013] Description of the Drawings

[0014] Figure 1. The drawing showing the flow diagrams of the method of the invention.

[0015] Figure 2. The drawing showing the schematic view of the system of the invention.

[0016] Figure 3. The drawing showing the schematic view of the system of the invention.

[0017] Description of the References in the Drawings

[0018] 1 . Cyber threat prediction system

[0019] 2. Network device

[0020] 3. Router monitoring device

[0021] 10. Pre-modeling module

[0022] 11 . Module for collecting data from network sensors

[0023] 12. Data pre-processing module

[0024] 13. Time-dependent IDS data flow generation module

[0025] 14. Graph modeling module specific to IDS networks 20. Modeling module

[0026] 21. Edge embedding representations generation module

[0027] 211 . Trained Deep Graph Infomax (DGI) model

[0028] 22. Learning encoder

[0029] 221 . Corruption Function

[0030] 222. Encoder Function

[0031] 223. Corrupted Embedded Graph

[0032] 224. Embedded Graph

[0033] 23. Learning Decoder

[0034] 231 . Discriminator Function

[0035] 232. Read Function

[0036] 30. Advanced-Modeling Module

[0037] 31. IDS Intrusion Detection Prediction Module

[0038] 311 . Trained E-GraphSAGE Model

[0039] 312. Trained CatBoost Model

[0040] 32. IDS Intrusion Detection Prediction Explanation Module

[0041] 321 . PGExplainer Model 322. Critical Subgraph

[0042] 323. Non-Critical Subgraph

[0043] 1001. Initiating operation of the module (11 ) for collecting data from network sensors

[0044] 1002. Installing network traffic monitoring tools, updating firmware, and adjusting firewall and Intrusion Prevention System (IPS) / IDS systems for continuous monitoring of network devices in the network topology

[0045] 1003. Periodically collecting network data on traffic volume, connection times, and packet rates associated with intrusions on the network that are confirmed and deemed relevant by experts associated with intrusions on the network

[0046] 1004. Recording data in a real-time database by the module (11 ) for collecting data from network sensors

[0047] 2001. Initiating the operation of the data pre-processing module (12)

[0048] 2002. Removing single-valued attributes from the data by the data pre-processing module (12), which do not contain any information about time-varying network conditions for network intrusion detection

[0049] 2003. Detecting and removing instances by the data pre-processing module (12), where numeric attributes contain alphabetic characters due to errors during data recording

[0050] 2004. Removing non-numeric, categorical attributes from IDS data by the data preprocessing module (12) using “one-hot encoding” and “label encoding” algorithms

[0051] 2005. Detecting non-existing values in the dataset by the data pre-processing module (12) and filling them using backfill (bfill) and / or forwardfill (ffill) methods

[0052] 2006. Generating a correlation matrix by the data pre-processing module (12) to understand the relationship between attributes, and keeping only one of the highly correlated attributes in the dataset and removing the other one to reduce computation time and increase the generalization capability of the models

[0053] 2007. Determining the acceptable limits for attribute-based sparsity ratios by domain experts, keeping the attributes within these limits in the data, and eliminating the others by the data pre-processing module (12)

[0054] 2008. Completing the data pre-processing module

[0055] 3001. Initiating the operation of the time-dependent IDS data flow generation module

[0056] (13)

[0057] 3002. Converting instantaneous network packets, collected from network devices and pre-processed, into network flows in a cumulative manner over a defined period of time by the time-dependent IDS data flow generation module (13)

[0058] 3003. Completing the Time-dependent IDS Data Flow Generation Module

[0059] 3004. Initiating the operation of the graph modeling (14) module specific to IDS networks

[0060] 3005. Converting network devices to represent nodes and network flows to represent edges into multi-graph data over the generated network flows by the graph modeling

[0061] (14) module specific to IDS networks

[0062] 3006. Completing the graph modeling module specific to IDS networks

[0063] 4001. Initiating the operation of the edge embedding representations generation module (21)

[0064] 4002. Obtaining pre-labels by the edge embedding representations generation module (21 ) through self-learning by running the DGI algorithm

[0065] 4003. Initiating the operation of the Learning Encoder (22) 4004. Using the trained DGI edge embedding representations as input, obtaining edge embedding representations for the graph by the Learning Encoder (22) through the graph neural network-based Encoder element E-GraphSAGE algorithm

[0066] 4005. Designing the encoder layer as a 1 -layer E-GraphSAGE with the DGI algorithm by Learning Encoder (22)

[0067] 4006. Completing the Learning Encoder (22) and initiating the operation of the learning decoder (23)

[0068] 4007. Determining a corruption function by the learning decoder (23) to generate a negative (corrupted) graph representation from the input graph

[0069] 4008. Generating edge embedding representations from both the input graph and the intentionally corrupted graph by the learning decoder (23) with the encoder layer of the DGI

[0070] 4009. Averaging and aggregating the edge embedding representations by the learning decoder (23) with the read function at the center of the DGI, and then processing these through a sigmoid function to compute a global graph summary

[0071] 4010. Evaluating the true and corrupted edge embedding representations of the discriminator layer of the DGI by the learning decoder (23), using its global summary as a guide

[0072] 4011. Providing a score between 0 and 1 by the learning decoder (23) with the help of the binary cross-entropy loss objective function, based on the comparisons made by the discriminator layer, and distinguishing the edge embedding representations of the true edge and the corrupted edge to train the encoder

[0073] 4012. Obtaining the trained E-GraphSAGE model trained, capable of correctly making graph edge embedding representations, by the learning decoder (23) and recording the model 4013. Completing the Learning Decoder (23) and completing the edge embedding representations generation module

[0074] 5001. Initiating the operation of the IDS intrusion detection prediction module (31 )

[0075] 5002. Investigating the values for the basic variables of the method with hyperparameter optimization to find the CatBoost method specific to the IDS structure by the IDS intrusion detection prediction module (31 ), and reaching the most accurate prediction model

[0076] 5003. Comparing the models obtained with the researched values through the performance evaluation variables by the IDS intrusion detection prediction module (31), and deciding on the CatBoost model specific to the IDS structure

[0077] 5004. Recording a CatBoost intrusion detection model designed specifically for the problem by the IDS intrusion detection prediction module (31)

[0078] 5005. Detecting new incoming traffic flows as intrusion or normal by the IDS intrusion detection prediction module (31) through the CatBoost model designed specifically for the problem

[0079] 5006. Generating an alert log in the system of network traffic flows detected as intrusions by the IDS intrusion detection prediction module (31)

[0080] 5007. Generating an information log in the system of network traffic flows detected as normal by the IDS intrusion detection prediction module (31 )

[0081] 5008. Completing the IDS intrusion detection prediction module (31)

[0082] 6001. Initiating the operation of the IDS intrusion detection prediction explanation module (32)

[0083] 6002. Investigating the most accurate values for the basic variables of the method with hyper-parameter optimization to find the most accurate PGExplainer method by the IDS intrusion detection prediction explanation module (32) 6003. Deciding on the PGExplainer model by the IDS intrusion detection prediction explanation module (32) based on the performance metrics of fidelity and sparsity

[0084] 6004. Generating global explanation with PGExplainer for intrusion detection system by IDS intrusion detection prediction explanation module (32)

[0085] 6005. Generating local explanation with PGExplainer for instances detected as intrusions by the IDS intrusion detection prediction explanation module (32)

[0086] 6006. Completing the IDS intrusion detection prediction explanation module (32)

[0087] Detailed Description of the Invention

[0088] The invention relates to a new explainable and graph neural networks-based cyber threat prediction system designed for IDS environments that utilizes the structural advantages of GNNs to efficiently process network traffic data, and also adapts a new XAI methodology, and to a method thereof.

[0089] The invention consists of three modules: Pre-modeling, Modeling, and Advanced- modeling.

[0090] The cyber detection system proposed in the invention uses network flows along with many critical network metrics. Pre-Modeling gathers information from network flows and then reviews the data using a data cleansing approach. It deletes empty and incomplete data and also makes the raw data meaningful thanks to data preprocessing. Then, it processes the collected data in accordance with the time series pattern. After these processes, the network data is modeled in the form of a graph by Pre-Modeling using graph theory. During this modeling, graph edge embedding representations are generated between each node. Here, nodes represent network devices, such as routers, in a network topology. The edges, on the other hand, represent the network flows between the network routers. Pre-Modeling processes these edges as embedding s and generates the graph model for use in the next module, Modeling module. Pre-Modeling also introduces distinctive threat models by modeling flow-based network data in the form of graphs and using graph edges. In this way sophisticated cyber intrusions, especially the four attacks named BruteForce, DDoS, DoS, and Bot in the current literature are successfully detected.

[0091] In the modeling stage, which is another module of this invention, embedding representations of graph edges are obtained. For this, two algorithms available in the literature are used. One of these is an algorithm named E-GraphSAGE, and the other is an algorithm named DGI. This invention generates a graph input containing nodes and edges necessary for the operation of the E-GraphSAGE algorithm, thanks to the previous module, Pre-Modeling. The invention then uses the edge features between each node together with the node features to propagate the message. This invention utilizes the E-GraphSAGE algorithm to generate new vector representations for each node and each edge. The network data flow-based IDS datasets of this invention comprise only edge features. For this reason, the E-GraphSAGE algorithm has been modified so that the node feature vectors consist of constant values. This invention samples a fixed number of neighbors for each node in the graph data and aggregates information from the sampled neighbors to create embedding representations for the target node and edge. Then, these embedding representations are refined using the DGI algorithm. The DGI algorithm is an approach that enables self supervised learning from unlabeled data. This model aims to learn the underlying features of the data through pre-labels that are automatically derived from the data. These pre-labels are generated by the model and are not the actual labels of the data. Teaching these labels contributes to the generation of good representations for the data and to increase the predictive performance of the supervised learning model to be used later. This invention improves the E-GraphSAGE algorithm by using it within the DGI algorithm. Therefore, it provides self-learning by maximizing the mutual knowledge between local and global information.

[0092] The trained encoder and edge embedding representations obtained through the modeling phase of the invention are used as input in the advanced-modeling module. The invention then uses the CatBoost classification algorithm described in the literature in the advanced-modeling module. The CatBoost classification algorithm is a gradient boosting machine learning algorithm known for its high prediction accuracy and speed. Using techniques such as gradient boosting and oblivious trees, it effectively processes a variety of data types and reduces overfitting. These methods help the model produce more general and reliable results, while enabling it to work quickly on large and complex datasets. This invention makes a binary decision for IDS estimation by utilizing the CatBoost algorithm in the decision phase of this module. This decision, if yes, means that there was an IDS cyber attack. If no, it means that there is no IDS cyber attack. Once the prediction has been made, this invention uses an approach named PGExplainer for explainability. PGExplainer provides explanations at a global level in many instances by developing a common network of explanations from the node and graph representations of the GNN model. It identifies a critical subgraph that contains the most important nodes in the graph to explain the decisions made by the trained model during prediction. In this process, nodes and attributes are removed and the effects of these operations on the output of the GNN model are analyzed. Removing nodes also requires removing the edges at the endpoints of that node from the graph; this makes it easier to identify important edges. In this invention, using the PGExplainer, the graph is divided into two parts. One of these is another subgraph that represents the critical subgraph and does not comprise the edges that are redundant. The other graph is the subgraph with important edges. In this way, Advanced-Modeling explains the graph topologies in GNNs and determines the critical subgraph in a way that minimizes conditional entropy. In summary, owing to the three modules presented by this invention, the invention makes predictions with edge embedding representations using both local and global neighborhood information. With the new machine learning training model proposed by this invention, graph structure and edge embedding representations are used in IDS cyber attack predictions. In this way, better prediction results are obtained by using edge embedding representations within the graph structure, rather than directly predicting from the data itself as in current studies. By integrating a GNN-based detection process line with the CatBoost classifier, this invention achieves higher accuracy in intrusion detection compared to existing state-of- the-art solutions. The invention uses a GNN-based model in IDS cyber attack prediction systems and provides both local and global insights by adapting the PGExplainer algorithm during the explainability phase. The use of PGExplainer outperforms basic explainability models, especially since it can operate in inductive environments, offers a non-black box evaluation approach, and can explain multiple instances collectively. This makes the invention highly suitable for IDS network data. In addition, thanks to the self-learning edge embedding representations used by the invention, the invention is compatible with changing topologies. Cyber threat prediction system (1) is used in intrusion detection systems. For internet service providers, IDS provides cyber threat prediction and prediction explanation. The cyber threat prediction system (1) consists of a pre-modeling module (10), a modeling module (20), and an advanced-modeling module (30).

[0093] Network devices (2) ensure communication on the network and the forwarding of network packets to the correct delivery point.

[0094] The router monitoring device (3) communicates with the network device (2) and enables the monitoring of the devices in the network.

[0095] It consists of a pre-modeling module (10), a module (11 ) for collecting data from network sensors, a data pre-processing module (12), a time-dependent IDS data flow generation module (13), and a graph modeling module (14) specific to IDS networks.

[0096] The module (11 ) for collecting data from network sensors enables the observation and collection of network parameters without tiring the system with data flow traffic and without affecting its performance. The main criteria for detecting a cyber intrusion is to comprehensively identify the relevant dataset at the beginning and to subject the system to continuous monitoring.

[0097] The data pre-processing module (12) deletes the empty and incomplete data in the data obtained from the network sensors through the module (11) for collecting data, makes the raw data meaningful, and then processes the collected data in accordance with the time series pattern. The data pre-processing module (12) includes attribute engineering studies specific to the problem of cyber intrusion detection in communication networks. For the artificial intelligence-based IDS algorithm, it is responsible for preparing the data in accordance with the artificial intelligence algorithm proposed by the invention. Without this module, even any learning algorithm may not be able to produce the expected performances.

[0098] The data pre-processing module (12) implements six steps of data pre-processing and cleaning. In the first step, it removes single-valued attributes from the data which do not contain any information about time-varying network conditions for network intrusion detection. In the second step, if the numeric attributes contain alphabetic characters due to some errors that occur during the recording of the data, it detects these examples and removes them from the dataset. In the third step, non-numeric attributes, i.e. categorical attributes, are digitized with "One-Hot-Encoding" and "Label Encoding" algorithms. In the fourth step, it detects the values that do not exist in the dataset and fills in the values that do not exist by using the bfill and / or ffill methods. This step is crucial for this invention as the IDS dataset used by this invention is a time series data, and direct deletion of this data causes data loss. In the fifth step, it generates the correlation matrix to understand the relationship between the attributes. This step is crucial for this invention as it keeps only one of the highly correlated attributes in the dataset and removes the others in order to reduce computation time and increase the generalization capability of the models. In the sixth and final step, the attribute sparsity ratio is calculated. The attribute sparsity ratio is the number of zeros in an attribute column divided by the number of instances that the attribute has. It sets acceptable limits for the attribute-based sparsity ratio and ensures that the attributes within these limits are kept in the data and the others are eliminated.

[0099] The time-dependent IDS data flow generation module (13) generates network flow data from cumulative network packets over a period of time using instantaneous network packets. In this way, long-term IDS activities are prevented from corrupting the entire data due to instantaneous anomalies. The Time-Dependent IDS Data Flow Generation Module (13) prepares the IDS prediction data for the learning algorithm in the next modules.

[0100] Graph Modeling Module (14) Specific to IDS Networks transforms data into graph modeling. In other words, it converts network data into multi-graph data, with routers representing nodes and data flows representing edges.

[0101] The modeling module (20) is the element where the ML model for IDS prediction is trained. Modeling module (20) consists of three modules. Modeling module (20) consists of edge embedding representations generation module (21 ), learning encoder (22), and the learning decoder (23) modules.

[0102] The edge embedding representations generation module (21) provides self supervised learning from unlabeled data, based on the DGI algorithm, and includes a problemspecific trained deep graph infomax (DGI) model (211). With this module, the underlying features of the data are learned through pre-labels that are automatically derived from the data. These pre-labels are generated by a trained deep graph infomax model (211) and are not actual labels of the data. Teaching these labels contributes to the generation of good representations for the data and to increase the predictive performance of the supervised learning model to be used later.

[0103] The Learning Encoder (22) uses the edge embedding representations of the trained deep graph infomax (DGI) model (211 ) as input and obtains the edge embedding representations for the graph using the E-GraphSAGE algorithm. The self-learning DGI model (211) trained with the Edge Embedding Representations Generation Module (21 ) is adjusted with an E-GraphSAGE encoder (222) in the Learning Encoder (22). The trained deep graph infomax (DGI) model (211) encodes in a 1 -layer E-GraphSAGE model using an average aggregate function. E-GraphSAGE uses a hidden layer size of 256 units and Rectified Linear Unit (ReLU) is used as the activation function. The Learning Encoder (22) determines a Corruption Function (221) (i.e., of the input graph that adds or subtracts nodes from the adjacency matrix) to generate a negative (corrupted) graph representation from the input graph. It then generates edge embedding representations from both the Embedded Graph (224) (i.e. preserved edge representations) and the Corrupted Embedded Graph (223) with the encoder layer of the trained deep graph infomax (DGI) model (211 ). To construct the spherical graph summary, the edge embedding representations are averaged and passed through a sigmoid function. Binary Cross-Entropy (BCE) was used as the loss function and Adam Optimizer with a learning ratio of 0.001 and gradient descent were used for backpropagation. A trained DGI model (211) provides edge embedding representations in the graph with the encoder. These embedding representations are obtained by collecting information from the edges of neighboring nodes. Therefore, it maximizes the mutual knowledge between local and global information. With this stage, the Learning Encoder (22) study is completed.

[0104] In the Learning Decoder (23), Adam Optimizer and BCE Loss approaches are used. The gradient descent optimization supported by these two approaches uses the Discriminator Function (231) of the DGI algorithm to optimize its values in an iterative manner. With the Read Function (232) at the center of the DGI algorithm, it averages and aggregates the edge embedding representations, and then processes them through a sigmoid function to compute a spherical graph summary (a single vector embedding representation of the entire graph). The DGI algorithm then evaluates the true and corrupted edge embedding representations of the discriminator layer by using its global summary as a guide. Comparisons made by the Discriminator Function (231 ) are given a score between 0 and 1 by means of the BCE Loss function. The Learning Decoder (23) sub-module therefore distinguishes between the embedding representations of the Embedded Graph (224) edge and the Corrupted Graph (223) edge in order to train the encoder. The Learning Decoder (23) completes its operation and the Trained E-GraphSAGE Model (311) is obtained, which can accurately make graph edge embedding representations. This model (311 ) is recorded for later use. The Modeling Module (20) in this invention improves the E-GraphSAGE algorithm using the DGI algorithm. Therefore, it provides self-learning by maximizing the mutual knowledge between local and global information. Finally, the last module of the cyber threat prediction system (1 ), the advanced-modeling module (30), operates. The trained deep graph infomax (DGI) model (211) encoder and edge embedding representations obtained from the modeling (20) phase are used as input in this module (30).

[0105] The Advanced-Modeling (30) module runs the first phase, the IDS intrusion detection prediction module (31). This module (31 ) uses the CatBoost classification algorithm, which has been described in the literature. The CatBoost Classification algorithm is a gradient boosting machine learning algorithm known for its high prediction accuracy and speed. The IDS Intrusion Detection Prediction Module (31 ) detects values for the method's primitive variables (e.g. min samples leaf, min samples split, and max depth) to find the CatBoost method specific to the IDS structure. It compares the obtained CatBoost Classification models with the detected values over performance evaluation variables (F1 -Macro, Accuracy, Detection Rate). The IDS Intrusion Detection Prediction Module (31 ) decides on the best performing CatBoost Classification model. The IDS intrusion detection prediction module (31 ) records the Trained CatBoost model (312) intrusion detection model designed specifically for the problem. Then, using this recorded model (312), it identifies new incoming traffic flows collected from the network flow traffic as intrusion or normal. This invention makes a binary decision for IDS estimation by utilizing the CatBoost algorithm in the decision phase of this module. Since this decision made with the trained CatBoost Model (312) is made using edge embedding representations, it outperforms the meta-features of the data, i.e., raw data. As a result this decision, if yes, means that there was an IDS cyber attack. The IDS intrusion detection prediction module (31 ) generates an alert log in the system of network traffic flows detected as intrusions. If the decision is no, it means that there is no IDS cyber attack. The IDS intrusion detection prediction module (31 ) generates an information log in the system of network traffic flows detected as normal. Then, the Advanced-Modeling (30) module runs the last phase, the IDS Intrusion Detection Prediction Explanation Module (32). Once the prediction has been made, this module (32) uses an approach named PGExplainer for explainability. PGExplainer provides explanations at a global level in many instances by developing a common network of explanations from the node and graph representations of the graph neural network (GNN) model. The IDS Intrusion Detection Prediction Explanation Module (32) identifies a critical subgraph that contains the most important nodes in the graph to explain the decisions made by the trained model during prediction. In this process, the IDS Intrusion Detection Prediction Explanation Module (32) removes nodes and attributes and analyzes the effects of these operations on the output of the GNN model. Removing nodes also requires removing the edges at the endpoints of that node from the graph; this enables identifying important edges. In this invention, using the PGExplainer, the graph is divided into two parts. One of these is the Critical Subgraph (322), which represents the critical subgraph, and the Non-Critical Subgraph (323), which comprises the edges that are redundant. In this way, the edges (data flows) which contribute the most to the prediction and which are critical in the network topology are identified. PGExplainer determines the Critical Subgraph (322) by maximizing mutual knowledge using the entropy term. The purpose of the PGExplainer Model (321) approach is to specifically describe the graph topologies in GNNs and to define the Critical Subgraph (322) in a way that minimizes conditional entropy. However, because there are too many candidate values for the Critical Subgraph (322), direct optimization is not possible. Therefore, a relaxation approach is used, assuming that the Critical Subgraph (322) follows the random graph distribution by the Gilbert approach and that the edge selections from the original input graph are conditionally independent of each other. With this relaxation, the PGExplainer Model (321) is reshaped and its optimization becomes more manageable. In this way, Advanced-Modeling Module (30) explains the graph topologies in GNNs and determines the critical subgraph in a way that minimizes conditional entropy. The use of PGExplainer Model (321) outperforms basic explainability models, especially since it can operate in inductive environments, offers a non-black box evaluation approach, and can explain multiple instances collectively. This makes the invention highly suitable for IDS network data. In addition, thanks to the self-learning edge embedding representations used by the invention, the invention is compatible with changing topologies.

[0106] The cyber threat prediction method includes the following process steps;

[0107] Initiating operation of the module (11 ) for collecting data from network sensors (1001 ),

[0108] Installing network traffic monitoring tools, updating firmware, and adjusting firewall and IPS / IDS systems for continuous monitoring of network devices in the network topology (1002)

[0109] Periodically collecting network data on traffic volume, connection times, and packet rates associated with intrusions on the network that are confirmed and deemed relevant by experts associated with intrusions on the network (1003)

[0110] Recording data in a real-time database by the module (11 ) for collecting data from network sensors (1004)

[0111] Initiating the operation of the data pre-processing module (12) (2001 )

[0112] Removing single-valued attributes from the data by the data pre-processing module (12), which do not contain any information about time-varying network conditions for network intrusion detection (2002)

[0113] Detecting and removing from the data the instances where numeric attributes contain alphabetic characters due to errors during data recording by the data pre-processing module (12) (2003)

[0114] Removing non-numeric, categorical attributes from IDS data by the data preprocessing module (12) using “One-Hot Encoding” and “Label Encoding” algorithms (2004),

[0115] Detecting non-existing values in the dataset by the data pre-processing module (12) and filling them using bfill and / or ffill methods (2005), Generating a correlation matrix by the data pre-processing module (12) to understand the relationship between attributes, and keeping only one of the highly correlated attributes in the dataset and removing the other one to reduce computation time and increase the generalization capability of the models (2006),

[0116] In the process step of generating a correlation matrix by the data pre-processing module (12) to understand the relationship between attributes, and keeping only one of the highly correlated attributes in the dataset and removing the other one to reduce computation time and increase the generalization capability of the models (2006), provided that the correlation ratio is 0.95 or higher, only one of these attributes is kept in the dataset and the other is removed.

[0117] Determining the acceptable limits for attribute-based sparsity ratios by domain experts, keeping the attributes within these limits in the data, and eliminating the others by the data pre-processing module (12) (2007),

[0118] In the process step of determining the acceptable limits for attribute-based sparsity ratios by domain experts, keeping the attributes within these limits in the data, and eliminating the others by the data pre-processing module (12) (2007), while it is desired that this sparsity ratio determined by the domain experts is 0 (zero), as a result of the observations made by the domain experts, it was found acceptable for this value to be less than 0.15, and the attributes that are within these limits accepted by the experts are kept in the data and the others are eliminated.

[0119] Completing the data pre-processing module (2008),

[0120] Initiating the operation of the time-dependent IDS data flow generation module (13) (3001),

[0121] Converting instantaneous network packets, collected from network devices and pre- processed, into network flows in a cumulative manner over a defined period of time by the time-dependent IDS data flow generation module (13) (3002),

[0122] Completing the Time-dependent IDS Data Flow Generation Module (3003), Initiating the operation of the graph modeling (14) module specific to IDS networks (3004),

[0123] Converting network devices (network routers) to represent nodes and network flows to represent edges into multi-graph data over the generated network flows by the graph modeling (14) module specific to IDS networks (3005),

[0124] Completing the Graph Modeling Module Specific to IDS Networks (3006),

[0125] Initiating the operation of the edge embedding representations generation module (21 ) (4001),

[0126] Obtaining pre-labels by the edge embedding representations generation module (21) through self-learning by running the deep graph infomax (DGI) algorithm (4002),

[0127] Initiating the operation of the Learning Encoder (22) (4003),

[0128] Using the trained DGI edge embedding representations as input, obtaining edge embedding representations for the graph by the Learning Encoder (22) through the GNN-based Encoder element E-GraphSAGE algorithm (4004),

[0129] Designing the encoder layer as a 1 -layer E-GraphSAGE with the DGI algorithm by Learning Encoder (22) (4005),

[0130] Completing the Learning Encoder (22) and initiating the operation of the learning decoder (23) (4006),

[0131] Determining a corruption function (i.e., of the input graph that adds or subtracts nodes from the adjacency matrix) by the learning encoder (23) to generate a negative (corrupted) graph representation from the input graph (4007),

[0132] Generating edge embedding representations from both the input graph (i.e. preserved edge representations) and the intentionally corrupted graph by the learning decoder (23) with the encoder layer of the DGI (4008), Averaging and aggregating the edge embedding representations by the learning decoder (23) with the read function at the center of the DGI, and then processing these through a sigmoid function to compute a global graph summary (a single vector embedding representation of the whole graph) (4009),

[0133] Evaluating the true and corrupted edge embedding representations of the discriminator layer of the DGI by the learning decoder (23), using its global summary as a guide (4010),

[0134] Providing a score between 0 and 1 by the learning decoder (23) with the help of the BCE loss objective function, based on the comparisons made by the discriminator layer, and distinguishing the edge embedding representations of the true edge and the corrupted edge to train the encoder (4011),

[0135] Obtaining the trained E-GraphSAGE model trained, capable of correctly making graph edge embedding representations, by the learning decoder (23) and recording the model (4012),

[0136] Completing the Learning Decoder (23) and completing the edge embedding representations generation module (4013),

[0137] Initiating the operation of the IDS intrusion detection prediction module (31) (5001),

[0138] Investigating the values for the basic variables (e.g. min samples leaf, min samples split, and max depth) of the method with hyper-parameter optimization to find the CatBoost method specific to the IDS structure by the IDS intrusion detection prediction module (31), and reaching the most accurate prediction model (5002),

[0139] Comparing the models obtained with the researched values through the performance evaluation variables (F1 -Macro, Accuracy, Detection Rate) by the IDS intrusion detection prediction module (31 ), and deciding on the CatBoost model specific to the IDS structure (5003),

[0140] Recording a CatBoost intrusion detection model designed specifically for the problem by the IDS intrusion detection prediction module (31) (5004), Detecting new incoming traffic flows as intrusion or normal by the IDS intrusion detection prediction module (31 ) through the CatBoost model designed specifically for the problem (5005),

[0141] Generating an alert log in the system of the network traffic flows detected as intrusions by the IDS intrusion detection prediction module (31 ) (5006),

[0142] Generating an information log in the system of the network traffic flows detected as normal by the IDS intrusion detection prediction module (31 ) (5007),

[0143] Completing the IDS intrusion detection prediction module (31 ) (5008),

[0144] Initiating the operation of the IDS intrusion detection prediction explanation module (32) (6001 ),

[0145] Investigating the most accurate values for the basic variables of the method with hyperparameter optimization to find the most accurate PGExplainer method by the IDS intrusion detection prediction explanation module (32) (6002),

[0146] Deciding on the PGExplainer model by the IDS intrusion detection prediction explanation module (32) based on the performance metrics of fidelity and sparsity (6003),

[0147] Generating global explanation with the PGExplainer for the intrusion detection system by the IDS intrusion detection prediction explanation module (32) (6004),

[0148] Generating local explanation with the PGExplainer for the instances detected as intrusions by the IDS intrusion detection prediction explanation module (32) (6005),

[0149] Completing the IDS intrusion detection prediction explanation module (32) (6006).

Claims

CLAIMS1. A computer-aided cyber threat prediction system (1) based on explainable and graph neural networks, including a multi-core processor and a graphics processor, characterized in that it comprises;- At least one pre-modeling module (10) comprising at least one module (11) for collecting data from network sensors that monitors and collects network parameters in the data received by the router monitoring device (3) from at least one network device (2); at least one data preprocessing module (12) that performs data pre-processing and cleaning of data obtained through the module (11) for collecting data from network sensors; at least one time-dependent IDS data flow generation module (13) that generates network flow data from cumulative network packets between a period of time using instantaneous network packets; at least one graph modeling module (14) specific to IDS networks that transforms data into graph modeling,- A modeling module (20), where a machine learning model is trained to predict existing intrusion detection systems, comprising at least one edge embedding representations generation module (21 ) that enables self-learning from unlabeled data, based on the deep graph infomax algorithm, including a trained deep graph infomax model (211) that learns by generating pre-labels; a learning encoder (22), which uses the E-GraphSAGE algorithm and trained DGI edge embedding representations as inputs, obtains edge embedding representations of the graph neural networks-based encoder element for the graph with the E-GraphSAGE algorithm, and while obtaining these representations, proceeds by collecting information from the edges of neighboring nodes, and provides self-learning by maximizing the mutual information between local and global information, comprising the corruption function (221 ), which determines the input graph that adds or removes nodes from the adjacency matrix, corrupts the baseline graph according to the determined function, the encoder function (222), which encodes the E- GraphSAGE model using the corrupted graph and the uncorrupted graph, the corrupted embedded graph (223), which is the graphcorrupted by the corruption function (221 ), and the embedded graph (224), which is the uncorrupted graph with edge embeddings generated by the trained deep graph infomax model (211 ); a learning decoder (23) which uses two edge embedding representations obtained with the learning encoder (22) as input, assigns negative values to the deliberately corrupted edge representations and positive values to the preserved edge representations from these two edge representations, and obtains a trained graph edge embedding learning model by running discrimination, comprising the discriminator function (231), which discriminates between the corrupted embedded graph (223) and the embedded graph (224), and the read function (232), which reads using the uncorrupted embedded graph and the output of the discriminant function (231) and trains the E-GraphSAGE model,- The advanced modeling module (30), which provides real-time detection prediction and explanation using trained models obtained in the modeling module (20) and the pre-modeling module (10), comprising an IDS intrusion detection prediction module (31 ) utilizing edge embedding representations including the trained E-GraphSAGE model (311 ) and the trained CatBoost model (312) to detect and predict whether flows in the network are intrusions or normal using the CatBoost approach; an IDS intrusion detection prediction explanation module (32) that detects data flows using the PGExplainer approach, one of the explainability methods for graphs, comprising the learned PGExplainer model (321) that divides the graph into two parts: the critical subgraph (322), which represents the important subgraph in the baseline graph, and the non- critical subgraph (323), which consists of edges that do not provide high utility in prediction.

2. The cyber threat prediction method, characterized in that it comprises the process steps of;Initiating operation of the module (11) for collecting data from network sensors (1001),- Installing network traffic monitoring tools, updating firmware, and adjusting firewall and IPS / IDS systems for continuous monitoring of network devices in the network topology (1002)- Periodically collecting network data on traffic volume, connection times, and packet rates associated with intrusions on the network that are confirmed and deemed relevant by experts associated with intrusions on the network (1003)- Recording data in a real-time database by the module (11) for collecting data from network sensors (1004)- Initiating the operation of the data pre-processing module (12) (2001 )- Removing single-valued attributes from the data by the data preprocessing module (12), which do not contain any information about time-varying network conditions for network intrusion detection (2002)- Detecting and removing from the data the instances where numeric attributes contain alphabetic characters due to errors during data recording by the data pre-processing module (12) (2003)- Removing non-numeric, categorical attributes from IDS data by the data pre-processing module (12) with “one-hot encoding” and “label encoding” algorithms (2004)- Detecting non-existing values in the dataset by the data pre-processing module (12) and filling them using bfill and / or ffill methods (2005)- Generating a correlation matrix by the data pre-processing module (12) to understand the relationship between the attributes, and keeping only one of the highly correlated attributes in the dataset and removing the other one to reduce computation time and increase the generalization capability of the models (2006)- Determining the acceptable limits for attribute-based sparsity ratios by domain experts, keeping the attributes within these limits in the data, and eliminating the others by the data pre-processing module (12) (2007)- Completing the data pre-processing module (2008)- Initiating the operation of the time-dependent IDS data flow generation module (13) (3001)- Converting instantaneous network packets, collected from network devices and pre-processed, into network flows in a cumulative mannerover a defined period of time by the time-dependent IDS data flow generation module (13) (3002)- Completing the Time-dependent IDS Data Flow Generation Module (3003)- Initiating the operation of the graph modeling (14) module specific to IDS networks (3004)- Converting network devices to represent the nodes and network flows to represent the edges into multi-graph data over the generated network flows by the graph modeling (14) module specific to IDS networks (3005)- Completing the graph modeling module specific to IDS networks (3006)- Initiating the operation of the edge embedding representations generation module (21) (4001),- Obtaining pre-labels by the edge embedding representations generation module (21) through self-learning by running the deep graph infomax algorithm (4002)- Initiating the operation of the Learning Encoder (22) (4003)- Using the trained DGI edge embedding representations as input, obtaining edge embedding representations for the graph by the Learning Encoder (22) through the graph neural network-based Encoder element, E-GraphSAGE algorithm (4004)- Designing the encoder layer as a 1 -layer E-GraphSAGE with the DGI algorithm by Learning Encoder (22) (4005)- Completing the Learning Encoder (22) and initiating the operation of the learning decoder (23) (4006)- Determining a corruption function by the learning decoder (23) to generate a negative (corrupted) graph representation from the input graph (4007)- Generating edge embedding representations from both the input graph and the intentionally corrupted graph by the learning decoder (23) with the encoder layer of the DGI (4008)- Averaging and aggregating the edge embedding representations by the learning decoder (23) with the read function at the center of the DGI, and then processing these through a sigmoid function to compute a global graph summary (4009)- Evaluating the true and corrupted edge embedding representations of the discriminator layer of the DGI by the learning decoder (23), using its global summary as a guide (4010)- Providing a score between 0 and 1 by the learning decoder (23) with the help of the binary cross-entropy loss objective function, based on the comparisons made by the discriminator layer, and distinguishing the edge embedding representations of the true edge and the corrupted edge to train the encoder (4011 )- Obtaining the trained E-GraphSAGE model capable of correctly making graph edge embedding representations, by the learning decoder (23) and recording the model (4012)- Completing the Learning Decoder (23) and completing the edge embedding representations generation module (4013)- Initiating the operation of the IDS intrusion detection prediction module (31) (5001)- Investigating the values for the basic variables of the method with hyperparameter optimization to find the CatBoost method specific to the IDS structure by the IDS intrusion detection prediction module (31 ), and reaching the most accurate prediction model (5002)- Comparing the models obtained with the investigated values through the performance evaluation variables by the IDS intrusion detection prediction module (31), and deciding on the CatBoost model specific to the IDS structure (5003)- Recording a CatBoost intrusion detection model designed specifically for the problem by the IDS intrusion detection prediction module (31 ) (5004),- Detecting new incoming traffic flows as intrusion or normal by the IDS intrusion detection prediction module (31) through the CatBoost model designed specifically for the problem (5005),- Generating an alert log in the system of the network traffic flows detected as intrusions by the IDS intrusion detection prediction module (31) (5006),- Generating an information log in the system of the network traffic flows detected as normal by the IDS intrusion detection prediction module (31) (5007),- Completing the IDS intrusion detection prediction module (31) (5008),- Initiating the operation of the IDS intrusion detection prediction explanation module (32) (6001),- Investigating the most accurate values for the basic variables of the method with hyper-parameter optimization to find the most accurate PGExplainer method by the IDS intrusion detection prediction explanation module (32) (6002),- Deciding on the PGExplainer model by the IDS intrusion detection prediction explanation module (32) based on the performance metrics of fidelity and sparsity (6003),- Generating global explanation with the PGExplainer for the intrusion detection system by the IDS intrusion detection prediction explanation module (32) (6004),- Generating local explanation with the PGExplainer for the instances detected as intrusions by the IDS intrusion detection prediction explanation module (32) (6005),- Completing the IDS intrusion detection prediction explanation module (32) (6006).

3. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a data pre-processing module (12) that removes single-valued attributes from the data, which do not contain any information about the network conditions that change over time in network intrusion detection.

4. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a data pre-processing module (12) that detects and removes from the dataset the instances where the numeric attributes contain alphabetic characters due to errors during data recording.

5. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a data pre-processing module (12) that removes non-numeric attributes, i.e. categorical attributes, from the dataset with one-hot coding and label coding algorithms.

6. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a data pre-processing module (12) that detects non-existing values in the dataset and fills the non-existing values using bfill and / or ffill methods.

7. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a data pre-processing module (12) that generates a correlation matrix to understand the relationship between attributes and keeps one of the attributes that show a high correlation relationship in the dataset and removes the other one to reduce the computation time and increase the generalization capability of the models.

8. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a data pre-processing module (12) that calculates the attribute sparsity ratio by dividing the number of zeros in an attribute column by the number of instances the attribute has, determines the acceptable limits for attribute-based sparsity ratios, and keeps the attributes within these limits in the data, while eliminating the others.

9. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a learning encoder (22) that determines a corruption function (221 ) to generate a negative graph illustration from the input graph.

10. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a trained deep graph infomax model (211 ) that generates edge embedding representations from both the embedded graph (224) and the corrupted embedded graph (223).

11. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a trained deep graph infomax model (211) that averages the edge embedding representations and passes them through a sigmoid function to generate the global graph summary.

12. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a learning decoder (23) that averages and aggregates theedge embedding representations with the read function (232) and then processes them through a sigmoid function to compute a global graph summary.

13. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises an IDS intrusion detection prediction explanation module (32) that removes nodes and attributes and analyzes the effects of these operations on the output of the GNN model.

14. The cyber threat prediction system (1 ) according to claim 1 , characterized in that it comprises a PGExplainer model (321 ) that determines the critical subgraph (322) by maximizing the mutual information with the help of the entropy term.

15. The cyber threat prediction method according to claim 2, characterized in that in the process step of investigating the values for the basic variables of the method with hyper-parameter optimization to find the CatBoost method specific to the IDS structure by the IDS intrusion detection prediction module (31 ), and reaching the most accurate prediction model (5002), the basic variables are minimum samples leaf, minimum samples split, and maximum depth.

16. The cyber threat prediction method according to claim 2, characterized in that in the process step of comparing the models obtained with the researched values through the performance evaluation variables by the IDS intrusion detection prediction module (31 ), and deciding on the CatBoost model specific to the IDS structure (5003), the performance evaluation variables are F1 -Macro, accuracy, and detection rate.

17. The cyber threat prediction method according to claim 2, characterized in that in the process step of generating a correlation matrix by the data pre-processing module (12) to understand the relationship between attributes, and keeping only one of the highly correlated attributes in the dataset and removing the other one to reduce computation time and increase the generalization capability of the models (2006), provided that the correlation ratio is 0.95 or higher, it comprises keeping only one of these attributes in the dataset and removing the other.

18. The cyber threat prediction method according to claim 2, characterized in that in the process step of determining the acceptable limits for attribute-based sparsity ratios by domain experts, keeping the attributes within these limits in the data, and eliminating the others by the data pre-processing module (12)(2007), the acceptable sparsity ratio is less than 0.15.

Citation Information

Patent Citations

  • Structural graph neural networks for suspicious event detection

    US20210067527A1

  • Anomaly Detection Using Graph Neural Networks

    US20230025826A1

  • Method and apparatus for anomaly detection on graph

    WO2023010502A1