Network protection method and system based on attack and defense game model
By constructing a dynamic network graph and using a graph neural network model to generate a Top-K attack hypothesis list, and combining it with MCTS to solve real-time defense strategies, the defense decision-making difficulties in existing technologies under incomplete information and large-scale network environments are solved, and efficient real-time defense strategy generation is achieved.
Patent Information
- Application Number
- CN202510987749.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing technologies make it difficult to make efficient network protection decisions in the context of incomplete information, large network scale and real-time requirements. Especially in a dynamic game environment, it is difficult for defenders to accurately assess attack risks and formulate optimal defense strategies.
By obtaining original logs, asset information, and threat intelligence, a dynamic network graph is constructed, and a graph neural network model is used to generate a Top-K attack hypothesis list. The Bayesian belief network is updated to generate a sub-game model. MCTS is used to solve real-time defense strategies and generate executable defense strategies.
It achieves efficient real-time defense decision-making in incomplete information and large-scale network environments, alleviates the problems of information incompleteness and complexity, can quickly identify potential threats and formulate optimal defense strategies, and avoids the infeasibility of exhaustive calculations.
Smart Images

Figure CN120512302B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network protection, and more specifically, to a network protection method and system based on an attack-defense game model. Background Art
[0002] As cyberattacks become increasingly complex and intelligent, traditional rule-based or signature-based network defense methods are no longer effective. Attackers often exploit zero-day vulnerabilities and social engineering techniques to lurk, infiltrate, and move laterally within networks, making it difficult for defenders to detect and block attacks immediately. This attack-defense dynamic is essentially a dynamic game, with both attackers and defenders constantly adjusting their strategies to maximize their own gains and minimize those of the other side. Therefore, it is crucial to develop a network defense solution that can simulate this dynamic game and make real-time, efficient defense decisions based on it.
[0003] However, existing network defense solutions based on attack-defense game models have yet to gain widespread adoption. This is primarily due to challenges in real-world network environments, such as incomplete information, large network scale, and high real-time requirements. First, incomplete information is a common phenomenon in network attack-defense games. For example, when a security information and event management (SIEM) platform issues a suspicious connection alert, defenders often lack a complete understanding of the attacker's identity, skill level, attack intent, and specific attack path. This information asymmetry makes it difficult for defenders to accurately assess attack risks and develop optimal defense strategies. Second, defense decision-making in large-scale network environments faces immense complexity. An enterprise network may contain tens of thousands of devices and connections. Faced with a suspicious alert, defenders may have multiple possible defense strategies, each with varying costs and risks. Using an exhaustive approach to analyze the impact of each decision on different attacker types would require constructing a massive game tree, a computational complexity that is impractical to complete within a few minutes. Finally, network attacks require extremely high real-time performance. Once suspicious activity is detected, defenders must make and execute decisions swiftly, as any delay could allow the attacker to further penetrate the network and cause greater damage. For example, if a core server has a suspicious connection, isolating the server may cause business interruption, but if it is not isolated in time, the attacker may have infiltrated the intranet, causing more serious consequences.
[0004] Therefore, how to make efficient real-time defense decisions for large-scale networks in the case of incomplete information has become a key technical issue that needs to be urgently addressed in the current field of network security. Summary of the Invention
[0005] Taking into account the above limitations in application, according to one aspect of the present application, a network protection method based on an attack-defense game model is provided, which includes: obtaining original logs, asset information and threat intelligence; constructing a dynamic network graph based on the original logs, asset information and threat intelligence; inputting the dynamic network graph into a trained graph neural network model to obtain a Top-K attack hypothesis list; using each attack hypothesis in the Top-K attack hypothesis list as new evidence, updating the prior Bayesian belief network to obtain a posterior Bayesian belief network; based on the attacker's strategy space, the defender's strategy space and the profit function, generating a sub-game model for the i-th attack hypothesis in the Top-K attack hypothesis list to obtain the i-th sub-game model; based on the posterior Bayesian belief network, solving the i-th sub-game model based on MCTS real-time defense strategy to obtain a recommended defense strategy; translating the recommended defense strategy into executable instructions and executing the strategy.
[0006] According to another aspect of the present application, a network protection system based on an attack-defense game model is provided, which includes: a data acquisition module for acquiring original logs, asset information and threat intelligence; a dynamic network graph construction module for constructing a dynamic network graph based on the original logs, asset information and threat intelligence; a dynamic network graph encoding module for inputting the dynamic network graph into a trained graph neural network model to obtain a Top-K attack hypothesis list; a network update module for updating the prior Bayesian belief network using each attack hypothesis in the Top-K attack hypothesis list as new evidence to obtain a posterior Bayesian belief network; a sub-game model generation module for generating a sub-game model for the i-th attack hypothesis in the Top-K attack hypothesis list based on the attacker's strategy space, the defender's strategy space and the payoff function to obtain the i-th sub-game model; a defense strategy recommendation module for performing a real-time defense strategy solution based on MCTS for the i-th sub-game model based on the posterior Bayesian belief network to obtain a recommended defense strategy; and a strategy translation and execution module for translating the recommended defense strategy into executable instructions and performing strategy execution.
[0007] Compared to existing technologies, this application provides a network protection method and system based on an attack-defense game model, aiming to address the problem of how to efficiently make network protection decisions in the context of incomplete information, large network scale, and real-time requirements. First, by integrating raw logs, asset information, and threat intelligence, a dynamic network graph is constructed to correlate fragmented information and initially alleviate information incompleteness. Next, a trained graph neural network model is used to identify and predict potential top-K attack hypotheses from the dynamic graph. This is equivalent to quickly focusing on possible threat paths within massive amounts of data, effectively addressing the complexity of large-scale networks. Subsequently, these attack hypotheses are used as new evidence to update the Bayesian belief network, further improving the understanding of the attacker's intentions and capabilities. Crucially, a sub-game model is generated for each attack hypothesis, and based on the updated Bayesian belief network, MCTS is used to solve real-time defense strategies. The introduction of MCTS enables efficient exploration and recommendation of optimal defense strategies under incomplete information and complex decision spaces, avoiding the infeasibility of exhaustive computation and thus achieving real-time and efficient defense decisions for large-scale networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 The flowchart of the network protection method based on the attack and defense game model according to the embodiment of the present application.
[0010] Figure 2 Schematic diagram of data flow of a network protection method based on an attack-defense game model according to an embodiment of the present application.
[0011] Figure 3 This is a flowchart of step S2 in the network protection method based on the attack and defense game model according to an embodiment of the present application.
[0012] Figure 4 4 is a block diagram of a network protection system based on an attack-defense game model according to an embodiment of the present application. DETAILED DESCRIPTION
[0013] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. While the drawings illustrate certain embodiments of the present disclosure, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0014] In view of the limitations of the existing background technology, this application proposes a network protection method based on an attack-defense game model. Figure 1 The flowchart of the network protection method based on the attack and defense game model according to the embodiment of the present application. Figure 2 Schematic diagram of data flow of the network protection method based on the attack and defense game model according to the embodiment of the present application. Figure 1 and Figure 2 As shown, according to the network protection method based on the attack and defense game model of the embodiment of the present application, the method includes: S1, obtaining original logs, asset information and threat intelligence; S2, constructing a dynamic network graph based on the original logs, asset information and threat intelligence; S3, inputting the dynamic network graph into the trained graph neural network model to obtain a Top-K attack hypothesis list; S4, using each attack hypothesis in the Top-K attack hypothesis list as new evidence, updating the prior Bayesian belief network to obtain a posterior Bayesian belief network; S5, based on the attacker's strategy space, the defender's strategy space and the profit function, generating a sub-game model for the i-th attack hypothesis in the Top-K attack hypothesis list to obtain the i-th sub-game model; S6, based on the posterior Bayesian belief network, solving the i-th sub-game model based on MCTS real-time defense strategy to obtain a recommended defense strategy; S7, translating the recommended defense strategy into executable instructions and executing the strategy.
[0015] In step S1, raw logs, asset information, and threat intelligence are obtained. Understandably, the complexity of network attack and defense requires defenders to have as comprehensive a grasp of the network situation as possible in order to effectively identify potential threats, assess risks, and develop reasonable defense strategies. It's worth noting that raw logs provide the most direct record of network activity and are the source of abnormal behavior; asset information clarifies the value and vulnerability of various resources in the network, helping to assess the potential damage caused by an attack; and threat intelligence provides contextual information on known external threats, helping defenders predict the attacker's possible means and goals. The fusion of this multi-source, heterogeneous information is the cornerstone for constructing dynamic network maps, conducting attack hypothesis reasoning, and subsequent game decision-making. It can effectively alleviate the information incompleteness issue mentioned in the background technology and provide sufficient data support for subsequent intelligent analysis and decision-making.
[0016] Optionally, step S1 can be implemented in the following manner: Raw logs are real-time or historical data from various security devices and servers on the network, such as connection logs from firewalls, alarm logs from intrusion detection systems (IDS), system call logs from host intrusion detection systems (HIDS), and access logs from application servers. These logs contain a large amount of network activity records, such as source IP addresses, destination IP addresses, ports, protocols, timestamps, and event types. This log data is typically stored in various formats and requires preliminary collection and preprocessing, such as real-time ingestion or periodic synchronization from various data sources using a log collection agent or API.
[0017] Asset information originates from the enterprise's configuration management database (CMDB) or other asset management system. It details the attributes of all assets on the network, including but not limited to the server's operating system type, version, open ports, installed application services, and its importance to business processes, such as the data value level. This information is crucial for assessing the potential damage caused by an attack and determining defense priorities. For example, a database server hosting core business data might have a high asset value, while a standard office PC might have a low asset value. This asset information is typically stored as structured data and can be retrieved directly through database queries or APIs.
[0018] Threat intelligence is structured information about known or potential cyber threats, provided by third-party security vendors or intelligence agencies. It includes malicious IP addresses, malicious domain names, known vulnerability reports (such as Common Vulnerability Event (CVE) numbers), attacker organizational information, and attack TTPs (Tactics, Techniques, and Procedures). This intelligence can help identify known attack patterns and threat sources, providing prior knowledge for subsequent attack hypothesis generation. For example, when a connection to a known malicious IP appears in raw logs, threat intelligence can immediately provide a record of that IP's historical malicious behavior. Threat intelligence is updated and accessible in real time via subscription services or APIs.
[0019] In step S2, a dynamic network graph is constructed based on the original logs, asset information and threat intelligence. Accordingly, it is difficult for traditional log analysis to fully reveal the complex relationships between network entities and potential attack paths. Therefore, in order to convert the scattered and heterogeneous original data into a unified, intuitive and semantically rich network topology, thereby solving the problems of incomplete information and large-scale network complexity mentioned in the background technology. This application constructs a dynamic network graph to abstract the devices, users, applications, etc. in the network as nodes, and abstract the communication, access and other behaviors between them as edges, and integrates asset value and threat intelligence, so that defenders can understand the network status from a global perspective, identify abnormal behavior, and provide structured input for subsequent attack hypothesis generation and game decision-making. This graphical representation method greatly improves the efficiency and accuracy of network situation awareness.
[0020] In an optional embodiment of the present application, Figure 3 FIG. 1 is a flow chart of step S2 in the network protection method based on the attack-defense game model according to an embodiment of the present application. Figure 3 As shown, step S2, based on the original log, asset information and threat intelligence, constructs a dynamic network graph, including: S21, parsing, standardizing and correlation analysis of the original log to identify suspicious activities from the original log; S22, combining asset information and threat intelligence to construct a dynamic graph; S23, marking the suspicious activities on the corresponding nodes or edges of the dynamic graph to obtain the dynamic network graph.
[0021] Step S2 can be specifically implemented in the following ways: First, execute S21. The original log formats are diverse. For example, the firewall log may record time, source IP, destination IP, source port, destination port, protocol, and action, while the IDS alarm log may contain time, alarm ID, alarm level, source IP, destination IP, and attack type. In the parsing stage, for different types of logs, a dedicated parser is used, such as Logstash's grok plug-in or Fluentd's parser, to convert these different types of text logs such as Syslog, JSON, key-value pairs, etc. into unified structured data in JSON format. In the standardization stage, the fields representing the same concept in different log sources are uniformly named. For example, source_ip, client.ip, and src_addr are all uniformly mapped to a standard field source_ip. The correlation analysis stage is the core of identifying suspicious activities. By setting rules, isolated events are connected into meaningful sequences. For example, the rule engine can discover that within one minute, if a host's HIDS log reports the creation of a WebShell file (Alert_A), and the host's firewall log records that it has initiated a connection to a malicious IP address from threat intelligence (Alert_B), then the system will no longer report two independent low-risk alerts, but will merge them to generate a high-confidence Alert_Events: the host has been implanted with a WebShell and connected to the C2 server. Through these analyses, specific suspicious activities such as suspicious port scans, abnormal login attempts, and malicious file transfers can be identified.
[0022] Next, execute S22. First, use asset information to construct the basic skeleton of the graph. Asset information comes from the enterprise's configuration management database or asset inventory. In specific implementation, each independent asset in the network, such as a server, workstation, network device, application, user account, or service port, is abstracted as a node in the graph. Each node is assigned inherent attributes derived directly from the asset information, such as IP address, host name, operating system type, business function, business value rating, and a list of open ports. Simultaneously, edges are created between corresponding nodes based on the physical connections, logical connections, or dependencies described in the asset information. For example, if an application-layer access relationship exists between web server Web-01 and database server DB-01, an edge representing this access relationship is created between the Web-01 node and the DB-01 node, optionally accompanied by attributes such as protocol and port. Next, threat intelligence is integrated into the constructed asset graph. Threat intelligence includes malicious IP addresses, malicious domain names, known vulnerability information, and the latest attack techniques. In specific implementation, threat intelligence can be integrated in two main ways: first, by supplementing the attributes of existing asset nodes. For example, if threat intelligence shows that a known vulnerability affects a software version running on the Web-01 server, then a vulnerability attribute will be added to the Web-01 server node, along with the vulnerability ID and risk level. The second is to add it to the graph as an independent threat node and establish a connection with the relevant asset nodes. For example, if the threat intelligence contains a known command and control server IP address, it can be added to the graph as a malicious IP node. If log analysis finds that an internal server has communicated with this malicious IP, then an edge representing the communication is created between the internal server node and the malicious IP node and marked as suspicious. It is worth mentioning that the construction of this graph is ongoing. When new asset information or threat intelligence flows in, the graph database will update node attributes, add or delete nodes and edges accordingly.
[0023] Finally, execute S23. For each suspicious activity identified, it is necessary to parse the key entities involved in the activity. These entities include source IP address, destination IP address, affected host name, user account, specific service port or communication protocol, etc. Then, locate the nodes or edges corresponding to these key entities on the constructed dynamic graph. If the suspicious activity directly points to a certain network asset, for example, WebShell file creation occurs on a specific server, then the activity will be marked on the node representing the server. The specific operation is to add a new field, such as a suspicious activity list, to the properties of the server node, such as the Web-01 node, and use the WebShell file creation activity and related information such as timestamp, confidence, and original log ID as the value of the field. In this way, the node carries its own asset attributes and current security event information. If the suspicious activity describes an abnormal interaction or communication behavior between two entities, then the activity will be marked on the edge connecting the two entities. For example, if a suspicious port scan is identified as communication from the attacker's IP address to the target server's IP address, the dynamic graph will be constructed to identify the node representing the attacker's IP address and the node representing the target server's IP address. A suspicious activity attribute will be added to the existing edge between them, marking it as a suspicious port scan and providing detailed information such as the scanned port and time. This labeling process is ongoing, and as new suspicious activities are identified, the dynamic network graph will be updated in real time to ensure that it always reflects the latest network security status.
[0024] In step S3, the dynamic network graph is input into the trained graph neural network model to obtain a list of Top-K attack hypotheses. It should be understood that although the dynamic network graph provides a structured network situation, its original node and edge attributes are usually discrete and high-dimensional, and difficult to be directly processed by the machine learning model. Based on this, the present application can convert these complex graph structure data into low-dimensional, continuous vector representations, i.e., embedded coding, through the graph neural network model, thereby capturing the deep semantic relationships and topological structure information between nodes and edges. This enables the model to automatically learn and identify potential attack patterns and paths from massive and complex network data, generate Top-K attack hypotheses, greatly improve the efficiency and accuracy of attack predictions, and provide accurate input for subsequent game decisions.
[0025] In an optional embodiment of the present application, step S3, inputting the dynamic network graph into a trained graph neural network model to obtain a Top-K attack hypothesis list, includes: S31, passing the nodes and edges in the dynamic network graph through an embedding layer to obtain a dynamic network embedding coding graph; S32, inputting the dynamic network embedding coding graph into the trained graph neural network model to obtain a set of node final vectors; S33, extracting the node final vector corresponding to the suspicious activity from the set of node final vectors as the alarm node final vector; S34, performing link prediction on the alarm node final vector to obtain a link prediction result; S35, performing node classification on all node final vectors in the set of node final vectors to obtain a node classification result; S36, determining the Top-K attack hypothesis list based on the link prediction result and the node classification result.
[0026] In other words, the node and edge attributes in the original graph are often discrete or difficult to directly process by machine learning models. The embedding layer learns the deep semantic and topological structure of these complex attributes and maps them into a unified vector space. This encoding method not only captures the characteristics of the nodes and edges themselves, but also implicitly reflects their contextual relationships in the network, providing high-quality, computable input for subsequent graph neural network models.
[0027] Step S31 can be implemented in the following manner: First, feature extraction is performed on each node and edge in the dynamic network graph, converting their raw attributes into numerical vectors. For nodes, raw attributes may include IP address, host name, operating system type, asset value rating, open port list, and suspicious activity information such as webshell implant alerts. For edges, raw attributes may include protocol type (TCP / UDP), port number, traffic flow, connection status (allowed / denied), and suspicious activity information such as port scan alerts. These raw attributes are of mixed types, including both categorical attributes (such as operating system type and protocol) and numerical attributes (such as asset value rating and traffic flow).
[0028] Next, these raw attributes are vectorized using an embedding layer. The purpose of the embedding layer is to map high-dimensional, sparse, or discrete features into a low-dimensional, dense, continuous vector space. For categorical features, one-hot encoding can be used followed by a fully connected layer for embedding, or the embedding matrix can be used directly for lookup. For example, the operating system type "Windows Server 2019" can be mapped to a specific embedding vector. For numerical features, they can be directly included as part of the embedding vector or input after normalization. Suspicious activity information can be included as an additional attribute of a node or edge and assigned a specific embedding vector, or it can be concatenated with the embedding vectors of existing attributes. For example, the raw attribute vector of a node Web-01 might include its IP address, operating system, asset value, etc. When it is marked as "WebshellImplant," the alert information is converted into an alert embedding vector and concatenated or weighted summed with the original attribute embedding vector of Web-01 to form a comprehensive embedding vector for the node. Specifically, the node embedding layer generates a fixed-dimensional node embedding vector for each node in the graph. This vector not only contains the node's own attribute information but also implicitly encodes its topological position in the network and its connections with other nodes. For example, a server node might be embedded as a 128-dimensional vector, which contains its characteristics as a web server, a database server, and information about its high asset value rating. Similarly, the edge embedding layer generates a fixed-dimensional edge embedding vector for each edge in the graph. This vector represents the characteristics of the communication or interaction connecting two nodes. For example, the embedding vector of an edge representing communication on TCP port 443 will reflect the characteristics of HTTPS traffic. It is worth noting that the weights and bias parameters of these embedding layers are automatically learned during the training of the graph neural network model through backpropagation and optimization algorithms. Ultimately, after processing by the embedding layer, the entire dynamic network graph is transformed into a dynamic network embedding encoding graph consisting of node embedding vectors and edge embedding vectors, where each node and edge is represented by a continuous numeric vector.
[0029] Accordingly, while dynamic network embedding coding graphs vectorize raw attributes, these vectors primarily reflect the local characteristics of the nodes and edges themselves. Cyberattacks are often complex, multi-stage, multi-hop behaviors that require understanding the global context of the node within the entire network topology. To address this, graph neural networks are required to propagate and aggregate information across the graph structure, incorporating each node's neighbor information, multi-hop connectivity information, and edge attributes into the node representation. This generates a final node vector that reflects the node's potential role in the attack path.
[0030] Step S32 can be implemented in the following manner: The pre-trained graph neural network model adopts a multi-layer structure, such as a graph convolutional network, a graph attention network, or GraphSAGE. Taking a graph convolutional network as an example, its basic operation is to update the feature representation of each node layer by layer. Specifically, for each node in the graph, its feature vector at the current layer is updated to the feature vector at the next layer by aggregating the features of its neighboring nodes. This aggregation process can be understood as each node collecting information from its directly connected neighbors and processing it in combination with its own information to form a new, richer feature representation. In a multi-layer graph convolutional network, information is propagated from neighboring nodes to the current node layer by layer. For example, the first layer of the graph convolutional network aggregates features of direct neighbors, the second layer aggregates features of two-hop neighbors, and so on. By stacking multiple graph convolutional layers, the final vector of each node incorporates rich information from its multi-hop neighbors, thereby capturing a broader network context. This graph neural network model is trained on a large amount of historical network data. During training, the model's internal parameters (e.g., weights and biases used for feature transformation and aggregation) are automatically adjusted through the learning process. After training, when a new dynamic network embedding code graph is input into the model, the model performs multi-layer information aggregation and transformation on each node's embedding vector based on its learned parameters, ultimately outputting a set of final vectors for all nodes. These final vectors are high-dimensional, continuous numerical representations of each node in the current network state, incorporating its own attributes and neighboring contextual information.
[0031] As you can understand, suspicious network activity has been identified and flagged in S21 and S23, occurring at specific nodes or edges. The final node vector set contains rich information about all nodes in the network, but not all nodes are directly related to the current suspicious activity. By extracting the final vectors of nodes directly associated with suspicious activity, namely the alerting nodes, irrelevant information can be effectively filtered out, focusing subsequent link prediction and node classification computational resources and attention on the starting points or key links most likely to constitute attack paths. This improves the efficiency and accuracy of attack hypothesis generation and mitigates the computational burden brought by the complexity of large-scale networks.
[0032] Step S33 can be specifically implemented in the following manner: First, the suspicious activities have been marked in S23 on the corresponding nodes or edges of the dynamic network graph. Each suspicious activity is associated with one or more specific networks, i.e., nodes. For example, if a suspicious activity created by a webshell file is identified and occurs on the Web-01 server, then the Web-01 server is the node corresponding to the suspicious activity. Similarly, if an abnormal login attempt is identified as occurring on the User-A account, then the node corresponding to User-A is the alarm node. Similarly, if an abnormal login attempt is identified as occurring on the User-A account, then the node corresponding to User-A is the alarm node. Next, from the set of node final vectors, based on the node identifiers associated with these suspicious activities, such as IP addresses, host names, user IDs, etc., the final vectors of these specific nodes are accurately searched and extracted. For example, if the set of node final vectors is a dictionary or list, where the keys are node identifiers and the values are corresponding final vectors, then the corresponding vectors can be retrieved from the set based on the node IDs involved in the marked suspicious activities. For example, let's assume the following suspicious activities are flagged: a webshell implant alert on node Web-01 (IP: 192.168.1.10); and an abnormal database access alert on node DB-01 (IP: 10.0.0.5). From the node final vector set, the final vectors corresponding to nodes Web-01 and DB-01 are extracted as the final vectors for the alert nodes. These extracted vectors constitute the set of alert node final vectors.
[0033] Accordingly, network attacks are often not isolated events, but rather comprise a series of interconnected steps. Link prediction can predict which potential connections or paths are most likely to form part of an attack path based on the current state of the alerting node and its connections to neighboring nodes. This helps fill in missing links in attack paths when information is incomplete, allowing for a more complete attack hypothesis. This addresses the issues of incomplete information and large-scale network complexity mentioned in the background technology, providing more accurate attack path inferences for subsequent defense decisions.
[0034] In an optional embodiment of the present application, step S34, performing link prediction on the final vector of the alarm node to obtain a link prediction result, includes: S341, extracting the node final vector of a neighbor node of the final vector of the alarm node as the neighbor node final vector; S342, performing association coding based on a fully connected layer on the final vector of the alarm node and the final vector of the neighbor node, and inputting the association coding result into a classifier to obtain the link prediction result, which is used to indicate the possibility that the path from the alarm node to the neighbor node is part of the attack path.
[0035] Step S34 can be specifically implemented in the following way: First, execute S341. For example, if the Web-01 node (whose final vector has been used as the final vector of the alarm node) is connected to DB-01 and App-Server-01 in the graph, then the final vectors of DB-01 and App-Server-01 will be extracted as the final vector of the neighbor node of Web-01. Then, in S342. For each alarm node, for example, alarm node A, whose final vector is VA and one of its neighbor nodes, for example, neighbor node B, whose final vector is VB, these two vectors are concatenated or element-wise operated, such as the absolute value of element-wise multiplication or subtraction, to form a joint vector representing the relationship between the pair of nodes. For example, VA and VB can be simply concatenated to obtain a longer vector X=[VA;VB]. This joint vector is then input into one or more fully connected layers for associative encoding. The role of the fully connected layer is to learn and extract deep features from the joint vector through nonlinear transformations that can characterize whether the connection between the two nodes is part of the attack path. Each fully connected layer contains a weight matrix and the bias vector , which is calculated as ,in is the input vector, i.e. the concatenation vector of VA and VB, is a nonlinear activation function such as Reluctant Unit (ReLU). These weights and biases are learned during the model training phase to maximize link prediction accuracy. After associative encoding, the resulting encoding output, a fixed-dimensional vector, is input to a classifier. This classifier is a binary classifier, such as a single-layer fully connected network with a Sigmoid activation function or a multi-layer perceptron with a Softmax activation function. Based on the associative encoding results, the classifier outputs a probability value between 0 and 1, indicating the likelihood that the path from alerting node A to neighboring node B is part of the attack path. For example, an output probability of 0.9 indicates a high probability that the link is part of the attack path. Specifically, the weights and bias parameters of this classifier and the fully connected layer are learned jointly during the training of the graph neural network model. The final output is a series of link prediction results for each alerting node and all its neighboring nodes. Each result is a probability value indicating the likelihood that the corresponding link is part of the attack path.
[0036] It's worth noting that attackers' ultimate goal in lateral movement and infiltration within a network is to control or steal critical assets. Relying solely on link prediction to infer attack paths is insufficient; it's also necessary to assess the likelihood of each node along the path being the ultimate attack target. To this end, by classifying all nodes, each node can be assigned a target probability. This helps quickly locate high-value, high-risk potential attack targets in complex network environments, providing critical endpoint information for subsequent attack hypothesis generation.
[0037] In an optional embodiment of the present application, step S35, performing node classification on all node final vectors in the set of node final vectors to obtain a node classification result, including: inputting each node final vector in the set of node final vectors into a classifier to obtain the node classification result, wherein the node classification result is used to represent the probability that each node is the final attack target.
[0038] Step S35 can be implemented in the following manner: each final node vector in the set of final node vectors is input into a pre-trained classifier to obtain the node classification result. This classifier is a binary classification model designed to determine whether each node is a final attack target. The classifier can be a simple logistic regression model, a multi-layer perceptron, or a more complex neural network structure. Taking a simple multi-layer perceptron classifier as an example, its architecture may include one or more fully connected layers, each followed by a nonlinear activation function such as ReLU. The final layer is an output layer that uses a Sigmoid activation function to output a probability value between 0 and 1. In specific implementation, each final node vector in the set is used as input to the classifier. The classifier processes the input vector and extracts features through a series of linear transformations and nonlinear activations. These linear transformations are determined by the classifier's internal weights and bias parameters, which are learned during the model training phase. Ultimately, the output is a series of node classification results for all nodes in the network. Each result is a probability value, indicating the likelihood that the corresponding node is a final attack target. For example, if the node classification result of a server is 0.95, it means that there is a high probability that the server is the final attack target.
[0039] It should be understood that link prediction provides the attacker's next possible lateral movement direction, while node classification indicates which nodes are most likely to be the final targets of the attack. Link prediction alone can generate a large number of potential paths, while node classification alone cannot reveal the complete attack process. By combining the two, a complete attack chain can be constructed, starting from the alert node, passing through a series of high-probability links, and ultimately reaching the high-probability target node. This provides defenders with the most likely attack scenarios, thus supporting subsequent game-playing decisions and defense strategy formulation.
[0040] Step S36 can be implemented in the following manner: The link prediction results are a set containing the probability values of each potential link from each alarm node to its neighboring nodes being part of the attack path. The node classification results are a set containing the probability values of each node in the network being the final attack target. Based on these probability values, potential attack hypotheses are constructed and evaluated. An attack hypothesis can be defined as a complete path starting from an alarm node, passing through a series of high-probability links, and ultimately reaching a high-probability target node. This process can utilize a graph traversal algorithm, such as depth-first search (DFS) or breadth-first search (BFS), to explore potential attack paths starting from each alarm node. During the path exploration process, each time a potential link is traversed, its aggressiveness score is evaluated in combination with its link prediction results. For example, the link prediction probability can be directly used as the link's aggressiveness score. Furthermore, when the path reaches a node, the node classification results are combined to evaluate the likelihood of that node being the final attack target.
[0041] To determine the Top-K attack hypotheses, a comprehensive score is calculated for each potential attack path. This comprehensive score can be determined by multiplying the aggressiveness scores of all links along the path by the final attack target probability of the path's endpoint. For example, the comprehensive score for a path can be calculated as: Path Comprehensive Score = (Link 1 Aggressiveness Probability × Link 2 Aggressiveness Probability × Link N Aggressiveness Probability) × End Node Target Probability. Here, the link aggressiveness probability is the output of S34, and the end node target probability is the output of S35. This product form ensures that a path only achieves a high comprehensive score when all links along the path have high aggressiveness and the end node is also a high-value target. In practice, an upper limit on the path length can be set, for example, only considering attack paths with no more than five hops, to avoid generating overly long and unrealistic hypotheses. Furthermore, a minimum link aggressiveness probability threshold, such as 0.6, can be set. Only links with predicted probabilities above this threshold are considered, reducing computational effort and filtering out low-confidence paths.
[0042] Finally, all potential attack paths with calculated comprehensive scores are sorted, and the K paths with the highest scores are selected as the Top-K attack hypothesis list. The K value can be preset based on actual needs. For example, K can be set to 5 or 10, indicating that the 5 to 10 most likely attack scenarios need to be identified.
[0043] Preferably, in another embodiment of the present application, step S3, inputting the dynamic network graph into a trained graph neural network model to obtain a Top-K attack hypothesis list, includes: S3-1, passing the nodes and edges in the dynamic network graph through an embedding layer to obtain a dynamic network embedding coding graph; S3-2, inputting the dynamic network embedding coding graph into the trained graph neural network model to obtain a set of node final vectors; S3-3, performing local association-global association balance on the set of node final vectors to obtain a set of optimized node final vectors; S3-4, extracting the optimized node final vector corresponding to the suspicious activity from the set of optimized node final vectors as the alarm node final vector; S3-5, performing link prediction on the alarm node final vector to obtain a link prediction result; S3-6, performing node classification on all optimized node final vectors in the set of optimized node final vectors to obtain a node classification result; S3-7, determining the Top-K attack hypothesis list based on the link prediction result and the node classification result. In particular, the specific processing procedures of steps S3-1 and S3-2 are the same as S31 and S32 in step S3 in the above embodiment, and thus will not be elaborated here.
[0044] In particular, here, the focus is on step S3-3, which performs a local correlation-global correlation balance on the set of node final vectors to obtain a set of optimized node final vectors for detailed explanation. It is worth mentioning that the node final vectors output by the graph neural network model need to support two different types of prediction tasks at the same time: one is link prediction based on two node final vectors, which focuses more on the local correlation and interaction characteristics between nodes; the other is target prediction based on a single node final vector, which focuses more on the internal characteristics of the node itself. Since these two tasks have different requirements for the scale pattern of feature distribution, this balancing process is required to improve the consistency balance of local correlation and global correlation aggregation of each node final vector while retaining the network topology correlation structure, thereby optimizing subsequent link prediction and target prediction results. This ensures that the node final vector can accurately reflect its own local characteristics and effectively capture its global correlation information in the entire network topology.
[0045] Based on this, in an optional embodiment of the present application, step S3-3, performing local association-global association balance on the set of node final vectors to obtain a set of optimized node final vectors, includes: first, arranging the set of node final vectors in two dimensions to obtain a node topology matrix. It should be understood that although graph neural networks can process graph data, when performing deeper association analysis, it is more intuitive to represent the topological structure of the graph in matrix form, which is also convenient for subsequent mathematical operations. In other words, by arranging the set of node final vectors in two dimensions according to their connection relationship in the graph, a matrix reflecting the network topology can be explicitly constructed. In this way, a node topology matrix that clearly represents the connection relationship between nodes is obtained. This matrix is the basis for the subsequent calculation of global association information, and it presents the connection relationship between nodes in the network in a structured manner.
[0046] Next, a local-global low-dimensional scale calculation is performed on the set of the node final vectors and the node topology matrix to obtain the node local-global low-dimensional scale latent variable parameters, namely: ;in, is the node topology matrix, is the set of node final vectors The final vector of the nodes, It is calculated The F-norm of It is calculated The second norm of is a low-dimensional latent variable parameter that scales the local-global relationship between nodes. Consequently, when aggregating neighbor information, graph neural networks may prioritize local details over global structure, leading to unbalanced information fusion. Therefore, this application introduces this latent variable parameter to dynamically adjust this emphasis, thereby better coordinating local and global information. It is a scalar value, which is calculated by comparing the norm of the node topology matrix (representing global structural information) with the sum of the final vector norms of all nodes (representing local feature information), thus reflecting the relative weight or scale difference between local information and global information in the current graph. It provides a key balance factor for subsequent association strength calculations.
[0047] Then, based on the local-global low-dimensional scale latent variable parameters of the nodes, the information latent correlation strength between any two node final vectors in the set of node final vectors is calculated to obtain the node information latent correlation matrix, that is: ;in, is the set of node final vectors The final vector of the nodes, yes and The information implicit correlation strength between them. It should be understood that in order to construct an association metric that can comprehensively consider the node's own feature strength and its relative position in the global topology based on the balance parameter calculated above to capture the deeper, non-direct topological implicit correlation between nodes, and make up for the shortcomings of relying solely on direct connection relationships. This application obtains the node information implicit correlation matrix by calculating the information implicit correlation strength between any two node final vectors. That is to say, each element of this matrix represents the information implicit correlation strength between any two node final vectors. It is calculated by combining the local-global low-dimensional scale latent variable parameters and the norm of the final vectors of the two nodes. This matrix not only reflects the characteristic strength of the node itself, but also The adjustment takes into account the relative importance and relevance of nodes in the overall network, thereby realizing the integration of short-range and long-range association information between nodes, and at the same time indirectly retains the topological association structure through the projection characteristics of the vector norm.
[0048] The node information implicit association matrix is used as the association consistency measurement space, and the set of the node final vectors is mapped to obtain the set of the optimized node final vectors, that is: ;in, is the node information implicit association matrix, is vector multiplication, is the set of final vectors of the optimized node Optimize the final node vectors. That is, to ensure that each node's final vector can better support both link prediction (which requires considering associations with other nodes) and target prediction (which requires stronger representation of its own features), the original node final vectors are reshaped using the constructed information latent association matrix as the transformation space. This allows alignment of local consistency and global consistency to obtain a set of optimized node final vectors. Specifically, the final vector of each node is recalculated by mapping the original node final vectors to the node information latent association matrix. Each vector within this optimized set of vectors better integrates the node's own local features and its global association information across the entire network, thereby achieving a balance of aggregate consistency between local and global associations. This enables these optimized vectors to perform better in subsequent link prediction and node classification tasks because they can more comprehensively and accurately characterize the characteristics of the node and its role in complex networks. In particular, the specific processing steps S3-4 to S3-7 are the same as S33 to S36 in steps S3 of the above embodiment, and therefore will not be elaborated on here.
[0049] In step S4, the a priori Bayesian belief network is updated using each attack hypothesis in the Top-K attack hypothesis list as new evidence to obtain a posterior Bayesian belief network. Furthermore, in order to integrate the inferred attack hypothesis into a higher-level attack intention reasoning framework built based on expert knowledge and domain experience, the present application updates the a priori Bayesian belief network. That is, the Bayesian belief network can effectively process uncertain information and perform probabilistic reasoning. The Top-K attack hypothesis is a prediction of the attack path based on current network data, but these predictions themselves may be uncertain and fail to directly link to the attacker's higher-level intentions such as data theft, service interruption, etc. By using these attack hypotheses as new evidence in the Bayesian belief network, the probability distribution of each node in the Bayesian belief network representing the attack stage, attack technology, attack intention, etc. can be dynamically updated, thereby inferring what the attacker wants to do from what has happened, providing more accurate attack intention judgment for subsequent attack and defense games.
[0050] Optionally, step S4 can be implemented in the following manner: First, a pre-built prior Bayesian belief network is required. This prior Bayesian belief network is a directed acyclic graph (DAG), whose nodes represent various events, states, or concepts in a cyberattack, such as initial access, lateral movement, privilege escalation, data theft, service disruption, and ransomware attack intent. Directed edges between nodes represent causal or conditional dependencies. Each node has a conditional probability table (CPT), which defines the probability of the node being in different states given the state of its parent node. These CPTs and network structures are pre-defined and constructed based on cybersecurity expert knowledge, historical attack event analysis, and frameworks such as MITRE ATT&CK. For example, the CPT for a successful lateral movement of a node would define the probability of its occurrence given successful initial access.
[0051] Next, the list of top-K attack hypotheses is fed into this prior Bayesian belief network as new evidence. Each attack hypothesis in the list describes an attack path from the alerting node to a potential target node. These attack hypotheses can be mapped to specific nodes or combinations of nodes in the prior Bayesian belief network as observed evidence. For example, if an attack hypothesis describes a SQL injection and data exfiltration path from a web server to a database server, the corresponding SQL injection and data exfiltration nodes in the prior Bayesian belief network will be set to observed or with a high probability of occurrence.
[0052] In practice, each attack hypothesis in the Top-K list is parsed and mapped to a corresponding evidence node in the prior Bayesian belief network. For example, if an attack hypothesis is Web-01 (webshell) -> DB-01 (SQL injection) -> Data exfiltration, then the nodes in the prior Bayesian belief network related to Web-01 being attacked, DB-01 being attacked, SQL injection occurring, and data exfiltration occurring will be set as evidence nodes. Based on the comprehensive score of the attack hypothesis in S36, the corresponding evidence strength will be assigned. For example, the higher the comprehensive score, the greater the evidence strength.
[0053] The prior Bayesian belief network is then updated using Bayesian inference algorithms such as variable elimination, belief propagation, or Monte Carlo sampling. The core of this update process is to recalculate the posterior probability distribution of all unobserved nodes in the prior Bayesian belief network, particularly those representing attack intent, based on new evidence, according to Bayes' theorem. For example, when SQL injection and data leakage are input as evidence, the posterior Bayesian belief network will infer, based on its internal conditional probability table and network structure, that the probability of data theft intent will increase significantly, while the probability of service disruption intent may remain unchanged or slightly decrease.
[0054] Ultimately, the output is a posterior Bayesian belief network. This posterior Bayesian belief network contains the latest probability distribution of attack phases, attack techniques, and attack intent for all nodes in the network, taking into account the current top-K attack hypotheses as evidence. In particular, it provides the probability of the attacker's most likely current attack intent, for example, the probability of data theft is 0.8 and the probability of ransomware attack intent is 0.1.
[0055] In step S5, based on the attacker's strategy space, the defender's strategy space, and the profit function, a sub-game model is generated for the i-th attack hypothesis in the Top-K attack hypothesis list to obtain the i-th sub-game model. It should be understood that network attack and defense is essentially a dynamic and adversarial process, and both attackers and defenders are constantly adjusting their strategies to maximize their own benefits. Traditional defense methods based on rules or signatures lack the ability to predict the attacker's intentions and behaviors, and cannot evaluate the actual effects of different defense strategies. Based on this, the present application constructs a sub-game model for each Top-K attack hypothesis, which can formally describe the possible actions, action costs, and gains / losses of the attacker and defender's actions, so that the game theory method can be used to predict the attacker's optimal strategy and find the optimal response strategy for the defender.
[0056] In an optional embodiment of the present application, step S5, based on the attacker strategy space, the defender strategy space and the profit function, generates a sub-game model for the i-th attack hypothesis in the Top-K attack hypothesis list to obtain the i-th sub-game model, including: for the i-th attack hypothesis in the Top-K attack hypothesis list, performing the following operations: S51, extracting relevant nodes, possible attack actions, and possible defense actions from the i-th attack hypothesis; S52, extracting attack actions corresponding to the possible attack actions from the attack knowledge base to obtain a set of attack actions; S53, extracting defense actions corresponding to the possible defense actions from the defense action library to obtain a set of defense actions; S54, based on the asset information and the cost of each defense action in the defense action library, calculating the set of attack actions and the profit / loss value between any pair of attack actions and defense actions in the set of defense actions to obtain the i-th sub-game model.
[0057] Step S5 can be implemented in the following manner: First, execute S51. The attack hypothesis is presented in a structured text format, describing a potential attack path from the initial alert node to the final target node and potentially including key techniques or behaviors used by the attacker along the path. For example, a typical attack hypothesis might be: Web-01 (Webshell) -> DB-01 (SQL Injection) -> Data Exfiltration. This indicates that the attacker first implants a webshell on the Web-01 server, then uses the webshell to launch a SQL injection attack against the DB-01 database, ultimately attempting to exfiltrate data. The implementation process for this input attack hypothesis consists of three main parts: The first step is to extract relevant nodes, identifying all network entities involved in the attack hypothesis. These entities will become the focus of the subsequent attack-defense game analysis. Specifically, the text description of the attack hypothesis is parsed to identify the names or identifiers of network nodes such as servers, databases, and user accounts explicitly mentioned. For example, in the hypothesis Web-01 (Webshell) -> DB-01 (SQL Injection) -> Data Exfiltration, the parsing clearly extracts the two nodes Web-01 and DB-01. Furthermore, if the attack hypothesis implies an intermediate springboard or other affected assets, such as data outflow through an intermediate proxy server, the proxy server node should also be identified and included in the relevant node set. These extracted nodes are key assets along the attack path, and their status and value will directly impact the payoff calculation of the attack-defense game. The second step is to extract possible attack actions, identifying the specific types of actions the attacker may take from the attack hypothesis. This is done by analyzing the attack techniques, vulnerability exploitation methods, or attack phases described in the attack hypothesis. For example, in the above assumptions, two attack techniques, Webshell and SQL injection, can be identified, along with the attack objective or phase, Data Exfiltration. These identified techniques and objectives are considered possible attack actions that an attacker might take in the current attack scenario. These actions form the basis of the attacker's strategy space, representing the actions an attacker might perform on a specific node or against a specific target. For example, identifying a Webshell means that the attacker might attempt to execute commands, upload files, or escalate privileges through the Webshell. The third step is to extract possible defensive actions. Based on the relevant nodes and possible attack actions identified in S51, the defender's possible response measures are inferred. The specific implementation relies on preset mapping rules or a security policy knowledge base.For example, if a webshell attack is identified on the Web-01 server, possible defender actions include isolating the Web-01 server, deleting webshell files, and updating web application firewall (WAF) rules to block webshell traffic. If a SQL injection attack is identified, possible defender actions include patching the SQL injection vulnerability, restricting database access permissions, and deploying a database firewall. The final output is a list of all relevant nodes, a list of all possible attack actions, and a list of all possible defense actions for the i-th attack hypothesis.
[0058] Next, execute S52. First, a pre-built attack knowledge base is required. This attack knowledge base is a structured database containing detailed information on a large number of known network attack techniques, tools, procedures, and strategies. It can organize and categorize attack behaviors based on widely recognized industry frameworks, such as MITRE ATT&CK. Each entry in the knowledge base details a specific attack action, including its name, description, possible targets, required conditions, expected effects, likelihood of detection, and the typical cost for an attacker to carry out the action, such as required resources and complexity. This information is collected and maintained based on historical attack event analysis, security research reports, and expert experience. During the extraction process, the system matches and queries each entry in the list of possible attack actions in the attack knowledge base. This matching process can be implemented in various ways, such as based on keyword matching, semantic similarity analysis, or predefined mapping relationships. For example, if S51 identifies the possible attack action of using a webshell for command execution, the system searches the attack knowledge base for specific attack techniques related to keywords such as webshell and command execution. The knowledge base may return a series of more fine-grained attack actions, such as executing system commands through Webshell, creating new user accounts through Webshell, downloading malicious payloads through Webshell, etc. For each specific attack action that is matched, the system will extract all its relevant attribute information from the attack knowledge base, including but not limited to: a detailed description of the action, the probability of its successful execution, the probability of being detected by security devices or logs, and the resources or costs required for the attacker to implement the action. The final output includes all specific and quantifiable attack actions that the attacker may take for the current attack hypothesis, as well as a collection of attack actions with detailed attributes of each action.
[0059] Next, S53 is executed. To obtain more detailed and specific defense actions, a pre-built defense action library is required. This defense action library is a structured database containing detailed information on a large number of known network security defense technologies, tools, strategies, and best practices. Each entry in the knowledge base details a specific defense action, including its name, description, applicable attack type, expected defense effect (e.g., probability of attack prevention, degree of attack success reduction), resources or costs required to implement the action (e.g., manpower investment, time consumption, impact on business continuity, funding required), and possible side effects. This information is collected and maintained based on industry standards, security best practices, historical defense event analysis, and expert experience. During the extraction process, the system matches and queries each entry in the list of possible defense actions against the defense action library. This matching process can be implemented in various ways, such as based on keyword matching, semantic similarity analysis, or predefined mapping relationships. For example, if S51 identifies the possible defense action of isolating the Web-01 server, the system searches the defense action library for specific defense technologies related to keywords such as "isolate" and "server." The knowledge base may return a series of more fine-grained defensive actions, such as disconnecting the Web-01 server from the network, migrating the Web-01 server to a quarantine zone, or restricting Web-01 server's external access ports. For each specific defensive action matched, the system extracts all relevant attribute information from the defensive action library, including but not limited to: a detailed description of the action, the probability of its effectiveness in preventing or mitigating the attack, and the cost of implementing the action (e.g., resource consumption, risk of business interruption, etc.). The final output is a set of defensive actions. This set includes all specific, quantifiable defensive actions that a defender could take based on the current attack hypothesis, as well as the detailed attributes of each action.
[0060] Finally, S54 is executed. The goal of this step is to construct a benefit / loss matrix for the i-th attack hypothesis. The rows of this matrix represent each possible attacker action, and the columns represent each possible defender action. Each cell in the matrix contains the benefit or loss value when the attacker and defender select the corresponding action. In specific implementation, the system iterates through each attack action in the set of attack actions and combines it with each defense action in the set of defense actions. For each attack / defense action combination—for example, if the attacker selects attack action A and the defender selects defense action D—the system calculates the respective benefits or losses for the attacker and defender under this combination. This calculation process relies on preset rules and parameters. First, the probability of attack success under the combined effects of attack action A and defense action D must be evaluated. This probability is not generated out of thin air but is determined based on preset attack / defense effectiveness evaluation rules or mitigation matrices. For example, if the attack action is SQL injection and the defense action is patching the SQL injection vulnerability, the probability of attack success will be very low, perhaps preset to 0.05. If the defense action is deploying a web application firewall, its SQL injection protection effectiveness may be preset to moderate, and the probability of attack success may be 0.3. These attack and defense effectiveness evaluation rules and probability values are pre-set and trained based on historical attack data, security vulnerability database, expert experience and industry best practices.
[0061] After determining the probability of attack success, the gains or losses for both attackers and defenders can be calculated. For the defender, the gains or losses primarily consist of two components: asset loss and defense costs. If the attack succeeds, the defender will suffer a loss in the value of the relevant assets. This asset value is derived from the asset information in S2. For example, the value of a core database might be preset at 10,000 units. Furthermore, regardless of whether the attack succeeds, the defender will incur the cost of implementing the selected defense action D. This cost is drawn from the defense action library. For example, the cost of isolating a server might be preset at 200 units, and the cost of updating firewall rules might be preset at 50 units. Therefore, the defender's loss can be calculated as: probability of attack success × asset loss + cost of the defense action. If the attack fails, the defender only bears the cost of the defense action.
[0062] For the attacker, their gains or losses are primarily composed of the gains from a successful attack and the costs of a failed attack or detection. If the attack is successful, the attacker receives a gain corresponding to the value of the attacked asset. If the attack fails or is detected by the defender, the attacker incurs certain costs, such as resource consumption, the risk of being tracked, and the exposure of their tools. These costs can also be preset to specific values. For example, a failed attack could result in a loss of 100 units for the attacker. Therefore, the attacker's gains can be calculated as: probability of attack success × asset value gain - probability of attack failure × attack cost. In particular, the attacker's gains are negatively correlated with the defender's losses: the greater the defender's losses, the greater the attacker's gains. By performing this calculation for all combinations of attack and defense actions, a complete gain or loss matrix is ultimately obtained. This matrix, together with the attacker and defender's strategy sets, constitutes the i-th sub-game model.
[0063] In step S6, based on the posterior Bayesian belief network, the MCTS-based real-time defense strategy solution is performed on the i-th sub-game model to obtain a recommended defense strategy. It is worth mentioning that network attack and defense is a highly dynamic and uncertain confrontation process. The defender needs to select the optimal response from many possible defense actions within a limited time. Traditional decision-making methods are difficult to effectively cope with this complexity and real-time requirements. MCTS is an efficient search algorithm, which is particularly suitable for decision-making problems with huge state spaces that are difficult to perform exhaustive searches. By integrating the attack intention probability provided by the posterior Bayesian belief network into MCTS, the search process can be made more focused on high-risk attack intentions, so that under real-time requirements, a defense strategy that can maximize the defender's benefits or minimize losses can be quickly found.
[0064] In an optional embodiment of the present application, step S6, based on the posterior Bayesian belief network, performing a real-time defense strategy solution based on MCTS for the i-th sub-game model to obtain a recommended defense strategy, includes: S61, based on the posterior Bayesian belief network, performing a real-time defense strategy solution based on MCTS for the i-th sub-game model to obtain the benefit value of each defense action; S62, selecting the defense action with the highest benefit value as the recommended defense strategy.
[0065] Step S6 can be implemented as follows: First, execute S6. It's worth noting that MCTS (Monte Carlo Tree Search) is an iterative search process that evaluates the pros and cons of different strategies by constructing and exploring a game tree. The process begins with a root node, which represents the current attack-defense situation, i.e., the starting state described by the i-th attack hypothesis.
[0066] Each iteration of MCTS consists of four main phases: 1. Selection: Starting from the root node, the currently constructed game tree is traversed downward. At each node, a child node is selected for exploration. This selection is based on a strategy that balances exploring underexplored paths with exploiting the most promising paths. For example, an upper confidence interval algorithm can be used, which prioritizes child nodes with high average payoffs and relatively few visits to ensure that the entire policy space is fully explored. 2. Expansion: When the selection phase reaches a node that has not yet been fully expanded (i.e., it has child actions that have not yet been added to the tree), an unvisited defense action or attack action is selected from that node and added to the game tree as a new child node. This new node represents the new state after taking that action. For example, if the current node represents a state where the attacker has implanted a webshell, the defender can choose the defensive action of isolating the web server, thereby expanding a new node in the tree. 3. Simulation: Starting from the newly expanded node, a simulated game or random deduction is performed until a terminal state is reached, such as attack success, attack failure, or the defender successfully blocking the attack. During this simulation, the attacker and defender's subsequent actions can be chosen randomly or based on simple heuristics. The posterior Bayesian belief network plays a key role in this phase: it provides a probability distribution of the attacker's most likely attack intentions. When simulating the attacker's actions, these intention probabilities can be used to bias the selection of attack actions. For example, if the probability of data theft is high, the simulated attacker may attempt actions related to data exfiltration more frequently. Simultaneously, the payoff or loss matrix defined in S54 is used to evaluate the final outcome of the simulated game. For example, if the simulation ultimately results in data theft, the defender calculates the loss based on the value of the stolen assets and the cost of the defensive actions taken during the simulation, while the attacker calculates the corresponding payoff. 4. Backpropagation: The outcome of the simulated game—the defender's payoff or loss—is propagated back from the final state to the root node of the game tree. During backpropagation, statistics for all traversed nodes (including total payoff and number of visits) are updated. This process enables MCTS to learn which paths are good and which are bad, thereby favoring high-payoff paths in subsequent iterations.
[0067] The four stages described above are repeated for a preset number of iterations, for example, 10,000. Setting the number of iterations requires a balance between computing resources and solution accuracy. In scenarios with high real-time requirements, the number of iterations is dynamically adjusted based on available time. After a sufficient number of iterations, MCTS constructs a game tree containing a large number of simulated paths, where each defensive action (a direct child of the root node) accumulates a wealth of payoff statistics. Ultimately, the output is the average payoff value for each possible defensive action for the i-th attack hypothesis. These payoff values reflect the average potential reward for the defender after taking that defensive action across a large number of simulated games.
[0068] Next, proceed to step S62. After calculating the payoffs for all possible defensive actions in S61, these values are simply compared and the largest one is selected. For example, if the payoff for defensive action A is -500 (indicating an average loss of 500 units), the payoff for defensive action B is -200, and the payoff for defensive action C is -800, then defensive action B will be selected as the recommended defense strategy because it represents the minimum loss the defender can tolerate under the current attack assumptions. This recommended defense strategy will serve as the final decision output, guiding the defender's actual actions.
[0069] In step S7, the recommended defense strategy is translated into executable instructions and executed. That is, while S6 provides optimal defense strategies, such as isolating the Web-01 server or clearing Webshell files, these remain high-level descriptions. To achieve effective protection in a real-world network environment, these strategies must be translated into precise, device-specific configuration commands or steps. This step is the key link between intelligent decision-making and actual action, ensuring a closed-loop defense process. This allows the system to respond to evolving attack threats in real time and automatically, thereby achieving the highly effective protection goals pursued by the patent.
[0070] Optionally, step S7 can be specifically implemented in the following manner: First, policy translation and instruction generation are performed. This requires a policy translation module, which is responsible for mapping high-level recommended defense policies to specific, low-level, device or tool-specific operational instructions. The core of this module is a pre-built policy-instruction mapping knowledge base. This knowledge base contains a large number of predefined mapping relationships, which associate abstract defense actions with specific commands, API calls or configuration scripts for different types of network devices such as firewalls, switches, security tools such as intrusion prevention systems, terminal detection and response tools, security information and event management platforms, and operating systems such as Linux and Windows. For example, if the recommended defense policy is to isolate the Web-01 server, and the constructed dynamic network map shows that the IP address of the Web-01 server is 192.168.1.10, and its network traffic is controlled by a Cisco ASA firewall, then the policy translation module will search the knowledge base for the corresponding relationship between the isolated server and the Cisco ASA firewall. The knowledge base may contain instruction templates. The module will fill the IP address of Web-01 (192.168.1.10) into the instruction template to generate specific firewall commands, such as access control list configuration commands for denying specific IP traffic.
[0071] Secondly, execute the instructions. After the executable instructions are generated, the next step is to deliver them to the target device or tool for actual execution. This can be achieved through a variety of automation interfaces. For modern security products, software-defined network (SDN) controllers, or cloud security services (such as cloud firewalls and security groups) that support APIs, the instruction generation module will directly call the corresponding API interface to programmatically issue configuration or operation instructions. This is the preferred way to achieve a high degree of automation. For traditional network devices and servers, you can log in to the device through SSH or Telnet protocols in combination with automation tools (such as Ansible and Python scripts) and execute the generated command line instructions. For policies that require host-level operations (such as file deletion, process termination, and registry modification), the instructions will be sent to the security agent deployed on the target host (such as EDR agent) for local execution by the agent. During the execution process, the security of the operation needs to be ensured. For example, the credentials for accessing network devices and servers must be securely stored and managed, and the principle of least privilege must be followed.
[0072] Finally, to ensure the effectiveness of the strategy and the robustness of the system, an execution confirmation and feedback mechanism can be introduced after the instruction is executed. This includes checking the execution results returned by the target device or tool to confirm whether the instruction was successfully issued and took effect. For example, checking whether the firewall's operating configuration has been updated or whether the file has been deleted. At the same time, network traffic, logs, or asset status can be further monitored to verify whether the defense strategy has achieved the desired effect, for example, whether malicious traffic has been blocked and suspicious activities have stopped. Furthermore, changes in network status after the defense strategy is executed (for example, a server is isolated and its connection status changes) are fed back to the dynamic network map to ensure that the map always reflects the latest network situation.
[0073] In summary, the network protection method based on the attack and defense game model of the embodiment of the present application is explained, which aims to solve the problem of how to make efficient network protection decisions in the context of incomplete information, large network scale and real-time requirements. First, by integrating raw logs, asset information and threat intelligence, a dynamic network map is constructed to associate fragmented information and initially alleviate information incompleteness. Then, the trained graph neural network model is used to identify and predict potential Top-K attack hypotheses from the dynamic map, which is equivalent to quickly focusing on possible threat paths in massive data and effectively coping with the complexity of large-scale networks. Subsequently, these attack hypotheses are used as new evidence to update the Bayesian belief network, further improving the understanding of the attacker's intentions and capabilities. The most critical thing is that for each attack hypothesis, a sub-game model is generated, and based on the updated Bayesian belief network, MCTS is used to solve the real-time defense strategy. The introduction of MCTS enables efficient exploration and recommendation of the optimal defense strategy under incomplete information and complex decision space, avoiding the infeasibility of exhaustive calculation, thereby realizing real-time and efficient defense decision-making for large-scale networks.
[0074] Figure 4 FIG is a block diagram of a network protection system based on an attack-defense game model according to an embodiment of the present application. Figure 4As shown, according to the embodiment of the present application, the network protection system 100 based on the attack and defense game model includes: a data acquisition module 110 for acquiring original logs, asset information and threat intelligence; a dynamic network graph construction module 120 for constructing a dynamic network graph based on original logs, asset information and threat intelligence; a dynamic network graph encoding module 130 for inputting the dynamic network graph into a trained graph neural network model to obtain a Top-K attack hypothesis list; a network update module 140 for using each attack hypothesis in the Top-K attack hypothesis list as new evidence to update the prior Bayesian belief. The network is updated to obtain a posterior Bayesian belief network; a sub-game model generation module 150 is used to generate a sub-game model for the i-th attack hypothesis in the Top-K attack hypothesis list based on the attacker strategy space, the defender strategy space and the profit function to obtain the i-th sub-game model; a defense strategy recommendation module 160 is used to solve the i-th sub-game model based on the posterior Bayesian belief network in real-time defense strategy to obtain a recommended defense strategy based on MCTS; a strategy translation execution module 170 is used to translate the recommended defense strategy into executable instructions and perform strategy execution.
[0075] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned network protection system based on the attack and defense game model have been referred to above. Figures 1 to 3 The network protection method based on the attack-defense game model has been introduced in detail, and therefore, its repeated description will be omitted.
Claims
1. A network protection method based on an attack-defense game model, characterized in that: include: Obtain original logs, asset information, and threat intelligence; build a dynamic network map based on original logs, asset information, and threat intelligence; Inputting the dynamic network graph into a trained graph neural network model to obtain a Top-K attack hypothesis list; using each attack hypothesis in the Top-K attack hypothesis list as new evidence, updating the prior Bayesian belief network to obtain a posterior Bayesian belief network; generating a sub-game model for the i-th attack hypothesis in the Top-K attack hypothesis list based on the attacker's strategy space, the defender's strategy space, and the payoff function to obtain the i-th sub-game model; solving the i-th sub-game model based on the MCTS real-time defense strategy based on the posterior Bayesian belief network to obtain a recommended defense strategy; translating the recommended defense strategy into executable instructions and executing the strategy; Among them, the dynamic network graph is input into a trained graph neural network model to obtain a Top-K attack hypothesis list, including: passing the nodes and edges in the dynamic network graph through an embedding layer to obtain a dynamic network embedding coding graph; inputting the dynamic network embedding coding graph into the trained graph neural network model to obtain a set of node final vectors; extracting the node final vector corresponding to the suspicious activity from the set of node final vectors as the alarm node final vector; performing link prediction on the alarm node final vector to obtain a link prediction result; performing node classification on all node final vectors in the set of node final vectors to obtain a node classification result; and determining the Top-K attack hypothesis list based on the link prediction result and the node classification result.
2. The network protection method based on the attack-defense game model according to claim 1 is characterized in that: Based on the original logs, asset information and threat intelligence, a dynamic network map is constructed, including: parsing, standardizing and correlation analyzing the original logs to identify suspicious activities from the original logs; combining asset information and threat intelligence to construct a dynamic map; marking the suspicious activities on the corresponding nodes or edges of the dynamic map to obtain the dynamic network map.
3. The network protection method based on the attack-defense game model according to claim 1 is characterized in that: Performing link prediction on the final vector of the alarm node to obtain a link prediction result, including: extracting a node final vector of a neighboring node of the final vector of the alarm node as the neighboring node final vector; performing association coding on the final vector of the alarm node and the final vector of the neighboring node based on a fully connected layer, and inputting the association coding result into a classifier to obtain the link prediction result, wherein the link prediction result is used to indicate the possibility that the path from the alarm node to the neighboring node is part of the attack path.
4. The network protection method based on the attack-defense game model according to claim 3 is characterized in that: Performing node classification on all node final vectors in the set of node final vectors to obtain a node classification result, including: inputting each node final vector in the set of node final vectors into a classifier to obtain the node classification result, wherein the node classification result is used to represent the probability that each node is the final attack target.
5. The network protection method based on the attack-defense game model according to claim 1 is characterized in that: The dynamic network graph is input into a trained graph neural network model to obtain a Top-K attack hypothesis list, including: passing the nodes and edges in the dynamic network graph through an embedding layer to obtain a dynamic network embedding coding graph; inputting the dynamic network embedding coding graph into the trained graph neural network model to obtain a set of node final vectors; performing local association-global association balance on the set of node final vectors to obtain a set of optimized node final vectors; extracting the optimized node final vector corresponding to the suspicious activity from the set of optimized node final vectors as the alarm node final vector; performing link prediction on the alarm node final vector to obtain a link prediction result; performing node classification on all optimized node final vectors in the set of optimized node final vectors to obtain a node classification result; and determining the Top-K attack hypothesis list based on the link prediction result and the node classification result.
6. The network protection method based on the attack-defense game model according to claim 5 is characterized in that: The method comprises the steps of: performing a local-global association balance on the set of node final vectors to obtain a set of optimized node final vectors, comprising: performing a two-dimensional arrangement on the set of node final vectors to obtain a node topology matrix; performing a local-global low-dimensional scale calculation on the set of node final vectors and the node topology matrix to obtain node local-global low-dimensional scale latent variable parameters; calculating the information latent association strength between any two node final vectors in the set of node final vectors based on the node local-global low-dimensional scale latent variable parameters to obtain a node information latent association matrix; and using the node information latent association matrix as an association consistency metric space, mapping the set of node final vectors to obtain the set of optimized node final vectors.
7. The network protection method based on the attack-defense game model according to claim 1 is characterized in that: Based on the attacker's strategy space, the defender's strategy space, and the profit function, a sub-game model is generated for the i-th attack hypothesis in the Top-K attack hypothesis list to obtain the i-th sub-game model, including: for the i-th attack hypothesis in the Top-K attack hypothesis list, the following operations are performed: relevant nodes, possible attack actions, and possible defense actions are extracted from the i-th attack hypothesis; attack actions corresponding to the possible attack actions are extracted from the attack knowledge base to obtain a set of attack actions; defense actions corresponding to the possible defense actions are extracted from the defense action library to obtain a set of defense actions; based on the asset information and the cost of each defense action in the defense action library, the profit / loss value between the set of attack actions and any pair of attack actions and defense actions in the set of defense actions is calculated to obtain the i-th sub-game model.
8. The network protection method based on the attack-defense game model according to claim 7 is characterized in that: Based on the posterior Bayesian belief network, the MCTS-based real-time defense strategy is solved for the i-th sub-game model to obtain a recommended defense strategy, including: based on the posterior Bayesian belief network, the MCTS-based real-time defense strategy is solved for the i-th sub-game model to obtain the benefit value of each defense action; and the defense action with the highest benefit value is selected as the recommended defense strategy.
9. A network protection system based on an attack-defense game model, characterized in that: include: Data acquisition module, used to obtain raw logs, asset information and threat intelligence; A dynamic network graph construction module is used to construct a dynamic network graph based on raw logs, asset information, and threat intelligence; a dynamic network graph encoding module is used to input the dynamic network graph into a trained graph neural network model to obtain a Top-K attack hypothesis list; a network updating module for updating the priori Bayesian belief network using each attack hypothesis in the Top-K attack hypothesis list as new evidence to obtain a posterior Bayesian belief network; a sub-game model generating module for generating a sub-game model for the i-th attack hypothesis in the Top-K attack hypothesis list based on the attacker's strategy space, the defender's strategy space, and the payoff function to obtain an i-th sub-game model; and a defense strategy recommendation module for performing a real-time defense strategy solution based on MCTS for the i-th sub-game model based on the posterior Bayesian belief network to obtain a recommended defense strategy; A strategy translation and execution module, configured to translate the recommended defense strategy into executable instructions and execute the strategy; Among them, the dynamic network graph is input into a trained graph neural network model to obtain a Top-K attack hypothesis list, including: passing the nodes and edges in the dynamic network graph through an embedding layer to obtain a dynamic network embedding coding graph; inputting the dynamic network embedding coding graph into the trained graph neural network model to obtain a set of node final vectors; extracting the node final vector corresponding to the suspicious activity from the set of node final vectors as the alarm node final vector; performing link prediction on the alarm node final vector to obtain a link prediction result; performing node classification on all node final vectors in the set of node final vectors to obtain a node classification result; and determining the Top-K attack hypothesis list based on the link prediction result and the node classification result.