Root cause analysis method, system and equipment based on large model causal reasoning
By adopting the root cause analysis method based on big model causal reasoning in the field of network security, the causal graph is constructed and optimized, and combined with Bayesian inference, the problem of insufficient causal dependency capture in the existing technology is solved, and the positioning accuracy of attack entry nodes is improved.
Patent Information
- Application Number
- CN202510495461.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When handling complex attack event chains, it is difficult for the prior art to accurately capture causal dependencies, resulting in insufficient positioning accuracy of attack entry nodes, especially when facing security events related to timing.
The root cause analysis method based on big model causal reasoning is adopted, and the second causal graph is obtained by extracting network entities and attack event chains from network traffic logs, the first causal graph is constructed, and the pre-trained large language model is used to optimize it to obtain the second causal graph. Then, based on Bayesian inference, the root cause probability of each node is calculated and the attack entry node with the highest outlier is located.
It improves the positioning accuracy of the attack entry node, can dynamically learn the dependencies of unknown entities, reason about unseen attack patterns, and does not need to update the rule base frequently, which reduces the false positive rate.
Smart Images

Figure CN120017426A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security, and in particular to a root cause analysis method, system and device based on large model causal reasoning. Background Art
[0002] Network traffic monitoring and security analysis are important components of modern network security. As network attack methods become increasingly complex and diverse, accurately identifying and locating attack sources has become a key link in ensuring network security. Through in-depth analysis of network traffic logs, potential security threats can be effectively discovered and protective measures can be taken in a timely manner, thereby protecting the stable operation of network systems and data security.
[0003] In the prior art, in order to effectively trace the source of network attacks, a variety of methods are usually used for analysis. For example, the rule-matching method detects abnormal traffic through preset security rules; the machine learning-based method uses classification algorithms to identify known attack patterns; the graph-based method analyzes the attack path by constructing a relationship graph between network entities. In addition, there are statistical analysis-based methods that determine abnormal behavior by calculating the probability distribution of traffic characteristics. These methods provide diverse solutions for tracing the source of network attacks from different perspectives.
[0004] However, the above methods have limitations when dealing with complex attack event chains, especially when facing time-related security events. It is difficult to accurately capture causal dependencies, resulting in insufficient positioning accuracy of attack entry nodes. Summary of the invention
[0005] In order to solve the above problems in the prior art, the present invention proposes a root cause analysis method, system and device based on large model causal reasoning, which improves the positioning accuracy of the attack entry node.
[0006] In a first aspect of the present invention, a root cause analysis method based on large model causal reasoning is proposed, the method comprising: Extracting network entities and attack event chains from network traffic logs; the attack event chains include security events with time-series correlation; According to the network entity and the attack event chain, a first causal graph is constructed based on a pre-trained graph neural network; wherein the nodes of the first causal graph represent the network entity or the security event, and the edge weights of the first causal graph represent the causal dependency strength dynamically learned through an attention mechanism; Using a pre-trained large language model to perform logical reasoning on the first causal graph, and then optimizing the first causal graph to obtain a second causal graph; The root cause probability of each node in the second causal graph is calculated based on Bayesian inference, and the attack entry node with the highest outlier value is located.
[0007] The present invention uses a pre-trained large language model to optimize the first causal graph, making the second causal graph used for positioning more accurate; the attention mechanism of the graph neural network can dynamically learn the dependencies of unknown entities; the generalization ability of the large language model can be used to infer unseen attack patterns (such as new vulnerability exploit chains), and emerging threats can be detected without frequently updating the rule base; the Bayesian network is used to locate the attack entry to reduce the false alarm rate. Therefore, the present invention improves the positioning accuracy of the attack entry node by combining the large language model, the graph neural network and the Bayesian network.
[0008] Preferably, the step of "extracting network entities and attack event chains from network traffic logs" includes: Extracting structured metadata fields from the network traffic log; the metadata fields include: timing features, network layer features, and transmission statistics; Standardizing the continuous metadata fields and embedding the discrete metadata fields to obtain preprocessed metadata; extracting a load feature matrix from the network traffic log; Input the preprocessed metadata and the load feature matrix into the first channel and the second channel of the dual-channel LSTM model respectively, capture the metadata time series pattern and the load semantic features, and output the fused time series features; Input the fused temporal features into the Transformer decoder to capture the dependencies between events, and output the confidence probability of each event chain type; The attack event chain is identified according to the confidence probability of each event chain type and a preset probability threshold.
[0009] Preferably, the step of “constructing a first causal graph based on a pre-trained graph neural network according to the network entity and the attack event chain” includes: Extracting the features of the network entity and the security event respectively, thereby constructing a node feature matrix; Constructing an adjacency matrix according to the communication relationship between the network entities, the temporal dependency between the security events in the attack event chain, and the triggering relationship between the network entities and the security events; The node feature matrix and the adjacency matrix are input into a pre-trained graph neural network to construct the first causal graph.
[0010] Preferably, the step of "using a pre-trained large language model to perform logical reasoning on the first causal graph, and then optimizing the first causal graph to obtain a second causal graph" includes: Converting the first causal graph into a text sequence; Generate prompt words based on domain knowledge rule base; Inputting the text sequence and the prompt word into the pre-trained large language model, and outputting an inference result; According to the inference result, the edges violating logic are deleted, the unreasonable edge weights are adjusted and / or the missing edges are supplemented in the first causal graph, so as to obtain an optimized second causal graph.
[0011] The present invention eliminates causal relationships that violate domain common sense based on the inference results, making the attack chain reasoning more consistent with actual security logic; adjusts edge weights (such as reducing the confidence of false alarm events and increasing the weight of real attack events), supplements potential attack paths missed by GNN (such as covert lateral movement), and is conducive to discovering complex attack chains ignored by traditional methods (such as multi-hop penetration in APT attacks).
[0012] Preferably, the step of "calculating the root cause probability of each node in the second causal graph based on Bayesian inference and locating the attack entry node with the highest outlier value" includes: Define a priori probability for each network entity node based on historical attack data, indicating the probability of normal behavior of the node in the absence of an attack; Define a priori probability for each security event node based on historical security event data, indicating the probability of the node occurring naturally without any associated attack; Constructing a Bayesian network according to the second causal graph, the prior probability of the network entity node and the prior probability of the security event node; Observing abnormal behavior from the network traffic log, and calculating the posterior probability of each node using the Bayesian network based on the abnormal behavior; The posterior probabilities of the network entity nodes are sorted, and the network entity node corresponding to the highest probability value is obtained, and is positioned as the attack entry node.
[0013] The present invention uses Bayesian networks to fuse prior probabilities with real-time abnormal behaviors and calculate posterior probabilities, which can resist attackers' obfuscation methods (such as traffic camouflage) and reduce false alarm rates, thereby further improving positioning accuracy. Example: Even if an attacker forges a large amount of "port scan" noise, the Bayesian network can still locate the real "Web vulnerability exploitation" entrance through logical consistency.
[0014] Preferably, the method further comprises: The posterior probabilities of the security event nodes are sorted to identify the most likely attack methods, analyze the attacker's behavior patterns, and assist in verifying the attack entry location results of the network entity nodes.
[0015] A second aspect of the present invention provides a root cause analysis system based on large model causal reasoning, the system comprising: An extraction module, used to extract network entities and attack event chains from network traffic logs; the attack event chains include security events with time-series correlation; A causal graph construction module, configured to construct a first causal graph based on a pre-trained graph neural network according to the network entity and the attack event chain; wherein the nodes of the first causal graph represent the network entity or the security event, and the edge weights of the first causal graph represent the causal dependency strength dynamically learned through an attention mechanism; A causal graph optimization module, configured to use a pre-trained large language model to perform logical reasoning on the first causal graph, and then optimize the first causal graph to obtain a second causal graph; The entry location module is used to calculate the root cause probability of each node in the second causal graph based on Bayesian inference, and locate the attack entry node with the highest outlier value.
[0016] According to a third aspect of the present invention, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute the method described above.
[0017] According to a fourth aspect of the present invention, a computer-readable storage device is provided, storing a computer program that can be loaded by a processor and execute the method as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the main steps of an embodiment of the root cause analysis method based on large model causal reasoning in the present invention. DETAILED DESCRIPTION
[0019] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] It should be noted that, in the description of the present invention, the terms "first" and "second" are only for the convenience of description, and do not indicate or imply the relative importance of the devices, elements or parameters, and therefore cannot be understood as limiting the present invention. In addition, the term "and / or" in the present invention is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally indicates that the associated objects before and after are in an "or" relationship.
[0022] Figure 1 This is a schematic diagram of the main steps of the root cause analysis method embodiment based on large model causal reasoning in the present invention. Figure 1 As shown, the root cause analysis method of this embodiment includes: Step S10: extracting network entities and attack event chains from network traffic logs.
[0023] Among them, network entities include devices, IP addresses, ports, etc.; security events include: vulnerability exploitation, privilege escalation, data leakage, etc.; the attack event chain is an attack process composed of multiple time-series related security events, which usually reflects the attacker's strategy of achieving the goal step by step, such as penetration, lateral movement, and data theft.
[0024] Specifically, this step may include steps S11-S16: Step S11: extracting structured metadata fields from the network traffic log.
[0025] The metadata fields include: timing characteristics, network layer characteristics and transmission statistics.
[0026] Timing features include: timestamp, packet interval, etc.; network layer features include: source / destination IP address, port number, protocol type, etc.; transmission statistics include: packet length, traffic rate, TCP flag bit combination, etc.
[0027] Step S12: Standardize the continuous metadata fields and embed the discrete metadata fields to obtain preprocessed metadata.
[0028] Step S13: extracting a load feature matrix from the network traffic log.
[0029] Step S14: input the preprocessed metadata and payload feature matrix into the first channel and the second channel of the dual-channel LSTM model respectively, capture the metadata temporal pattern and payload semantic features, and output the fused temporal features.
[0030] In this step, the first channel of LSTM is used to capture metadata temporal patterns (such as the progressiveness of port scanning), and the second channel is used to parse payload semantic features (such as malicious keywords in the attack payload). The final hidden states of the two channels are then concatenated to generate a fused temporal feature vector through a fully connected layer.
[0031] Step S15: Input the fused temporal features into the Transformer decoder to capture the dependencies between events and output the confidence probability of each event chain type.
[0032] In this step, the fused time series features are input into the Transformer decoder layer, and the self-attention mechanism is used to capture the long-range dependencies between events (such as the association between DDoS attacks and subsequent data leakage), and then the confidence probability of each event chain type is calculated through the Softmax layer. For example, port scanning → vulnerability exploitation (probability 78%), brute force cracking → privilege escalation (probability 65%).
[0033] Step S16: Identify the attack event chain according to the confidence probability of each event chain type and a preset probability threshold.
[0034] Step S20: construct a first causal graph based on the pre-trained graph neural network according to the network entities and the attack event chain.
[0035] Among them, each node of the first causal graph represents a network entity or a security event, and the edge weight of the first causal graph represents the causal dependency strength dynamically learned through the attention mechanism.
[0036] Specifically, this step may include steps S21-S23: Step S21: extract the features of network entities and security events respectively, so as to construct a node feature matrix.
[0037] The structure of the node feature matrix: Row: All network entities + security events (N nodes in total); Column: feature dimension (D dimension, entity and event features need to be aligned).
[0038] Alignment method: fill missing fields with zeros (for example, an event has no "device type" field) and perform Min-Max normalization on heterogeneous features.
[0039] Step S22: construct an adjacency matrix according to the communication relationship between network entities, the temporal dependency between security events in the attack event chain, and the triggering relationship between network entities and security events.
[0040] For example, if there is bidirectional traffic between two network entities in the past 5 minutes, an undirected edge is established, and the edge weight is initialized with the ratio of traffic bytes; if event A occurs earlier than event B and the time difference is less than a threshold (such as 2 minutes), a directed edge A→B is established, and the edge weight is initialized with the inverse of the time difference; if the entity IP appears in the "source IP" field of the event log, a directed edge is established, and the edge weight is initialized with a fixed value of 1.0.
[0041] Step S23: input the node feature matrix and the adjacency matrix into the pre-trained graph neural network to construct a first causal graph.
[0042] The graph neural network can choose GAT (Graph Attention Network), which can automatically learn the importance of edge weights through the attention mechanism. In GAT, the attention coefficient is calculated for each edge, and the weight in the adjacency matrix is fused with the attention coefficient.
[0043] Step S30: Use the pre-trained large language model to perform logical reasoning on the first causal graph, and then optimize the first causal graph to obtain a second causal graph.
[0044] Specifically, this step may include steps S31-S34: Step S31: convert the first causal graph into a text sequence.
[0045] The converted text sequence is as follows: [Node] Host A (Type: Web Server) → [Edge] Trigger (Weight: 0.8) → [Node] SQL Injection (Type: Security Event) [Node] User B (type: terminal) → [Edge] Login (weight: 0.6) → [Node] Host A.
[0046] Step S32: Generate prompt words based on the domain knowledge rule base.
[0047] In this embodiment, the MITRE ATT&CK framework is used as the domain knowledge rule base, which includes: logical constraints (such as "vulnerability exploitation must occur after port scanning") and permission constraints (such as "ordinary users cannot directly trigger firewall configuration changes").
[0048] In this step, the domain knowledge rules are input into LLM as part of the prompt, for example: Please check whether the following causal diagram is reasonable according to the rules: 1. SQL injection must be triggered by a Web request and cannot be initiated directly by the end user; 2. Firewall denial events must precede lateral movement within the intranet.
[0049] Output conflicting edges and suggested fixes.
[0050] Step S33: input the text sequence and the prompt word into the pre-trained large language model, and output the inference result.
[0051] The inference result can be in JSON format, for example: { "conflicts": [ { "edge": "User B → Host A", "reason": "End users do not have the authority to directly trigger SQL injection", "action": "delete" } ], "missing_edges": [ { "from": "Host A", "to": "Database", "reason": "SQL injection usually accompanies database access", "suggested_weight": 0.9 } ] } Step S34: According to the inference result, in the first causal graph, the edges that violate logic are deleted, the unreasonable edge weights are adjusted, and / or the missing edges are supplemented, so as to obtain an optimized second causal graph.
[0052] For example, with respect to the inference result of step S33 above, the edge "user B → host A" in the "conflicts" list may be removed, or the missing edge "host A → database" may be added according to the "missing_edges" list.
[0053] Step S40: Calculate the root cause probability of each node in the second causal graph based on Bayesian inference, and locate the attack entry node with the highest outlier value.
[0054] Specifically, this step may include steps S41-S45: Step S41: define a priori probability for each network entity node according to historical attack data, indicating the probability of normal behavior of the node in the absence of an attack.
[0055] Step S42: define a priori probability for each security event node based on historical security event data, indicating the probability of natural occurrence of the node in the absence of associated attacks.
[0056] Step S43: construct a Bayesian network according to the second causal graph, the prior probabilities of network entity nodes and the prior probabilities of security event nodes.
[0057] Step S44: observe abnormal behavior from the network traffic log, and calculate the posterior probability of each node using the Bayesian network based on the abnormal behavior.
[0058] Step S45: sort the posterior probabilities of the network entity nodes, obtain the network entity node corresponding to the highest probability value, and locate it as the attack entry node.
[0059] In an optional embodiment, the root cause analysis method may further include: Step S46: sort the posterior probabilities of the security event nodes to identify the most likely attack methods, analyze the attacker's behavior patterns, and assist in verifying the attack entry location results of the network entity nodes.
[0060] Although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art can understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of the present invention.
[0061] Based on the above method embodiment, the present invention also provides a root cause analysis system embodiment based on large model causal reasoning. The system of this embodiment includes: an extraction module, a causal graph construction module, a causal graph optimization module and an entry location module.
[0062] Among them, the extraction module is used to extract network entities and attack event chains from network traffic logs, and the attack event chains include time-series-related security events; the causal graph construction module is used to construct a first causal graph based on the network entities and attack event chains based on the pre-trained graph neural network; wherein the nodes of the first causal graph represent network entities or security events, and the edge weights of the first causal graph represent the causal dependency strength dynamically learned through the attention mechanism; the causal graph optimization module is used to use the pre-trained large language model to perform logical reasoning on the first causal graph, and then optimize the first causal graph to obtain a second causal graph; the entry positioning module is used to calculate the root cause probability of each node in the second causal graph based on Bayesian inference, and locate the attack entry node with the highest outlier value.
[0063] Furthermore, the present invention also provides an embodiment of an electronic device. The electronic device of this embodiment includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute the method described above.
[0064] Furthermore, the present invention also provides an embodiment of a computer-readable storage device, in which the storage device of this embodiment stores a computer program that can be loaded by a processor and execute the method described above.
[0065] The computer-readable storage device may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0066] Those skilled in the art should be able to appreciate that the method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0067] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A root cause analysis method based on large model causal reasoning, characterized in that: The method comprises: Extracting network entities and attack event chains from network traffic logs; the attack event chains include security events with time-series correlation; According to the network entity and the attack event chain, a first causal graph is constructed based on a pre-trained graph neural network; wherein each node of the first causal graph represents the network entity or the security event, and the edge weight of the first causal graph represents the causal dependency strength dynamically learned through an attention mechanism; Using a pre-trained large language model to perform logical reasoning on the first causal graph, and then optimizing the first causal graph to obtain a second causal graph; The root cause probability of each node in the second causal graph is calculated based on Bayesian inference, and the attack entry node with the highest outlier value is located.
2. The root cause analysis method based on large model causal reasoning according to claim 1 is characterized in that: The steps for "Extracting network entities and attack event chains from network traffic logs" include: Extracting structured metadata fields from the network traffic log; the metadata fields include: timing features, network layer features, and transmission statistics; Standardizing the continuous metadata fields and embedding the discrete metadata fields to obtain preprocessed metadata; extracting a load feature matrix from the network traffic log; Input the preprocessed metadata and the load feature matrix into the first channel and the second channel of the dual-channel LSTM model respectively, capture the metadata time series pattern and the load semantic features, and output the fused time series features; Input the fused temporal features into the Transformer decoder to capture the dependencies between events, and output the confidence probability of each event chain type; The attack event chain is identified according to the confidence probability of each event chain type and a preset probability threshold.
3. The root cause analysis method based on large model causal reasoning according to claim 1 is characterized in that: The step of "constructing a first causal graph based on a pre-trained graph neural network according to the network entity and the attack event chain" includes: Extracting the features of the network entity and the security event respectively, thereby constructing a node feature matrix; Constructing an adjacency matrix according to the communication relationship between the network entities, the temporal dependency between the security events in the attack event chain, and the triggering relationship between the network entities and the security events; The node feature matrix and the adjacency matrix are input into a pre-trained graph neural network to construct the first causal graph.
4. The root cause analysis method based on large model causal reasoning according to claim 1 is characterized in that: The step of "using a pre-trained large language model to perform logical reasoning on the first causal graph, and then optimizing the first causal graph to obtain a second causal graph" includes: Converting the first causal graph into a text sequence; Generate prompt words based on domain knowledge rule base; Inputting the text sequence and the prompt word into the pre-trained large language model, and outputting an inference result; According to the inference result, the edges violating logic are deleted, the unreasonable edge weights are adjusted and / or the missing edges are supplemented in the first causal graph, so as to obtain an optimized second causal graph.
5. The root cause analysis method based on large model causal reasoning according to claim 1 is characterized in that: The step of "calculating the root cause probability of each node in the second causal graph based on Bayesian inference and locating the attack entry node with the highest outlier value" includes: Define a priori probability for each network entity node based on historical attack data, indicating the probability of normal behavior of the node in the absence of an attack; Define a priori probability for each security event node based on historical security event data, indicating the probability of the node occurring naturally without any associated attack; Constructing a Bayesian network according to the second causal graph, the prior probability of the network entity node and the prior probability of the security event node; Observing abnormal behavior from the network traffic log, and calculating the posterior probability of each node using the Bayesian network based on the abnormal behavior; The posterior probabilities of the network entity nodes are sorted, and the network entity node corresponding to the highest probability value is obtained, and is positioned as the attack entry node.
6. The root cause analysis method based on large model causal reasoning according to claim 5 is characterized in that: The method further comprises: The posterior probabilities of the security event nodes are sorted to identify the most likely attack methods, analyze the attacker's behavior patterns, and assist in verifying the attack entry location results of the network entity nodes.
7. A root cause analysis system based on large model causal reasoning, characterized in that: The system comprises: An extraction module, used to extract network entities and attack event chains from network traffic logs; the attack event chains include security events with time-series correlation; A causal graph construction module, configured to construct a first causal graph based on a pre-trained graph neural network according to the network entity and the attack event chain; wherein the nodes of the first causal graph represent the network entity or the security event, and the edge weights of the first causal graph represent the causal dependency strength dynamically learned through an attention mechanism; A causal graph optimization module, configured to use a pre-trained large language model to perform logical reasoning on the first causal graph, and then optimize the first causal graph to obtain a second causal graph; The entry location module is used to calculate the root cause probability of each node in the second causal graph based on Bayesian inference, and locate the attack entry node with the highest outlier value.
8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute the method according to any one of claims 1 to 6.
9. A computer-readable storage device, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Host intrusion security detection method, system and device based on graph neural network
CN113225331A
Space thin-wall part spinning quality diagnosis method based on causal analysis
CN116502958A
Fault root cause analysis method and device and network equipment
CN116866149A
Emotion recognition method and device based on SoftmaxOne agency attention mechanism and medium
CN118964609A
Attack detection method and device, attack detection equipment and storage medium
CN119577744A
Cited By
Network attack attribution analysis and responsibility determination method based on causal reasoning
CN121012699A
Event aggregation and attribution method based on attack chain homologous analysis
CN121283742A
Website abnormal behavior detection method and device based on multi-source data fusion and medium
CN121664538A
A method, device and medium for detecting abnormal website behavior through multi-source data fusion
CN121664538B