A network event identification method, electronic equipment, storage medium and program product
Patent Information
- Application Number
- CN202610144228.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-02-02
Smart Images

Figure CN121841829B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and more specifically, to a method for identifying network events, an electronic device, a storage medium, and a program product. Background Technology
[0002] As the complexity and frequency of cyberattacks continue to increase, traditional security measures based on single alerts are no longer sufficient to meet the demands of modern cybersecurity. In this context, security operations platforms need an intelligent engine capable of automatically analyzing alert information in network events, transforming it into higher-level "security events," and interpreting these events. Currently, manual annotation is used to label the attack stages within network events. However, if a network event contains a complex attack chain, it cannot be identified, resulting in low accuracy and efficiency in identifying cyberattacks within network events. Summary of the Invention
[0003] The purpose of some embodiments of this application is to provide a method, electronic device, storage medium, and program product for identifying network events. Through the technical solutions of the embodiments of this application, a network event to be identified is obtained, wherein the network event to be identified includes at least a source address, a destination address, an event occurrence time, and an alarm type; based on the graph of the network event to be identified, related event information corresponding to the network event to be identified is determined using graph computation; the related event information is processed for confidence level, attack stage, and attack method to obtain the analysis result of the network event, wherein the analysis result of the network event includes at least a target attack stage, a target attack method, and a target confidence level; based on a large language model, target prompt words, and the analysis result of the network event, a natural language report corresponding to the network event to be identified is generated. In the embodiments of this application, by obtaining network events from multiple sources, automatic clustering of security alarms is achieved based on graph computation and multi-source evidence fusion to obtain related event information, and the related event information is processed for confidence level, attack stage, and attack method to obtain the analysis result of the network event. Combined with a large language model, a network event interpretation report corresponding to the analysis result is generated, significantly improving the automation level and analysis efficiency of attack methods in network events.
[0004] Firstly, some embodiments of this application provide a method for identifying network events, including: Obtain network events to be identified, wherein the network events to be identified include at least the source address, destination address, event occurrence time, and alarm type; Based on the graph of the network events to be identified, and using graph computation, the associated event information corresponding to the network events to be identified is determined. The associated event information is processed for confidence level, attack stage, and attack method to obtain the network event parsing result. The network event parsing result includes at least the target attack stage, target attack method, and target confidence level. Based on the large language model, target prompt words, and the parsing results of the network event, a natural language report corresponding to the network event to be identified is generated. Some embodiments of this application acquire network events from multiple sources, achieve automatic clustering of security alarms based on graph computing and multi-source evidence fusion, obtain related event information, process the related event information for confidence level, attack stage and attack method, obtain the analysis results of network events, and combine with a large language model to generate a network event interpretation report corresponding to the analysis results, which significantly improves the automation level and analysis efficiency of attack methods in network events.
[0005] Optionally, determining the associated event information corresponding to the network event to be identified based on the graph of the network event to be identified, using graph computation, includes: Based on the source address and the target address, establish graph data corresponding to the network event to be identified; The connected component algorithm is used to determine the associated event information in the graph data, wherein the associated event information includes a unique event identifier, a list of alarm identifiers, and start and end times. Optionally, the step of establishing the graph data corresponding to the network event to be identified based on the source address and the target address includes: The first terminal corresponding to the source address and the second terminal corresponding to the target address are designated as nodes; The alarm information and the time of occurrence of the network event to be identified are used as directed edges; Based on the nodes and the directed edges, establish graph data corresponding to the network events to be identified.
[0006] Some embodiments of this application construct an IP association graph through a specified time window and use the GraphX connected component algorithm to achieve automatic alarm clustering. Combined with standardized log mapping and multi-source evidence alignment, high-confidence associated events are formed.
[0007] Optionally, the step of processing the associated event information for confidence level, attack stage, and attack method to obtain the network event parsing result includes: Based on the pre-established confidence model, determine the target confidence level corresponding to the associated event information; Based on the pre-established attack phase mapping relationship, determine the target attack phase corresponding to the associated event information; Based on a pre-established attack pattern database, the target attack method corresponding to the associated event information is determined. Some embodiments of this application calculate the target confidence level of associated event information based on the confidence level module, and then determine the ATT&CK attack stage based on the attack stage mapping relationship. The RAG mechanism is used to semantically match the attack behavior summary with the vectorized APT knowledge base to achieve auditable advanced threat attribution. It can integrate ATT&CK attack knowledge, APT organization behavior patterns and user supplementary evidence into the security event analysis process. The technical solution supports flexible configuration of graph association windows, evaluation rules and interpretation templates according to the actual network environment, and can quickly adapt to different security operation scenarios. Optionally, determining the target confidence level corresponding to the associated event information based on a pre-established confidence model includes: The confidence model is determined based on the scores for the number of associated alarms, the time series, and the asset importance, as well as the weight values corresponding to the number of associated alarms, the time series, and the asset importance, respectively. The associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information. Some embodiments of this application combine the number of associated alarms, the time sequence, and the importance of the assets, setting different scores and corresponding weight values for different parameters to establish a confidence model. Then, the associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information. Optionally, determining the target attack stage corresponding to the associated event information based on a pre-established attack stage mapping relationship includes: Based on the correspondence between threat type, attack stage, attack identification information, and the order of occurrence, an attack stage mapping relationship is established, wherein the attack stage includes at least lateral movement, data leakage, or network attack. The associated event information and the attack stage mapping relationship are matched to obtain the target attack stage and target attack identification information corresponding to the associated event information.
[0008] Some embodiments of this application collect a set of TTPs corresponding to all alarms within an associated event, and construct a behavior time sequence chain by combining their occurrence time order. Based on predefined attack stage logic rules, the target attack stage and target attack identification information corresponding to the target event information are determined for subsequent response strategy matching and event priority sorting.
[0009] Optionally, the pre-established attack pattern database is obtained in the following manner: Acquire threat intelligence data, wherein the threat intelligence data includes at least APT reports or VirusTotal; Extract APT attack pattern data from the threat intelligence data. The APT attack pattern data includes at least the organization name, TTP sequence, attack tools, target industry, initial entry method, and lateral movement method. The APT attack pattern data is encoded using a preset encoder to obtain a first semantic vector corresponding to the APT attack pattern data. Based on the first semantic vector, the attack pattern database is constructed. Some embodiments of this application extract the attack patterns of mainstream APT organizations from publicly available threat intelligence, determine the field information under different attack patterns, concatenate the field information under different attack patterns, and use a preset encoder to convert the concatenated field information into a 768-dimensional semantic vector, which is then stored in a Milvus vector database, thus obtaining the attack pattern database. Evidence-based APT organization behavior comparison is achieved through a Retrieval-Augmented Generation (RAG) mechanism. Optionally, determining the target attack method corresponding to the associated event information based on the pre-established attack pattern database includes: The attack stage, attack identification information, alarm type, and terminal device information in the associated event information are concatenated to obtain concatenated data. The pre-defined encoder is used to encode the spliced data to obtain a second vector corresponding to the spliced data. The second vector is matched with the first vector in the pre-established attack pattern database; Based on the matching results, the target attack method corresponding to the associated event information is determined.
[0010] Optionally, determining the target attack method corresponding to the associated event information based on the matching result includes: If the first vector and the second vector match, calculate the similarity between the first vector and the second vector; If the similarity is greater than a preset threshold, the organization name corresponding to the first vector is determined as the target attack method corresponding to the associated event information.
[0011] Some embodiments of this application match a pre-established attack pattern database with vectors of associated event information and calculate similarity. Based on the matching results and similarity, the target attack method corresponding to the associated event information is determined. Secondly, some embodiments of this application provide a network event identification device, including: The acquisition module is used to acquire network events to be identified, wherein the network events to be identified include at least the source address, the destination address, the event occurrence time, and the alarm type; The determination module is used to determine the associated event information corresponding to the network event to be identified based on the graph of the network event to be identified and using graph computing. The processing module is used to process the associated event information by confidence level, attack stage and attack method to obtain the analysis result of the network event. The analysis result of the network event includes at least the target attack stage, the target attack method and the target confidence level. The reading module is used to generate a natural language report corresponding to the network event to be identified, based on the large language model, target prompt words, and the parsing results of the network event.
[0012] Some embodiments of this application acquire network events from multiple sources, achieve automatic clustering of security alarms based on graph computing and multi-source evidence fusion, obtain related event information, process the related event information for confidence level, attack stage and attack method, obtain the analysis results of network events, and combine with a large language model to generate a network event interpretation report corresponding to the analysis results, which significantly improves the automation level and analysis efficiency of attack methods in network events.
[0013] Optionally, the determining module is configured to: Based on the source address and the target address, establish graph data corresponding to the network event to be identified; The connected component algorithm is used to determine the associated event information in the graph data, wherein the associated event information includes a unique event identifier, a list of alarm identifiers, and start and end times. Optionally, the determining module is configured to: The first terminal corresponding to the source address and the second terminal corresponding to the target address are designated as nodes; The alarm information and the time of occurrence of the network event to be identified are used as directed edges; Based on the nodes and the directed edges, establish graph data corresponding to the network events to be identified.
[0014] Some embodiments of this application construct an IP association graph through a specified time window and use the GraphX connected component algorithm to achieve automatic alarm clustering. Combined with standardized log mapping and multi-source evidence alignment, high-confidence associated events are formed.
[0015] Optionally, the processing module is configured to: Based on the pre-established confidence model, determine the target confidence level corresponding to the associated event information; Based on the pre-established attack phase mapping relationship, determine the target attack phase corresponding to the associated event information; Based on a pre-established attack pattern database, the target attack method corresponding to the associated event information is determined. Some embodiments of this application calculate the target confidence level of associated event information based on the confidence level module, and then determine the ATT&CK attack stage based on the attack stage mapping relationship. The RAG mechanism is used to semantically match the attack behavior summary with the vectorized APT knowledge base to achieve auditable advanced threat attribution. It can integrate ATT&CK attack knowledge, APT organization behavior patterns and user supplementary evidence into the security event analysis process. The technical solution supports flexible configuration of graph association windows, evaluation rules and interpretation templates according to the actual network environment, and can quickly adapt to different security operation scenarios. Optionally, the processing module is configured to: The confidence model is determined based on the scores for the number of associated alarms, the time series, and the asset importance, as well as the weight values corresponding to the number of associated alarms, the time series, and the asset importance, respectively. The associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information. Some embodiments of this application combine the number of associated alarms, the time sequence, and the importance of the assets to set different scores and set corresponding weight values for different parameters to establish a confidence model. Then, the associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information.
[0016] Optionally, the processing module is configured to: Based on the correspondence between threat type, attack stage, attack identification information, and the order of occurrence, an attack stage mapping relationship is established, wherein the attack stage includes at least lateral movement, data leakage, or network attack. The associated event information and the attack stage mapping relationship are matched to obtain the target attack stage and target attack identification information corresponding to the associated event information.
[0017] Some embodiments of this application collect a set of TTPs corresponding to all alarms within an associated event, and construct a behavior time sequence chain by combining their occurrence time order. Based on predefined attack phase logic rules, the target attack phase and target attack identification information corresponding to the target event information are determined for subsequent response strategy matching and event priority sorting.
[0018] Optionally, the processing module is configured to: Acquire threat intelligence data, wherein the threat intelligence data includes at least APT reports or VirusTotal; Extract APT attack pattern data from the threat intelligence data. The APT attack pattern data includes at least the organization name, TTP sequence, attack tools, target industry, initial entry method, and lateral movement method. The APT attack pattern data is encoded using a preset encoder to obtain a first semantic vector corresponding to the APT attack pattern data. Based on the first semantic vector, the attack pattern database is constructed. Some embodiments of this application extract the attack patterns of mainstream APT organizations from publicly available threat intelligence, determine the field information under different attack patterns, concatenate the field information under different attack patterns, and use a preset encoder to convert the concatenated field information into a 768-dimensional semantic vector, which is then stored in the Milvus vector database, thus obtaining the attack pattern database. Evidence-based APT organization behavior comparison is achieved through a Retrieval-Augmented Generation (RAG) mechanism.
[0019] Optionally, the processing module is configured to: The attack stage, attack identification information, alarm type, and terminal device information in the associated event information are concatenated to obtain concatenated data. The pre-defined encoder is used to encode the spliced data to obtain a second vector corresponding to the spliced data. The second vector is matched with the first vector in the pre-established attack pattern database; Based on the matching results, the target attack method corresponding to the associated event information is determined.
[0020] Optionally, the processing module is configured to: If the first vector and the second vector match, calculate the similarity between the first vector and the second vector; If the similarity is greater than a preset threshold, the organization name corresponding to the first vector is determined as the target attack method corresponding to the associated event information.
[0021] Some embodiments of this application match vectors of pre-established attack pattern databases with those of associated event information and calculate similarity. Based on the matching results and similarity, the target attack method corresponding to the associated event information is determined.
[0022] Thirdly, some embodiments of this application provide an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the network event identification method as described in any embodiment of the first aspect.
[0023] Fourthly, some embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the network event identification method as described in any embodiment of the first aspect.
[0024] Fifthly, some embodiments of this application provide a computer program product, the computer program product including a computer program, wherein when the computer program is executed by a processor, it can implement the network event identification method as described in any embodiment of the first aspect. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of some embodiments of this application, the accompanying drawings used in some embodiments of this application will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A flowchart illustrating a method for identifying network events provided in an embodiment of this application; Figure 2 A flowchart illustrating a method for identifying network events provided in an embodiment of this application; Figure 3 A schematic diagram of the structure of a network event identification device provided in an embodiment of this application; Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of some embodiments of this application will now be described with reference to the accompanying drawings.
[0028] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0029] With the increasing complexity and frequency of cyberattacks, traditional security measures based on single alarms are no longer sufficient to meet the demands of modern network security. In this context, security operations platforms require an intelligent engine capable of automatically analyzing alarm information in network events, transforming it into higher-level "security events," and interpreting these events. Currently, manual annotation is used to label attack stages in network events. However, if a network event contains a complex attack chain, it cannot be identified. Therefore, the accuracy and efficiency of identifying network attacks in network events are low. In view of this, some embodiments of this application provide a method for identifying network events. This method includes acquiring a network event to be identified, wherein the network event to be identified includes at least a source address, a destination address, an event occurrence time, and an alarm type; determining associated event information corresponding to the network event to be identified based on a graph of the network event to be identified, using graph computation; and processing the associated event information for confidence, attack stage, and attack method to obtain the network event information. The analysis results of the network event include at least the target attack stage, the target attack method, and the target confidence level. Based on the large language model, target prompt words, and the analysis results of the network event, a natural language report corresponding to the network event to be identified is generated. In this embodiment, by acquiring network events from multiple sources, automatic clustering of security alarms is achieved based on graph computing and multi-source evidence fusion to obtain related event information. The confidence level, attack stage, and attack method of the related event information are then processed to obtain the analysis results of the network event. Combined with the large language model, a network event interpretation report corresponding to the analysis results is generated, which significantly improves the automation level and analysis efficiency of attack methods in network events.
[0030] like Figure 1 As shown, embodiments of this application provide a method for identifying network events, the method comprising: S101. Obtain the network event to be identified, wherein the network event to be identified includes at least the source address, destination address, event occurrence time, and alarm type; Specifically, the terminal device acquires the network event to be identified. The network event to be identified can be obtained from other devices, or by using a probe to acquire the network event, or by adding an extraction module to the device locally to acquire the network event. The specific acquisition method is not specifically limited in this application. The network event to be identified includes at least the source address, destination address, event occurrence, and alarm type.
[0031] S102. Based on the graph of the network events to be identified, determine the associated event information corresponding to the network events to be identified using graph computation. Specifically, the terminal device establishes a graph of network events based on the source and destination addresses of each network event to be identified. That is, the terminals corresponding to the source and destination addresses are used as nodes in the graph, and the event occurrence and alarm type are used as edges to construct the graph of the network events to be identified.
[0032] Then, a graph computation method is used to iteratively query the graph of the network events to be identified to obtain each attack event, and each attack event is an associated event, with each attack event corresponding to a connected subgraph. S103, the associated event information is processed for confidence, attack stage, and attack method to obtain the analysis results of the network events. The analysis results of the network events include at least the target attack stage, the target attack method, and the target confidence. Specifically, the terminal device determines the target confidence level corresponding to the associated event information based on a pre-established confidence level model; determines the target attack stage corresponding to the associated event information based on a pre-established attack stage mapping relationship; and determines the target attack method corresponding to the associated event information based on a pre-established attack pattern database. The target attack stage, target attack method, and target confidence level are then defined as the parsing results corresponding to the associated event information. S104: Based on the large language model, target prompt words, and the parsing results of the network event, a natural language report corresponding to the network event to be identified is generated.
[0033] Specifically, the terminal device acquires the parsing results corresponding to various related event information, including attack stage, related alarms and evidence, involved assets, confidence level, and ATT&CK TTPs. This parsing result is a structured event. Target prompts are pre-set, and the parsing result and target prompts are input into the Large Language Model (LLM). The LLM outputs a natural language report corresponding to the network event to be identified. In other words, an event interpretation module is added to the terminal device. Based on the structured security events output by the preceding module, the LLM generates natural language interpretations using predefined prompt templates, transforming the structured security events into human-readable and operable natural language summaries, thus improving human-machine collaboration efficiency. Some embodiments of this application acquire multi-source network events, achieve automatic clustering of security alarms based on graph computing and multi-source evidence fusion, obtain related event information, and process this related event information for confidence level, attack stage, and attack method to obtain the parsing results of the network events. Combined with the LLM, a network event interpretation report corresponding to the parsing results is generated, significantly improving the automation level and analysis efficiency of attack methods in network events.
[0034] Another embodiment of this application further supplements the description of the network event identification method provided in the above embodiments.
[0035] This application provides a security event intelligent analysis method that integrates graph algorithms and large language models (LLM). The method aims to automatically upgrade security alerts to higher-level, actionable security events. Deployed in a security operations platform, the method includes three core functional units: an input module, an event association module, an event evaluation module, and an event interpretation module. Through the collaborative work of multi-source data fusion, graph mining, rule reasoning, and semantic understanding, it significantly improves the identification capability and response efficiency of complex attacks.
[0036] like Figure 2 As shown, including; Step 1, Input Module: The terminal device receives multi-source security alarm information and supplementary evidence uploaded by the user (i.e., network events to be identified), and cleans them into a standard security log format.
[0037] Step 2, Event Association Module: Based on the network events to be identified, the terminal device constructs an IP directed graph, i.e., a graph spectrum, and then uses the connected component algorithm to cluster scattered alarms into associated events with attack chain semantics in the graph spectrum.
[0038] Step 3, Event Assessment Module: The terminal device automatically determines the confidence level and ATT&CK attack stage based on multi-dimensional event information and performs APT behavior semantic comparison. A multi-dimensional in-depth assessment is performed on each related event, including confidence quantification, ATT&CK stage determination, and APT behavior semantic comparison to determine whether it constitutes a real attack.
[0039] Step 4, Event Interpretation Module: The prompt words are combined with security event information to generate an interpretation report.
[0040] Optionally, based on the graph of the network events to be identified, and using graph computation, the associated event information corresponding to the network events to be identified is determined, including: Based on the source address and destination address, construct the graph data corresponding to the network event to be identified; The connected component algorithm is used to determine the associated event information in the graph data. The associated event information includes the event unique identifier, alarm identifier list, start time and end time. Optionally, based on the source address and destination address, a graph data corresponding to the network event to be identified is constructed, including: The first terminal corresponding to the source address and the second terminal corresponding to the destination address are designated as nodes; The alarm information and the time of occurrence of the network event to be identified are used as directed edges; Based on nodes and directed edges, construct graph data corresponding to the network events to be identified.
[0041] Specifically, in step 1, the terminal device receives security event alerts and supplementary evidence from multiple data sources as raw input for analysis. Security alerts can be obtained using probes or threat detection devices, and include, but are not limited to, SIEM platform logs, firewall logs, and EDR endpoint logs. Each alert contains structured fields, including but not limited to log_id (unique identifier), src_ip / dst_ip (source / destination IP address), time: event occurrence time, cat1 / cat2 (category tags such as "network attack" or "lateral movement"), and raw_log (raw log content).
[0042] Security alerts are specifically initial risk warnings automatically triggered by security detection systems (such as EDR, NDR, SIEM, firewalls, IDS / IPS, etc.) based on preset rules, abnormal behavior, or known threat characteristics. They are machine-generated, structured event notifications; typically containing fields such as timestamp, source / destination IP, user, process, TTPs tags, and alert level; representing that "the system believes there may be a security threat," but it has not yet been confirmed by manual or advanced analysis. Examples include "Suspicious PowerShell script execution detected (T1059)" and "SMB brute-force attack attempt (T1110)."
[0043] Supplementary evidence refers to raw logs or data fragments actively collected or associated to verify or enrich the context of security alerts. It typically originates from endpoints, networks, or identity systems. It is raw observation data, not the alert itself; used to support the authenticity of the alert and reconstruct attack chain details; it may include PCAP packets, EDR endpoint logs, vulnerability scan reports, etc. It originates from user uploads, retrieval by automated forensics modules, and cross-system correlation queries. For example, the PsExec execution record corresponding to the alert "Suspicious Lateral Movement".
[0044] All alarm information and supplementary evidence input data undergo unified cleaning and format conversion, mapping fields from different systems to the standard security log model. For example, process_cmdline="net use \\\\10.10.20.50\\C$" in EDR is parsed into structures such as cat1="lateral movement", cat2 = "SMB management share abuse", src_ip="10.10.20.51", and dst_ip="10.10.20.50".
[0045] The "standard security log model" refers to a unified, structured data representation specification used to map raw logs from heterogeneous security devices or systems (such as EDR, firewalls, PCAP, vulnerability scanners, etc.) to a common field system for subsequent analysis, correlation, and retrieval.
[0046] It does not refer to a specific product, but rather to an abstract data model design. It is a predefined structured event representation specification in the embodiments of this application, and its core goal is to solve the problem of "difficulty in integrating multi-source heterogeneous logs". For example, PCAP is network packets, EDR logs are process behavior, and vulnerability reports are asset risks. These logs have completely different formats. Only after the fields are aligned can cross-source behavior chain reconstruction be performed.
[0047] The specific implementation method of event association in step 2 is as follows: After acquiring a network event to be identified, the terminal device preprocesses the event to obtain a processed network event. The terminal device then enters the event correlation stage, aiming to cluster scattered alarms into "correlated events" with potential attack chains. This module utilizes the Apache Spark GraphX distributed graph computing framework to achieve efficient correlation analysis of large-scale alarms and discover related events.
[0048] The implementation steps are as follows: 1) Construct a directed attribute graph. The directed attribute graph is drawn according to the flow of IP addresses. Nodes represent devices in the network (identified by IP address), and directed edges represent the flow of data from one device to another. The directionality of the edges indicates the direction of communication / attack, for example, from the source IP address to the target IP address. Security alerts and logs are stored in an Elasticsearch (ES) cluster and indexed by time partition. End devices use configurable time windows (default is 1 day, suitable for general enterprise networks; users can also dynamically adjust the window length according to the actual network size, alert density, or security policy requirements, such as 1 hour for high-sensitivity real-time scenarios, or 7 days for long-term APT behavior backtracking) to batch pull alarm data to be processed from ES through the SparkES-Hadoop connector and convert it into a structured DataFrame with fields including src_ip, dst_ip, timestamp, alarm_id, cat2, time, etc. Using IP addresses as nodes, each alarm corresponds to a directed edge, with the source vertex being the source address src_ip and the destination vertex being the destination address dst_ip. Edge attributes include alarm identifier alarm_id, category label cat2 (scan, brute force, virus) and timestamp, constructing a dynamic graph.
[0049] 2) Connected Component Mining. The `ConnectedComponents.run()` connected component algorithm, built into GraphX (a distributed graph computing library), is invoked to perform connected component calculations on the entire graph for the day. This algorithm is based on the Pregel model (a graph distributed graph algorithm) and, through multiple rounds of message passing, assigns each vertex the smallest vertex ID of its connected component as a component ID. All vertices belonging to the same component ID and their connecting edges constitute a candidate for associated events. Attack events form a connected subgraph.
[0050] The use of the smallest vertex ID as the component identifier is standard behavior in the Connected Components algorithm (especially in GraphX's Pregel model-based implementation). In an undirected graph (or a weakly connected directed graph), a connected component is a set of nodes that are reachable from each other. The goal of the algorithm is to assign a unique "label" (i.e., Component ID) to all vertices within each connected component so that subsequent identification can be made of nodes belonging to the same associated event. Spark GraphX automatically selects the ID of the node with the smallest vertex ID in the connected component as the Component ID for the entire component.
[0051] The Component ID is not manually assigned, but rather a deterministic identifier automatically generated by the algorithm. Vertex IDs are typically assigned automatically by the system during graph construction (e.g., using IP address hashes, unique log event IDs, or globally increasing sequences), but once the graph is built, each vertex's ID is fixed. This ID is calculated entirely by the algorithm based on the graph's topology, requiring no manual intervention or predefinition, and is used to uniquely identify a cluster of potential multi-hop attack events.
[0052] 3) Generate associated events. Group the edges by component ID, and output each group as an "associated event" object, which includes attributes such as unique event ID, alarm ID list, start time, and end time. The object is then output to the ES security event index for consumption by the event evaluation module.
[0053] Some embodiments of this application construct an IP association graph through a specified time window and use the GraphX connected component algorithm to achieve automatic alarm clustering. Combined with standardized log mapping and multi-source evidence alignment, high-confidence associated events are formed.
[0054] Optionally, the related event information is processed according to confidence level, attack stage, and attack method to obtain the analysis results of the network events, including: Based on the pre-established confidence model, determine the target confidence level corresponding to the associated event information; Based on the pre-established attack phase mapping relationship, determine the target attack phase corresponding to the associated event information; Based on a pre-established attack pattern database, the target attack method corresponding to the associated event information is determined. Some embodiments of this application calculate the target confidence level of associated event information based on the confidence level module, and then determine the ATT&CK attack stage based on the attack stage mapping relationship. The RAG mechanism is used to semantically match the attack behavior summary with the vectorized APT knowledge base to achieve auditable advanced threat attribution. It can integrate ATT&CK attack knowledge, APT organization behavior patterns and user supplementary evidence into the security event analysis process. The technical solution supports flexible configuration of graph association windows, evaluation rules and interpretation templates according to the actual network environment, and can quickly adapt to different security operation scenarios. Optionally, based on a pre-established confidence model, a target confidence level corresponding to the associated event information is determined, including: The confidence model is determined based on the scores for the number of associated alarms, the time series, and the asset importance, as well as the weight values corresponding to the number of associated alarms, the time series, and the asset importance, respectively. The associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information.
[0055] Specifically, the terminal device comprehensively considers the quantity and timing of associated alarms, the importance of the assets involved (such as whether they are critical systems such as domain controllers or database servers), whether the target assets have known vulnerabilities, and whether there is supplementary terminal or network evidence from users (such as EDR process logs, firewall connection records, etc.) to build a confidence model. Then, the associated event information is input into the confidence model, and the confidence level of 0-1 is output.
[0056] The "confidence model" is a scoring mechanism used to quantify the probability that a related event is a real attack (rather than a false positive or coincidence). It can be understood as a configurable and interpretable weighted fusion model, belonging to the weighted rule model. Each factor is assigned a weight (which can be configured by security experts or calibrated through historical data). After normalization of each dimension, the weighted sum is calculated and then mapped to the [0,1] interval. The "Confidence Score" result will be one of the main attributes of the output security events, allowing security operations personnel to sort all security events in descending order based on confidence level, with high-confidence events being viewed first. Additionally, the "Event Interpretation" section also incorporates confidence level to improve interpretability. The large model will also guide event response based on confidence level; if the confidence level is >0.9, automatic isolation is recommended; if the confidence level is <0.6, the event will be archived for further investigation.
[0057] Some embodiments of this application combine the number of associated alarms, their timing, and asset importance, setting different scores and corresponding weight values for different parameters to establish a confidence model. Then, the associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information. Optionally, based on a pre-established attack stage mapping relationship, the target attack stage corresponding to the associated event information is determined, including: Based on the correspondence between threat type, attack stage, attack identification information, and the order of occurrence, an attack stage mapping relationship is established. Attack stages include at least lateral movement, data leakage, or network attack. Matching the associated event information with the attack stage mapping relationship yields the target attack stage and target attack identification information corresponding to the associated event information.
[0058] Specifically, the terminal device automatically identifies the MITRE ATT&CK attack lifecycle stage of the associated event. Each original alarm is associated with one or more ATT&CK technology identifiers (TTPs) through built-in mapping rules upon access. For example, an SMB brute-force attack alert is mapped to T1021.002 (SMB / Windows Admin Shares), and a suspicious PowerShell execution is mapped to T1059.001 (PowerShell).
[0059] The terminal device collects a set of TTPs corresponding to all alarms within the associated event and constructs a behavior time sequence chain by combining their occurrence order. Subsequently, based on predefined attack phase logic rules, it automatically infers the most likely main attack phase of the current event. This determination result is output as a structured field for subsequent response strategy matching and event priority ranking.
[0060] For example, if T1190 (Exploit Public-Facing Application) is followed immediately by T1059 (Command and Scripting Interpreter), it is determined to be "Initial Access → Execution", and the main attack phase most likely to be in the current event is automatically inferred, such as "lateral movement" or "data leakage".
[0061] Some embodiments of this application collect a set of TTPs corresponding to all alarms within an associated event, and construct a behavior time sequence chain by combining their occurrence time order. Based on predefined attack phase logic rules, the target attack phase and target attack identification information corresponding to the target event information are determined for subsequent response strategy matching and event priority sorting.
[0062] Optionally, the pre-established attack pattern database is obtained in the following way: Acquire threat intelligence data, which includes at least APT reports or VirusTotal; Extract APT attack pattern data from threat intelligence data. APT attack pattern data should include at least the organization name, TTP sequence, attack tools, target industry, initial entry method, and lateral movement method. A preset encoder is used to encode the APT attack pattern data to obtain the first semantic vector corresponding to the APT attack pattern data; An attack pattern database is constructed based on the first semantic vector.
[0063] Specifically, the terminal device first extracts various attack patterns of mainstream APT groups from publicly available threat intelligence (such as MITRE ATT&CK, APT reports, VirusTotal, etc.). Each record contains structured fields, including the organization name, commonly used TTPs sequences, attack tools (such as Cobalt Strike, Mimikatz), target industry (such as finance, energy), initial entry method, and lateral movement techniques. Subsequently, a preset encoder is used to convert the descriptive text concatenated from the above multiple fields into a 768-dimensional semantic vector and store it in the Milvus vector database to build an efficient Approximate Nearest Neighbor (ANN) index. The preset encoder includes at least a security-domain fine-tuned Sentence-BERT encoder.
[0064] Some embodiments of this application extract the attack patterns of mainstream APT organizations from publicly available threat intelligence, determine the field information under different attack patterns, concatenate the field information under different attack patterns, use a preset encoder to convert the concatenated field information into a 768-dimensional semantic vector, and store it in the Milvus vector database to obtain the attack pattern database. The evidence-based comparison of APT organization behavior is achieved through the Retrieval-Augmented Generation (RAG) mechanism.
[0065] Optionally, based on a pre-established attack pattern database, the target attack method corresponding to the associated event information is determined, including: The attack stage, attack identification information, alarm type and terminal device information in the related event information are spliced together to obtain spliced data; A preset encoder is used to encode the spliced data to obtain a second vector corresponding to the spliced data; The second vector is matched against the first vector in a pre-established attack pattern database; Based on the matching results, the target attack method corresponding to the associated event information is determined.
[0066] Optionally, based on the matching results, the target attack method corresponding to the associated event information is determined, including: If the first vector and the second vector match, calculate the similarity between the first vector and the second vector; If the similarity is greater than a preset threshold, the organization name corresponding to the first vector will be determined as the target attack method corresponding to the associated event information.
[0067] Specifically, to further uncover hidden threat patterns, terminal devices can optionally enable the large model-assisted analysis function during the event assessment phase. Its core lies in achieving evidence-based comparison of APT organization behaviors through the Retrieval-Augmented Generation (RAG) mechanism.
[0068] Specifically, the terminal device first automatically extracts key elements from the associated events. These key elements include ATT&CK TTPs, alarm types, and terminal or network evidence uploaded by the user. The key elements are then concatenated to obtain a well-structured and semantically complete summary text of the attack behavior. In other words, key elements are extracted from the associated events in a structured manner to form a summary text and a vector. Then, the record with the most similar vector is queried from the knowledge base to obtain the associated attack organization.
[0069] For example, the attacker uses SMB brute force to gain access to the file server and executes a lateral movement command after 7 seconds.
[0070] The terminal device converts the summary text into a high-dimensional vector (i.e., a second vector) using an embedding model (such as Sentence-BERT or a dedicated security domain encoder), and uses it as query input to a vector database (such as Milvus) to perform semantic similarity matching with a pre-built and regularly updated knowledge base of APT organization behavior.
[0071] The terminal device matches the second vector with the first vector in the vector database, i.e., the attack pattern database, returns the Top-1 matching result, and calculates the similarity score between the first vector and the second vector. If the similarity score exceeds a preset threshold (e.g., 0.85), the statement "suspected to be highly similar to the behavior of [organization name]" will be added to the security incident as enhanced evidence. The entire process does not rely on large language models to directly generate attribution conclusions, but is based on auditable search results to ensure that the output is objective, interpretable, and traceable.
[0072] a. The terminal device will integrate the results of event correlation and event assessment to form a structured security event, which includes fields such as associated alarms and evidence, confidence level, attack stage, TTPs, and suspected attacking organization.
[0073] Some embodiments of this application match vectors of pre-established attack pattern databases with those of associated event information and calculate similarity. Based on the matching results and similarity, the target attack method corresponding to the associated event information is determined.
[0074] For example, the network event identification method provided in this application includes: 1. Input security alerts and supplementary evidence: Specifically, the terminal device receives network alarms, EDR terminal logs, and supplementary evidence uploaded by analysts from the SIEM platform within a preset time period. For example, the preset time period could be one day, 8 hours, etc. This includes: A WAF alert: src_ip=203.0.113.45, dst_ip=10.10.10.100 (webserver01),cat2="Exploit", raw_log="Apache Log4j RCE attempt"; An EDR log entry: host=webserver01, process_cmdline="powershell -enc ...download payload.ps1", timestamp is 38 seconds after the alert; An internal firewall alert: src_ip=10.10.10.100, dst_ip=10.10.20.50(fileserver01), cat2=SMB brute-force attack, time=7 minutes after the alert; The terminal device obtains the supplementary PCAP package uploaded by the analyst and confirms that the SMB connection contains the command `net use \\fileserver01\C$`.
[0075] All data is preprocessed and uniformly mapped to standard security logs, with field alignment completed.
[0076] 2. Event Relationships: The terminal device constructs a directed IP graph based on all alarms for the day: nodes include 203.0.113.45, 10.10.10.100, and 10.10.20.50, and edges are generated from the above three records. After calling the Spark GraphX's ConnectedComponents.run() algorithm, the three are clustered into the same connected component (Component ID = hash(10.10.10.100)), generating a related event event_20251208_042, which contains all three alarm IDs and the time span (02:15–02:22).
[0077] 3. Incident Assessment: Confidence Calculation: Due to the inclusion of 3 reasonable time-series alarms, the involvement of critical web and file servers, and the support of dual evidence from EDR and PCAP, the system outputs a confidence level of 0.92; ATT&CK phase determination: The alarms are mapped to T1190 (Log4j exploit), T1059.001 (PowerShell execution), and T1021.002 (SMB lateral movement). Based on the timing sequence, the attack phase is inferred to be "initial access → execution → lateral movement". Agent Assessment: The system generated an attack behavior summary: "The attacker compromised the web server through the Log4j vulnerability, executed PowerShell to download the payload, and moved laterally to the file server via SMB 7 minutes later." After searching the APT29 knowledge base with RAG, the similarity was 0.88, and the behavior was marked as "suspected APT29 behavior".
[0078] 4. Event Interpretation The prompt can be designed as follows: For a senior cybersecurity analyst, please write a concise, objective, and actionable security incident analysis report based on the following complete factual data of the security incident.
[0079] [Input Data] Event ID: event_20251208_042 The attack phases involved (ATT&CK) are: Initial Access (T1190), Execution (T1059.001), and Lateral Movement (T1021.002). Key chain of evidence: 1. [2025-12-08 02:15:23] External IP 203.0.113.45 launched an exploit attack on webserver01 (Apache Log4j RCE); 2. [2025-12-08 02:16:01] webserver01 executed a suspicious PowerShell command to download payload.ps1; 3. [2025-12-08 02:22:17] webserver01 initiates an SMB connection to fileserver01 and executes net use \\\\fileserver01\\C$; Assets involved: webserver01 (Internet Web server), fileserver01 (Domain file server); Confidence level: 0.92 Additional evidence: EDR logs confirm that the PowerShell process was started by a child process of java.exe; Output Requirements 1. Using only the above input data, you must not speculate on the attacker's identity, motives, or unmentioned actions; 2. The report should be organized chronologically according to the attack timeline, clearly presenting the complete chain from the entry point to the lateral movement; 3. Includes: an overview of the attack path, analysis of key behaviors, affected assets, and specific handling recommendations; Output Format Please output the main body of the report directly; no title or additional description is required.
[0080] The structured events are fed into the interpretation module, and the LLM generates the following natural language report based on preset prompts: Attackers compromised the internet web server webserver01 using a Log4j remote code execution vulnerability (T1190), then executed a PowerShell script (T1059.001) to download a malicious payload. Six minutes later, they used the SMB protocol (T1021.002) to lateral move to the domain file server fileserver01. EDR logs confirmed that the malicious process was derived from java.exe. It is recommended to immediately isolate webserver01 and fileserver01, reset the local and domain account passwords on both hosts, check the log4j component version and apply patches, and audit the access records of sensitive files on fileserver01.
[0081] It should be noted that each of the implementable methods in this embodiment can be implemented individually or in any combination without conflict. This application does not limit this.
[0082] Another embodiment of this application provides a network event identification device for performing the network event identification method provided in the above embodiments.
[0083] like Figure 3 The diagram shown is a structural schematic of a network event identification device provided in an embodiment of this application. The network event identification device includes an acquisition module 301, a determination module 302, a processing module 303, and a reading module 304, wherein: The acquisition module 301 is used to acquire network events to be identified, wherein the network events to be identified include at least the source address, destination address, event occurrence time, and alarm type; The determination module 302 is used to determine the associated event information corresponding to the network event to be identified based on the graph of the network event to be identified and using graph computation. The processing module 303 is used to process the confidence level, attack stage and attack method of the associated event information to obtain the analysis result of the network event. The analysis result of the network event includes at least the target attack stage, the target attack method and the target confidence level. The reading module 304 is used to generate a natural language report corresponding to the network event to be identified based on the parsing results of the large language model, target prompt words and network events.
[0084] Regarding the apparatus in this embodiment, the specific manner in which each module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0085] Some embodiments of this application acquire network events from multiple sources, achieve automatic clustering of security alarms based on graph computing and multi-source evidence fusion, obtain related event information, process the related event information for confidence level, attack stage and attack method, obtain the analysis results of network events, and combine with a large language model to generate a network event interpretation report corresponding to the analysis results, which significantly improves the automation level and analysis efficiency of attack methods in network events.
[0086] Another embodiment of this application further supplements the description of the network event identification device provided in the above embodiments.
[0087] Optionally, a module is defined for: Based on the source address and destination address, construct the graph data corresponding to the network event to be identified; The connected component algorithm is used to determine the associated event information in the graph data. The associated event information includes the event unique identifier, alarm identifier list, start time and end time. Optionally, a module is defined for: The first terminal corresponding to the source address and the second terminal corresponding to the destination address are designated as nodes; The alarm information and the time of occurrence of the network event to be identified are used as directed edges; Based on nodes and directed edges, construct graph data corresponding to the network events to be identified.
[0088] Some embodiments of this application construct an IP association graph through a specified time window and use the GraphX connected component algorithm to achieve automatic alarm clustering. Combined with standardized log mapping and multi-source evidence alignment, high-confidence associated events are formed.
[0089] Optionally, the processing module is used for: Based on the pre-established confidence model, determine the target confidence level corresponding to the associated event information; Based on the pre-established attack phase mapping relationship, determine the target attack phase corresponding to the associated event information; Based on a pre-established attack pattern database, the target attack method corresponding to the associated event information is determined. Some embodiments of this application calculate the target confidence level of associated event information based on the confidence level module, and then determine the ATT&CK attack stage based on the attack stage mapping relationship. The RAG mechanism is used to semantically match the attack behavior summary with the vectorized APT knowledge base to achieve auditable advanced threat attribution. It can integrate ATT&CK attack knowledge, APT organization behavior patterns and user supplementary evidence into the security event analysis process. The technical solution supports flexible configuration of graph association windows, evaluation rules and interpretation templates according to the actual network environment, and can quickly adapt to different security operation scenarios. Optionally, the processing module is used for: The confidence model is determined based on the scores for the number of associated alarms, the time series, and the asset importance, as well as the weight values corresponding to the number of associated alarms, the time series, and the asset importance, respectively. The associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information. Some embodiments of this application combine the number of associated alarms, timing, and asset importance to set different scores and set corresponding weight values for different parameters to establish a confidence model. Then, the associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information.
[0090] Optionally, the processing module is used for: Based on the correspondence between threat type, attack stage, attack identification information, and the order of occurrence, an attack stage mapping relationship is established. Attack stages include at least lateral movement, data leakage, or network attack. Matching the associated event information with the attack stage mapping relationship yields the target attack stage and target attack identification information corresponding to the associated event information.
[0091] Some embodiments of this application collect a set of TTPs corresponding to all alarms within an associated event, and construct a behavior time sequence chain by combining their occurrence time order. Based on predefined attack phase logic rules, the target attack phase and target attack identification information corresponding to the target event information are determined for subsequent response strategy matching and event priority sorting.
[0092] Optionally, the processing module is used for: Acquire threat intelligence data, which includes at least APT reports or VirusTotal; Extract APT attack pattern data from threat intelligence data. APT attack pattern data should include at least the organization name, TTP sequence, attack tools, target industry, initial entry method, and lateral movement method. A preset encoder is used to encode the APT attack pattern data to obtain the first semantic vector corresponding to the APT attack pattern data; Based on the first semantic vector, an attack pattern database is constructed. Some embodiments of this application extract the attack patterns of mainstream APT organizations from publicly available threat intelligence, determine the field information under different attack patterns, concatenate the field information under different attack patterns, and use a preset encoder to convert the concatenated field information into a 768-dimensional semantic vector, which is then stored in the Milvus vector database, thus obtaining the attack pattern database. The evidence-based comparison of APT organization behavior is achieved through the Retrieval-Augmented Generation (RAG) mechanism.
[0093] Optionally, the processing module is used for: The attack stage, attack identification information, alarm type and terminal device information in the related event information are spliced together to obtain spliced data; A preset encoder is used to encode the spliced data to obtain a second vector corresponding to the spliced data; The second vector is matched against the first vector in a pre-established attack pattern database; Based on the matching results, the target attack method corresponding to the associated event information is determined.
[0094] Optionally, the processing module is used for: If the first vector and the second vector match, calculate the similarity between the first vector and the second vector; If the similarity is greater than a preset threshold, the organization name corresponding to the first vector will be determined as the target attack method corresponding to the associated event information.
[0095] Some embodiments of this application match vectors of pre-established attack pattern databases with associated event information and calculate similarity. Based on the matching results and similarity, the target attack method corresponding to the associated event information is determined.
[0096] Regarding the apparatus in this embodiment, the specific manner in which each module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0097] It should be noted that each of the implementable methods in this embodiment can be implemented individually or in any combination without conflict. This application does not limit this.
[0098] This application also provides a computer-readable storage medium storing a computer program thereon. When the program is executed by a processor, it can implement the operation of any of the methods corresponding to the network event identification methods provided in the above embodiments.
[0099] This application also provides a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operation of any of the methods corresponding to the network event identification methods provided in the above embodiments.
[0100] like Figure 4 As shown, some embodiments of this application provide an electronic device 400, which includes: a memory 410, a processor 420, and a computer program stored in the memory 410 and executable on the processor 420. When the processor 420 reads the program from the memory 410 via a bus 430 and executes the program, it can implement any of the methods included in the above-described network event identification method.
[0101] Processor 420 can process digital signals and may include various computing architectures. For example, it may be a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements multiple instruction set combinations. In some examples, processor 420 may be a microprocessor.
[0102] Memory 410 can be used to store instructions executed by processor 420 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all of the functions of one or more modules described in the embodiments of this application. The processor 420 of this disclosure embodiment can be used to execute instructions in memory 410 to implement the methods shown above. Memory 410 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.
[0103] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0105] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for identifying network events, characterized in that, The method includes: Obtain network events to be identified, wherein the network events to be identified include at least the source address, destination address, event occurrence time, and alarm type; Based on the graph of the network events to be identified, the associated event information corresponding to the network events to be identified is determined using graph computing. This includes: taking the terminals corresponding to the source address and destination address as nodes of the graph, and then taking the event occurrence and alarm type as edges to construct the graph of the network events to be identified. Using graph computing, the graph of the network events to be identified is iteratively queried to obtain each attack event, and each attack event is an associated event, with one attack event corresponding to one connected subgraph. The associated event information is processed for confidence level, attack stage, and attack method to obtain the network event parsing result. The network event parsing result includes at least the target attack stage, target attack method, and target confidence level. Based on the large language model, target prompt words, and the parsing results of the network event, a natural language report corresponding to the network event to be identified is generated; in, The process of processing the associated event information for confidence level, attack stage, and attack method to obtain the network event parsing results includes: Based on a pre-established confidence model, a target confidence level corresponding to the associated event information is determined, wherein the confidence model is determined based on the score of the number of associated alarms, the score of the time sequence, the score of the asset importance, and the corresponding weights. Based on the pre-established attack phase mapping relationship, the target attack phase corresponding to the associated event information is determined. The attack phase mapping relationship is determined based on the correspondence between threat type, attack phase, attack identification information and the order of occurrence. The attack phase includes at least lateral movement, data leakage or network attack. Based on a pre-established attack pattern database, the target attack method corresponding to the associated event information is determined, wherein determining the target attack method corresponding to the associated event information based on the pre-established attack pattern database includes: The attack stage, attack identification information, alarm type, and terminal device information in the associated event information are concatenated to obtain concatenated data. A preset encoder is used to encode the spliced data to obtain a second vector corresponding to the spliced data; The second vector is matched with the first vector in the pre-established attack pattern database; Based on the matching results, the target attack method corresponding to the associated event information is determined.
2. The method for identifying network events according to claim 1, characterized in that, The associated event information includes a unique event identifier, a list of alarm identifiers, and start and end times.
3. The method for identifying network events according to claim 1, characterized in that, The step of determining the target confidence level corresponding to the associated event information based on a pre-established confidence model includes: The associated event information is input into the confidence model to obtain the target confidence level corresponding to the associated event information.
4. The method for identifying network events according to claim 1, characterized in that, The step of determining the target attack stage corresponding to the associated event information based on the pre-established attack stage mapping relationship includes: The associated event information and the attack stage mapping relationship are matched to obtain the target attack stage and target attack identification information corresponding to the associated event information.
5. The method for identifying network events according to claim 1, characterized in that, The pre-established attack pattern database is obtained in the following way: Acquire threat intelligence data, wherein the threat intelligence data includes at least APT reports or VirusTotal; Extract APT attack pattern data from the threat intelligence data. The APT attack pattern data includes at least the organization name, TTP sequence, attack tools, target industry, initial entry method, and lateral movement method. The APT attack pattern data is encoded using a preset encoder to obtain a first semantic vector corresponding to the APT attack pattern data. The attack pattern database is constructed based on the first semantic vector.
6. The method for identifying network events according to claim 5, characterized in that, The step of determining the target attack method corresponding to the associated event information based on the matching results includes: If the first vector and the second vector match, calculate the similarity between the first vector and the second vector; If the similarity is greater than a preset threshold, the organization name corresponding to the first vector is determined as the target attack method corresponding to the associated event information.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it can implement the network event identification method according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, characterized in that, when the program is executed by a processor, it can implement the network event identification method according to any one of claims 1-6.
9. A computer program product, said computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it can implement the network event identification method according to any one of claims 1-6.
Citation Information
Patent Citations
Threat event identification method and related device
CN116827564A
LLM-driven industrial network intrusion detection method and response system
CN118381627A