A Network Security Alert Method, Device, Equipment, and Storage Medium
By building an alarm aggregation rule tree, clustering and merging rule tree, and generating directed association graphs, the problem of alarm storm in network security events is solved, and the rule-level logical relationship maintenance and efficient handling of security events is achieved.
Patent Information
- Application Number
- CN202510327002.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-19
AI Technical Summary
When handling network security incidents, the prior art is prone to alarm storms due to repeated or conflicting alarm logs, making it difficult for security operators to conduct overall analysis and effective handling, thereby expanding the losses of security incidents.
By extracting the alarm meta-verb set from structured knowledge and semi-structured knowledge, constructing alarm aggregation rules and generating rule trees, determining the similarity between rule trees for clustering and merging, generating directed correlation graphs based on rule relationships, and finally aggregating alarm data based on rule trees and directed correlation graphs to generate network security events.
Reduce the repetition of rules, obtain and maintain logical relationships, avoid duplication or suppression of triggering security events, thereby improving the processing efficiency and accuracy of security events.
Smart Images

Figure CN119854044B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security, and particularly to a network security warning method, a network security warning device, a network security warning equipment, and a computer-readable storage medium. Background Art
[0002] In the scenario of security operation, security personnel are responsible for monitoring, protecting, and maintaining the overall security of the security system of the target organization. For different types of network attacks, security personnel need to handle different types of security devices, such as firewalls, intrusion detection systems, vulnerability scanning systems, bastion hosts, honeypots, and so on. When a security event occurs, multiple devices may simultaneously generate logs that are repetitive, similar, or conflicting with each other. The warning storm caused by a large number of repeated warnings submerges the attention of security operators in a vast number of warnings, making it impossible to conduct an overall analysis of security events and take effective disposal measures, thereby further expanding the losses caused by security events.
[0003] The prior art generally aggregates warnings into security events through rule aggregation, but there may be problems of similarity, overlap, and opposition between the rules used for aggregation, resulting in the occurrence of duplicate or suppressed triggering of security events in the end. Summary of the Invention
[0004] The purpose of the present invention is to provide a network security warning method, device, equipment, and storage medium, which are applied to the field of network security. This method constructs a rule tree through warning aggregation rules, merges similar rule trees, and associates the merged rule trees based on rule relationships, reducing the duplication degree between rules, obtaining and maintaining logical relationships from the rule level, and avoiding the occurrence of duplicate or suppressed triggering of security events.
[0005] To solve the above technical problems, the present invention provides a network security warning method, including:
[0006] Extracting a warning primitive set from structured knowledge and semi-structured knowledge, and constructing a warning aggregation rule based on the warning primitive set;
[0007] Constructing a rule tree based on the warning aggregation rule; the root node in the rule tree is the warning primitive set, the child nodes are entities in the warning primitive set, and the node parent-child relationship is the pre-order of the internal entities in the warning primitive set;
[0008] Determining the inter-tree similarity between each rule tree, clustering the rule trees through a clustering algorithm based on the inter-tree similarity, and merging the rule trees in each cluster;
[0009] After the merging is completed, a directed association graph is generated by associating the rule trees based on the rule relationships; the vertices of the directed association graph are the rule trees, and the edges are the rule relationships;
[0010] Obtain alarm data, and aggregate the alarm data based on the rule tree and the directed association graph to generate a network security event.
[0011] Optionally, determine the similarity between the rule trees, and cluster the rule trees through a clustering algorithm based on the similarity between the trees, including:
[0012] Extract node features from the node attributes of the rule tree based on a feature extraction function;
[0013] Determine the hash feature value of the node features based on the SimHash algorithm, and determine the first similarity between the rule trees based on the hash feature value;
[0014] Determine a first similarity threshold, and classify the rule trees with the first similarity greater than the first similarity threshold to obtain a set of initially screened similar trees;
[0015] In the set of initially screened similar trees, determine the tree edit distance between the rule trees based on the Zhang-Shasha algorithm, and determine the second similarity between the rule trees based on the tree edit distance;
[0016] Determine a second similarity threshold, and cluster the rule trees in each set of initially screened similar trees through the clustering algorithm until the similarity between the clusters is lower than the second similarity threshold.
[0017] Optionally, merge the rule trees in each cluster, including:
[0018] Determine a reference tree in each cluster, and align the remaining rule trees in each cluster with the reference tree based on a tree alignment algorithm;
[0019] Merge the aligned nodes, and retain the feature branches and feature nodes of each rule tree.
[0020] Optionally, the network security alarm method further includes:
[0021] When a new rule tree is generated, determine the hash feature values of the node features of the new rule tree and the rule tree based on the SimHash algorithm;
[0022] Determine the first similarity between the new rule tree and the rule tree based on the hash feature value, and determine the maximum first similarity from the first similarities;
[0023] When the maximum first inter-tree similarity is greater than the first similarity threshold, merge the new rule tree with the corresponding most similar rule tree and update the directed association graph;
[0024] When the maximum first inter-tree similarity is less than the first similarity threshold, determine the tree edit distance between the rule trees based on the Zhang-Shasha algorithm, and determine the second inter-tree similarity between the new rule tree and the rule trees based on the tree edit distance;
[0025] Determine the maximum second inter-tree similarity from the second inter-tree similarities. When the second inter-tree similarity is greater than the second similarity threshold, merge the new rule tree with the corresponding most similar rule tree and update the directed association graph.
[0026] Optionally, aggregating the alarm data based on the rule tree and the directed association graph to generate a network security event includes:
[0027] Convert the rule tree and the directed association graph into SQL-like rule statements;
[0028] Create an alarm aggregation engine and load the rule statements into the alarm aggregation engine;
[0029] Obtain the alarm data, input the alarm data into the normalizer and shunt of the alarm aggregation engine for processing, so as to shunt the normalized alarm data to the corresponding alarm aggregation processors in the alarm aggregation engine;
[0030] Aggregate the normalized alarm data based on the alarm aggregation processors to generate the network security event.
[0031] Optionally, the basic syntax structure of the rule statement includes: a WITH clause for defining intermediate variables and aliases within the alarm aggregation rule; a FROM clause for defining the alarm source; a WHERE clause for defining the pre-filtering conditions of the alarm; a SELECT clause for performing field projection and conversion; a GROUP BY clause for grouping alarms according to a preset dimension; a HAVING clause for defining the trigger conditions of the network security event.
[0032] Optionally, the network security alarm method further includes:
[0033] Determine a handling solution for the network security event based on a predefined security knowledge graph and send the handling solution to the operation and maintenance terminal;
[0034] When receiving the handling completion information of the network security event, update the aggregation period of the alarm aggregation engine.
[0035] To solve the above technical problems, the present invention provides a network security warning device, including:
[0036] A first module, configured to extract a warning primitive set from structured knowledge and semi-structured knowledge, and construct a warning aggregation rule based on the warning primitive set;
[0037] A second module, configured to construct a rule tree based on the warning aggregation rule; the root node in the rule tree is the warning primitive set, the child nodes are the entities in the warning primitive set, and the parent-child relationship between the nodes is the pre-order of the internal entities in the warning primitive set;
[0038] A third module, configured to determine the similarity between the rule trees, cluster the rule trees through a clustering algorithm based on the similarity between the rule trees, and merge the rule trees in each cluster to obtain a merged rule tree;
[0039] A fourth module, configured to generate a directed association graph by associating the merged rule tree based on a rule relationship; the vertices of the directed association graph are the merged rule trees, and the edges are the rule relationships;
[0040] A fifth module, configured to obtain warning data, and aggregate the warning data based on the merged rule tree and the directed association graph to generate a network security event.
[0041] To solve the above technical problems, the present invention provides a network security warning device, including:
[0042] A memory, configured to store a computer program;
[0043] A processor, configured to implement the above-mentioned network security warning method when executing the computer program.
[0044] To solve the above technical problems, the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, the above-mentioned network security warning method is implemented.
[0045] It can be seen that the method of the present invention extracts an alarm primitive set from structured knowledge and semi-structured knowledge, constructs an alarm aggregation rule based on the alarm primitive set; constructs a rule tree based on the alarm aggregation rule; the root node in the rule tree is the alarm primitive set, the child nodes are entities in the alarm primitive set, and the parent-child relationship between the nodes is the pre-order of the entities inside the alarm primitive set; determines the similarity between the rule trees, clusters the rule trees through a clustering algorithm based on the similarity between the rule trees, and merges the rule trees in each cluster; after the merging is completed, associates the rule trees based on the rule relationship to generate a directed association graph; the vertices of the directed association graph are the rule trees, and the edges are the rule relationships; obtains alarm data, and aggregates the alarm data based on the rule trees and the directed association graph to generate a network security event.
[0046] By constructing a rule tree through the alarm aggregation rule, merging similar rule trees, and associating the merged rule trees based on the rule relationship, the duplication between the rules is reduced, the logical relationship is obtained and maintained at the rule level, and the situation of repeated or suppressed triggering of security events is avoided. Brief Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0048] Figure 1 It is a flowchart of a network security alarm method provided by an embodiment of the present invention;
[0049] Figure 2 It is a structural framework diagram of an alarm aggregation engine provided by an embodiment of the present invention;
[0050] Figure 3 It is an example diagram of an alarm aggregation process provided by an embodiment of the present invention;
[0051] Figure 4 It is a structural block diagram of a network security alarm device provided by an embodiment of the present invention. Detailed Embodiments
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] In the scenario of security operation, security personnel are required to monitor, protect and maintain the overall security of the target organization's security system. To address different types of cyberattacks, security personnel need to handle different types of security devices, such as firewalls, intrusion detection systems, vulnerability scanning systems, bastion hosts, honeypots, etc. When security incidents occur, multiple devices may simultaneously generate logs that are repetitive, similar or conflicting with each other. The alert storm caused by a large number of repeated alerts overwhelms the attention of security operators in a vast amount of alerts, preventing them from conducting a holistic analysis of security incidents and taking effective countermeasures, thus further expanding the losses caused by security incidents.
[0054] In a complete and effective lifecycle of security incident handling, the generation of an alert only marks the start of the process. Subsequently, it is necessary to aggregate related alerts into a security incident, establish the correlation relationships between events, and then evaluate indicators such as effectiveness, threat level, and impact scope for the incident. Next, assign the handling measures and responsible persons for the incident, and finally record the handling results in the incident for subsequent post-mortem review of the incident. In the alert aggregation stage, it is necessary to accurately identify the correlation between alerts, aggregate repeated alerts, and associate alerts with logical relationships to form a disposable security incident. To improve the response speed and response rate of security operators, the incident needs to provide corresponding handling measures, and at the same time, the incident should be sorted according to indicators such as threat level and impact scope so that security operators can give priority to handling incidents with greater harm.
[0055] To solve the alert storm problem, in existing technical solutions, a rule engine is generally used to dynamically load and execute alert aggregation rules, and at the same time, methods for the distribution, storage, loading, and status update of rules are declared. The above solutions mainly aggregate alerts with relatively single attributes in the power industry, such as the magnitude and change period of voltage. However, in the network security industry, the attributes of alerts are more complex, and there are more complex correlation relationships between alerts. More correlation analysis and dynamic adjustment of aggregation status are required during the aggregation process to improve the aggregation effect. This solution cannot effectively handle the repetition of alert aggregation rules, the need for dynamic update of aggregation status, and the correlation problem between aggregated alerts. The fields designed for alert aggregation in this solution are relatively single, and the method of directly triggering security incidents by individual rules is not applicable to the complex alert logs in the network security scenario. Network security alerts have the characteristics of complex alert attributes and complex correlation relationships between alerts. Therefore, this solution cannot be well migrated from the power industry to the network security industry.
[0056] In the prior art, there are also solutions that use rules + a streaming engine (Flink). For network security scenarios, the relevance analysis ability is improved by classifying aggregation rules. The classification includes configuration-based, statistical, correlation-based, and feature-based rules. Feature-based rules represent rules that directly generate security event notifications. Once these aggregation rules trigger a threshold, a security event will be generated. Statistical and correlation-based are aggregation types used in combination. The statistical type is responsible for aggregating quantitative relationships based on meta-information (such as network addresses, ports, etc.), and the correlation-based type performs a retrospective analysis based on the triggered statistical type to further aggregate multiple related alarms. The above solution enhances the degree of correlation analysis between alarms and improves the aggregation effect by splitting the alarm aggregation process into statistical and correlation-based types and performing secondary aggregation. However, this solution does not effectively manage the composition of the rules, and there may be problems such as similarity, overlap, and opposition between the rules, resulting in duplicate or suppressed triggering of security events in the end. At the same time, during secondary aggregation, the correlation analysis requires a retrospective analysis of all the alarms included in the current aggregation type, resulting in a large computational overhead and unable to guarantee real-time performance. Using Flink as the streaming engine for processing also cannot dynamically track the aggregation status of triggered security events and cannot dynamically update the status of the remaining security events based on the security events that have been processed by security operation personnel.
[0057] To solve the problems in the current discovery and handling process of network security alarms, such as the single and difficult-to-maintain duplicate alarm aggregation rules, the existence of duplicates and conflicts between alarm aggregation rules, the inability to perform real-time hierarchical aggregation or the large performance overhead of aggregation, and the inability to dynamically change the aggregation status, etc., this application, out of the practical needs of real-time aggregation, hierarchical notification, and disposal of alarms, proposes a network security alarm method to solve the above-mentioned alarm technical problems.
[0058] The following combines Figure 1 , Figure 1 which is a flowchart of a network security alarm method provided by an embodiment of the present invention. The method may include:
[0059] S101: Extract an alarm primitive set from structured knowledge and semi-structured knowledge, and construct an alarm aggregation rule based on the alarm primitive set.
[0060] A knowledge graph is a structured semantic knowledge base used to describe concepts in the physical world and their interrelationships in symbolic form. The basic unit of a knowledge graph is an "entity-relationship-entity" triple, as well as entity and its related attribute-value pairs. Entities are interconnected through relationships to form a networked knowledge structure. A knowledge graph can realize the transformation of the Web from web page links to concept links, support users to retrieve by topic rather than by string, and truly realize semantic retrieval. A search engine based on a knowledge graph can feedback structured knowledge to users in a graphical way, and users can accurately locate and deeply obtain knowledge without having to browse a large number of web pages.
[0061] This embodiment does not limit the data source of the alarm aggregation rule. Generally, the alarm aggregation rule can be extracted from structured knowledge and semi-structured knowledge. Generally, structured knowledge can include target rules constructed by domain experts, structured threat intelligence data, and STIX (Structured Threat Information Expression, a standardized language for describing and sharing cyber threat information) 2.1 threat intelligence data, etc. Semi-structured knowledge can include domain expert experience, security forum and blog content, MITRE (a cybersecurity agency) knowledge base, and national information security vulnerability database, etc.
[0062] For structured knowledge, this embodiment can directly obtain entities, relationships, and attributes through the access algorithm of structured data, and then convert them into alarm meta-languages of relevant security devices through a mapping table, and further form an alarm aggregation rule based on the alarm meta-language set.
[0063] For semi-structured knowledge, this embodiment can use natural language processing technology to extract entities and relevant relationships and attributes in the knowledge, and process them according to the process of semi-structured knowledge after semi-structured conversion.
[0064] For example, for the report content of EternalBlue of CVE-2017-0143 (a high-risk vulnerability), the attack sequence extracted through natural language technology includes: scanning port 443 or 139, brute-forcing port 443 or 139, constructing and transmitting an SMB (Server Message Block, a network sharing protocol) remote overflow attack, exploiting the SMB service vulnerability, and subsequent possible Trojan implantation and lateral movement actions. These alarms are respectively mapped into alarm meta-languages of corresponding devices through a mapping table. For example, scanning port 443 corresponds to the firewall alarm firewall_alert_id of 1, the SMB remote overflow attack corresponds to the IDS (Intrusion Detection Systems) device alarm ids_alert_id of 2, and lateral movement corresponds to the APT (Advanced Persistent Threat) detection device alarm apt_id of 3. Combining these information can obtain an alarm meta-language set.
[0065] This embodiment can construct an alarm aggregation rule based on the extracted alarm meta-language set. Due to the rich knowledge source, this embodiment can construct multiple alarm aggregation rules.
[0066] This embodiment does not limit the specific form of a complete alarm aggregation rule. Generally, it can include: an alarm meta-language set, a condition set, the size of the time window, and the threshold of the target field, etc.
[0067] S102: Construct a rule tree based on the alarm aggregation rules; the root node in the rule tree is the alarm primitive set, the child nodes are the entities in the alarm primitive set, and the parent-child relationship between the nodes is the pre-order of the internal entities in the alarm primitive set.
[0068] Since the knowledge sources that make up the alarm aggregation rules are diverse, there will inevitably be duplication and inclusion relationships between the rules. At the same time, the continuous update of knowledge related to security events means that the alarm aggregation rules need to be continuously updated. In the existing technology, there is no step to merge or establish mutual association relationships between the alarm aggregation rules. Therefore, when new knowledge is added, the established rules cannot be synchronized and updated in a timely and efficient manner. For example, a new attack variant has emerged in the EternalBlue family. From the perspective of handling security events, this aggregation rule should be merged with the existing EternalBlue security alarms, but the existing technology cannot quickly and efficiently implement this function.
[0069] This embodiment can construct a rule tree based on the obtained alarm aggregation rules. The root node in the rule tree is the alarm primitive set, and the child nodes are the entities in the alarm primitive set, such as port scanning, port blasting - 443, and vulnerability exploitation - Payload.
[0070] The parent-child relationship between the nodes is the pre-order of the internal entities in the alarm primitive set. For example, in the EternalBlue security event, there must be an alarm for SMB vulnerability exploitation before the SMB blasting alarm can be aggregated. Therefore, the SMB blasting node in the rule tree needs to be a child node of the SMB vulnerability exploitation node.
[0071] This embodiment can define the rule tree as T=(V, E, R), where V represents the set of nodes in the rule tree, V={v 0 ,v 1 ,…,v n}, v 0 represents the root node (i.e., the alarm primitive set), and v i (i>0) represents each entity node.
[0072] E represents the set of edges, E ⊆V×V, E={(v i ,v j )}, v i is the parent node of v j .
[0073] R is a mapping function from the set of nodes to the set of attributes P, R:V→P, P={p 1 ,p 2 ,…,p m} is the set of attributes of P for any v i ∈V, R(v i ) is the set of attributes of P.
[0074] S103: Determine the similarity between rule trees, cluster the rule trees through a clustering algorithm based on the similarity between trees, and merge the rule trees in each cluster.
[0075] After constructing the rule trees of each alarm aggregation rule, this embodiment can determine the similarity between rule trees, cluster the rule trees through a clustering algorithm based on the similarity between trees, and merge the rule trees in each cluster to obtain the merged rule trees.
[0076] Specifically, the constructed rule trees can be combined into a rule tree set {T 1 , T 2 , …, T n}, and a single rule tree can be represented as T i = (V i , E i , R i ), where the subscript is the number.
[0077] This embodiment does not limit the specific process of determining the similarity between rule trees. Generally, it can be divided into preliminary calculation and accurate calculation. Specifically, node features are extracted from the node attributes of the rule tree based on a feature extraction function; the hash feature value of the node features is determined based on the SimHash algorithm, and the first similarity between rule trees is determined based on the hash feature value; a first similarity threshold is determined, and the rule trees with the first similarity greater than the first similarity threshold are classified to obtain a preliminary screening similar tree set; in the preliminary screening similar tree set, the tree edit distance between rule trees is determined based on the Zhang-Shasha algorithm, and the second similarity between rule trees is determined based on the tree edit distance.
[0078] This embodiment does not limit the specific type of the feature extraction function, which can be determined based on the actual application. Through the feature extraction function F, features F(T i ) can be extracted from the node attributes of the rule tree T i , F(T i ) = {f 1 , f 2 , …, f m}, where f j is the node feature extracted from the node attributes of the rule tree T i .
[0079] SimHash itself belongs to a locality-sensitive hashing algorithm, and the hash signature it generates can represent the similarity of the original content to a certain extent. Therefore, this embodiment can use the SimHash algorithm to determine the hash feature value H(T i ) = SimHash(F(Ti ))。
[0080] This embodiment can determine the first similarity between rule trees based on hash feature values. Specifically, the hash feature values of two rule trees can be input into the first similarity calculation function to obtain the first similarity between the rule trees. This embodiment does not limit the specific type of the similarity function, which can generally be:
[0081] Sim 1 (T i ,T j ) = 1 – HammingDistance(H(T i ), H(T j )) ÷ k;
[0082] In the formula, Sim 1 (T i , T j ) is the first similarity between two rule trees T i and T j . H(T i ) is the hash feature value of T i . HammingDistance is the Hamming distance calculation function. H(T j ) is the hash feature value of T j . k is the number of bits of SimHash.
[0083] After calculating the first similarity between the rule trees, this embodiment can set the first similarity threshold θ 1 , and classify the similar rule trees based on the first similarity threshold to obtain the initially screened similar tree set.
[0084] Specifically, when the first similarity between two rule trees is greater than the first similarity threshold, it can be determined that the above two rule trees are similar rule trees after preliminary screening. The corresponding formula is: Define the total set of initially screened similar trees S = {S 1 , S 2 , …, S p}, where the initially screened similar tree set S q = {T i , T j , …}, satisfying T i , T j ∈ S q , Sim 1 (T i , T j ) ≥ θ 1 .
[0085] After the preliminary screening, this embodiment can further perform precise calculation of the similarity between trees. Specifically, the Zhang-Shasha algorithm is used to calculate the precise tree edit distance for the tree pairs after the preliminary screening. The tree edit distance can represent the minimum number of operations required to convert one tree into another tree, that is, for each set S of initially screened similar trees q In this set, the tree edit distance between each pair of regular trees in the set is calculated using the Zhang-Shasha algorithm, and the second similarity between the regular trees is determined based on the tree edit distance.
[0086] This embodiment does not limit the specific method for determining the second similarity between regular trees based on the tree edit distance. Generally, the tree edit distance can be input into the second similarity calculation function to determine the second similarity. The expression of the second similarity calculation function can be:
[0087] Sim 2 (T i ,T j )=1-(EditDistance(T i ,T j ) / MaxEditDistance);
[0088] In the formula, Sim 2 (T i ,T j ) is the second similarity between two regular trees T i and T j , EditDistance(T i ,T j ) is the tree edit distance between two regular trees T i and T j , and MaxEditDistance is the maximum possible tree edit distance between the regular trees T i and T j .
[0089] In this embodiment, the similarity calculation is performed twice successively mainly considering efficiency. SimHash is a locality-sensitive hashing algorithm that maps high-dimensional features to a low-dimensional space and can quickly judge similarity, but there will be certain misjudgments. Zhang-Shasha is an algorithm for calculating the tree edit distance. The tree edit distance represents the minimum number of edit operations (inserting, deleting, replacing nodes) required to convert one tree into another tree. This distance can more accurately reflect the similarity of the tree structure. In the first stage, SimHash and the Hamming distance are used to calculate the similarity, and the calculation complexity is low, which can quickly filter out trees that are obviously dissimilar. In the second stage, the Zhang-Shasha algorithm is used to calculate the similarity. Although the complexity increases, it can accurately determine the similarity between trees.
[0090] SimHash may produce false positives (judging as similar when actually not), but it is less likely to produce false negatives (judging as not similar when actually similar). The Zhang-Shasha algorithm not only gives the distance but also can give the specific edit operation sequence. The effect of the two-stage filtering depends on the setting of the first similarity threshold of SimHash, and a trade-off between accuracy and efficiency is needed, which can be set based on the actual application scenario.
[0091] Furthermore, after calculating the similarity between the second trees, in this embodiment, a second similarity threshold θ can be set. 2 Based on the similarity between the second trees, the rule trees in each initially screened similar tree set are clustered by a clustering algorithm until the similarity between clusters is lower than the second similarity threshold.
[0092] The second similarity threshold in this embodiment is used to determine whether two rule trees are exactly similar trees. When Sim 2 (T i , T j ) ≥ θ 2 , it can be determined that T i and T j are exactly similar trees.
[0093] Therefore, in this embodiment, the rule trees in the initially screened similar tree set can be clustered based on the second similarity threshold. In each initially screened similar tree set, the hierarchical clustering algorithm is used to group similar trees, and the bottom-up agglomerative hierarchical clustering method is adopted. The clustering stops when the similarity between clusters is lower than this threshold.
[0094] After clustering is completed, in this embodiment, the merging of rule trees in the clusters can be performed. This embodiment does not limit the way of merging rule trees. Generally, a reference tree can be determined in each cluster, and the remaining rule trees in each cluster are aligned with the reference tree based on the tree alignment algorithm; the aligned nodes are merged, and the characteristic branches and characteristic nodes of each rule tree are retained to obtain the merged rule tree.
[0095] This embodiment does not limit the way of determining the reference tree. Generally, the tree with the largest length or the largest number of nodes can be selected. Further, tree alignment algorithms such as the maximum common subtree algorithm can be used to align the other rule trees in the cluster with the base tree, merge the aligned nodes, and retain all unique branches and nodes.
[0096] In this embodiment, if there is a single rule tree that is not a similar rule tree with any other rule trees, there is no need to merge this single rule tree.
[0097] This embodiment does not limit the merging process of rule trees. Generally, the total cluster set J = {J 1 , J2 , …, J p}, for any cluster J q , determine J q the union set V of the node sets of all the rule trees in union = ∪V i |T i ∈ J q .
[0098] For each tree node v in V union , define the support Support(v) = |{T i |v ∈ V i , Ti ∈ S q}| ÷ |S q |.
[0099] In this embodiment, the support Support(v) can represent the proportion of the trees containing the node v in J q .
[0100] Furthermore, a merged tree T merged = (V merged , E merged , R merged ) can be constructed, where V merged = v | v ∈ V union , Support(v) ≥ σ, σ is the node retention threshold, E merged = {(v i , v j ) | T k ∈ J q , (v i , v j ) ∈ E k , v i , v j ∈ V merged , and R merged (v) = MergeAttributes(R i (v) | T i ∈ J q , v ∈ V i ).
[0101] In this embodiment, MergeAttributes is an attribute merging function. Specifically, the MergeAttributes function can represent an empirical process of attribute merging, including but not limited to: when merging the attribute sets of the same node v in multiple rule trees, if the attribute is numerical (such as time window, threshold), take the maximum value (to ensure coverage of all scenarios); if the attribute is enumerative (such as alarm type), take the union. Specifically, the attribute merging method of MergeAttributes can be determined according to the actual empirical process.
[0102] S104: After the merging is completed, based on the rule relationships, the rule trees are associated to generate a directed association graph; the vertices of the directed association graph are the merged rule trees, and the edges are the rule relationships.
[0103] In this embodiment, after the merging of similar rule trees is completed, this embodiment can associate the rule trees based on the rule relationships to generate a directed association graph; the vertices of the directed association graph are the merged rule trees, and the edges are the rule relationships.
[0104] Specifically, this embodiment can construct and maintain the association relationships between these rule trees from the rule relationships extracted from structured knowledge and semi-structured knowledge.
[0105] There may be various associations between rules, such as temporal relationships, causal relationships, combination relationships, etc. Temporal relationships represent the occurrence order between rules, causal relationships indicate that the triggering of some rules depends on the satisfaction of other rules, and combination relationships mean that multiple rules together constitute different aspects of a certain type of security event.
[0106] This embodiment does not limit the specific method of rule relationship extraction and storage. Generally, a directed association graph can be used as the basic data structure. The directed association graph can be represented as G=(M,N), where the vertex M represents the rule tree and the edge N represents the relationship between rules. The implementation of the directed association graph uses an adjacency list structure, and each vertex maintains its out-edge and in-edge lists for quick search of relevant rules.
[0107] Each edge N records the type of rule relationship, and these types can be directly determined by prior knowledge. The relationship types include pre-triggering (requiring a strict temporal sequence), simultaneous triggering (needing to occur together within a specified time window), mutual exclusion (rules cannot be triggered simultaneously), etc. For example, for rule tree T i and T j , if it is known that the behavior described by T i is a necessary prerequisite for T j and must occur within a specific time window, then an edge can be added in the directed association graph from T i to T j , and T i is marked as T jThe pre-trigger relationship.
[0108] Furthermore, in this embodiment, the construction of upper-layer rules can be carried out, and closely related rule subsets can be identified based on the structural characteristics of the directed association graph. By analyzing the relationship types of the edges, the rules with a complete attack link or constituting a specific security event are grouped together. For each rule subset, an upper-layer composite rule is constructed according to the relationship type of the edges, and its trigger logic is determined, including required rules, optional rules, and mutually exclusive rules. This structure can accurately express specific attack scenarios and complete security events. Finally, specific rule content can be converted according to the rule format required by the target analysis engine.
[0109] Furthermore, in the prior art, when new knowledge is added, the established rules cannot be synchronized and updated in a timely and efficient manner. Therefore, in this embodiment, when a new alarm aggregation rule is added, a new rule tree can be constructed based on the new alarm aggregation rule, and the directed association graph can be updated.
[0110] Specifically, when generating a new rule tree, the hash feature values of the new rule tree and the node features in the rule tree are determined based on the SimHash algorithm; the first inter-tree similarity between the new rule tree and the rule tree is determined based on the hash feature values, and the maximum first inter-tree similarity is determined from the first inter-tree similarities; when the maximum first inter-tree similarity is greater than (or equal to) the first similarity threshold, the new rule tree is merged with the corresponding most similar rule tree, and the directed association graph is updated.
[0111] When the maximum first inter-tree similarity is less than the first similarity threshold, the tree edit distance between the rule trees is determined based on the Zhang-Shasha algorithm, and the second inter-tree similarity between the new rule tree and the rule tree is determined based on the tree edit distance; the maximum second inter-tree similarity is determined from the second inter-tree similarities, and when the second inter-tree similarity is greater than (or equal to) the second similarity threshold, the new rule tree is merged with the corresponding most similar rule tree, and the directed association graph is updated.
[0112] To improve the calculation efficiency, when a new rule tree is added, in this embodiment, it can first be preliminarily determined whether the new rule tree can be merged with the existing rule trees through the SimHash algorithm. The hash feature values of the new rule tree and the node features in the rule tree are determined based on the SimHash algorithm, and the first inter-tree similarity between the new rule tree and the rule tree is determined based on the hash feature values.
[0113] In this embodiment, the new rule tree can be merged with the corresponding most similar rule tree, that is, the maximum first inter-tree similarity is determined from the first inter-tree similarities; when the maximum first inter-tree similarity is greater than the first similarity threshold, the new rule tree is merged with the most similar rule tree corresponding to the maximum first inter-tree similarity.
[0114] When it is determined by the SimHash algorithm that there is no tree similar to the new rule tree in the existing tree, the Zhang-Shasha algorithm can be further used for determination. That is, when the maximum similarity between the first trees is less than the first similarity threshold, the tree edit distance between the rule trees is determined based on the Zhang-Shasha algorithm, and the second similarity between the new rule tree and the rule trees is determined based on the tree edit distance.
[0115] In this embodiment, the new rule tree can be merged with the most similar rule tree, that is, the maximum second similarity between the trees is determined from the second similarity between the trees. When the second similarity between the trees is greater than the second similarity threshold, the new rule tree is merged with the most similar rule tree corresponding to the maximum second similarity between the trees, and the directed association graph is updated.
[0116] If the second similarity between the trees is less than the second similarity threshold, it can be considered that there is no similar relationship between the existing rule tree and the new rule tree, and there is no need to merge the new rule tree.
[0117] S105: Obtain alarm data, and aggregate the alarm data based on the rule tree and the directed association graph to generate a network security event.
[0118] In this embodiment, alarm data can be obtained, and the alarm data is aggregated based on the rule tree and the directed association graph to generate a network security event.
[0119] The rule of the security event is the basis for alarm aggregation. By matching the original alarm data with the rule tree and determining whether a security event is triggered, a large number of alarms can be reduced to specific security event notifications. The noise reduction ratio calculation formula: 1 - (the number of events / the number of original alarms) * 100%.
[0120] This embodiment does not limit the specific method of alarm aggregation. An alarm aggregation engine can be created, and the rule statements are loaded into the alarm aggregation engine; alarm data is obtained, and the alarm data is input into the normalizer and shunt of the alarm aggregation engine for processing, so as to shunt the normalized alarm data to the corresponding alarm aggregation processors in the alarm aggregation engine; the normalized alarm data is aggregated based on the alarm aggregation processors to generate a network security event.
[0121] Specifically, the overall process can be as Figure 2 shown, and the engine is mainly composed of the following core parts:
[0122] Data Input Layer: It includes two data channels, namely alarm data access and rule configuration. In the alarm data access channel, the alarm data is processed by a normalizer and a splitter. The output of the splitter is pre-checked by a pre-filter to identify and exclude non-critical or duplicate alarms, reducing the subsequent processing burden and ensuring that major alarms are promptly attended to and processed. In the rule configuration data channel, it is used to load and parse the converted rule content, that is, convert the alarm aggregation rule into a tree structure, including rule API (Application Programming Interface) parsing, syntax parsing, semantic parsing, and AST (Abstract Syntax Tree) construction, etc. The constructed rule tree needs to complete operations such as similarity calculation, aggregation, rule tree merging, and association in the aggregation primitive work pipeline.
[0123] Processing Core Layer: It consists of an alarm aggregation processor. The alarm aggregation processor processes the normalized alarm stream according to the loaded rule content, including the pre-filtering and aggregation logic defined by the rules. The alarm aggregation processor can adjust its processing strategy in real time according to the dynamic update of the rules.
[0124] Aggregation Output Layer: It includes links such as aggregation status maintenance, security event generation, event alarm notification, and event processing result feedback, forming a complete closed-loop for alarm processing.
[0125] In this embodiment, for the convenience of rule maintenance and intuitive display, and to ensure that the rules can be efficiently executed in the engine designed in this chapter, the graph structure and tree structure in this embodiment can be mapped to SQL-Like rule statements. This mapping not only preserves the semantic integrity of the rules but also provides a clear rule expression. The basic syntax structure of the rule statements can include: a WITH clause for defining intermediate variables and aliases within the alarm aggregation rule; a FROM clause for defining the alarm source; a WHERE clause for defining the pre-filtering conditions for alarms; a SELECT clause for field projection and conversion; a GROUP BY clause for grouping alarms according to preset dimensions; and a HAVING clause for defining the trigger conditions for the network security events.
[0126] SQL (Structured Query Language), as a declarative language, can clearly express the intention of data processing, making the semantics of rules easier to understand and maintain. Secondly, each component of the rule naturally corresponds to the processing flow of the alarm aggregation engine. FROM and WHERE correspond to the preprocessing and filtering stages, and GROUP BY and HAVING correspond to the aggregation processing stage. Finally, this structured rule format facilitates the verification, optimization, and transformation of rules. The engine can directly parse the rule statements and construct an efficient execution plan. Through this mapping method, both the expressive power of the rules is ensured, and the efficient execution of the rules in the engine is also ensured. At the same time, the modification and maintenance of the rules become more intuitive and convenient. Security analysts can focus on the semantic design of the rules without having to pay attention to the underlying implementation details.
[0127] The reference structure of the rule is as follows:
[0128] [WITH <alias definition>];
[0129] FROM <alarm source>;
[0130] [WHERE <filter condition>];
[0131] SELECT <projection definition>;
[0132] [GROUP BY <grouping field>];
[0133] [HAVING <security event trigger condition>].
[0134] In this embodiment, the alarm aggregation engine dynamically processes SQL-Like rule statements through a rule loading module (rule configuration data channel). The rule loading module first uses an AST syntax parser to parse the rule statements into an abstract syntax tree. The parsing process strictly follows the rule syntax definition to ensure the integrity of the rule semantics. The obtained AST tree structure reflects the logical order of rule processing: the FROM and WHERE nodes correspond to the alarm source selection and pre-filtering logic, the SELECT node represents the field projection and transformation operations, and the GROUP BY and HAVING nodes correspond to the grouping aggregation and event trigger judgment logic.
[0135] The alarm aggregation engine converts each processing node into a corresponding operator according to the structure of the AST tree and assembles these operators into a complete alarm aggregation processor in the execution order of the rule. The alarm aggregation processor, as the runtime carrier of the rule, is responsible for maintaining the state information during the rule execution process, including grouping count, time window, threshold statistics, etc.
[0136] When the alarm data stream passes through the alarm aggregation processor, the processor updates its internal state according to the logic defined by the rules and triggers a security event when the conditions defined in the HAVING clause are met. This rule loading mechanism based on SQL-Like syntax enables the engine to flexibly adapt to different alarm aggregation scenarios while ensuring the efficiency of rule execution and the reliability of state management.
[0137] This embodiment does not limit the specific process of aggregation. Generally, it can be as Figure 3 shown. First, normalization processing of the input source is performed. In this embodiment, multiple alarm input sources can be connected simultaneously. For example, protection devices such as firewalls, gateways, bastion hosts, and IDSs of different brands in an organization. This embodiment can use pre-trained machine learning models and natural language processing algorithms to extract and normalize the metadata in the original alarm text to unify the input entity types. The normalized data needs to conform to the unified data model defined by the engine to ensure the standardization of subsequent processing processes.
[0138] Furthermore, the input source is shunted. After the alarm data enters the alarm aggregation engine, the alarm aggregation engine shunts the normalized alarm data to the corresponding alarm aggregation processors through the shunt according to the definitions of the FROM and WHERE clauses of the loaded SQL-Like rules. The shunting process takes into account the enabled status and priority of the rules to ensure that the alarm data can be correctly routed to the corresponding alarm aggregation processors.
[0139] The alarm aggregation processor aggregates based on the obtained alarm data. During the aggregation process, alarm processing and status maintenance can be performed. The alarm aggregation processor, as the execution carrier of the rules, is responsible for maintaining various status information during the operation of the rules. These status information include: the alarm count statistically grouped by the grouping dimension defined in the GROUP BY clause, the aggregation window information (such as start time and duration) defined in the HAVING clause, the trigger conditions and processing status of security events, and the prerequisite condition status related to other rules, etc.
[0140] This embodiment can check whether the trigger conditions of the network security event are all in the triggered state through the alarm aggregation processor; if so, the corresponding network security event can be triggered and generated; if not, the alarm aggregation can continue based on the input data of the input source.
[0141] In the network security event trigger process, it includes the check of the trigger status of the previous security event. Specifically, based on the constructed directed association graph, the prerequisite dependency status of the corresponding rule will be checked when processing each alarm. This dependency relationship is declared through the WITH clause of the SQL-Like rule, making the state sharing between rules clearer.
[0142] For example, security events such as port scanning and brute force attacks, as basic events, are often the preconditions for other advanced threat events. When such basic events are triggered, their status will be recorded and shared with all subsequent rules that depend on them, thereby activating the detection logic of these rules and effectively avoiding duplicate calculations.
[0143] During the alarm aggregation process, this embodiment can update and check the aggregation status of the alarm aggregation processor. After the alarms are shunted, they enter the corresponding alarm aggregation processor, and the alarm aggregation processor updates its internal status according to the grouping conditions defined by GROUP BY in the rule and the aggregation conditions in the HAVING clause. These statuses include information such as count statistics, time windows, and threshold accumulations. After the status update is completed, the processor checks whether the trigger conditions defined by the rule are met. If they are met, it enters the network security event generation process; otherwise, it continues to maintain the aggregation status.
[0144] During the generation of network security alarm events, when the alarm aggregation processor confirms that all trigger conditions are met, it will generate corresponding security events. Further, this embodiment can also determine the handling solutions for network security events based on a predefined security knowledge graph, and send the handling solutions to the operation and maintenance side; when receiving the information indicating that the handling of the network security event is completed, it can reset the aggregation status of the alarm aggregation processor and update the aggregation period of the alarm aggregation processor.
[0145] The system automatically retrieves and associates applicable handling solutions and processes through the association of the predefined security knowledge graph in the rule, and notifies these information to the security operation personnel as the extended attributes of the event, improving the response efficiency.
[0146] After the operation personnel complete the handling of the event, they can mark the event as the handled status. At this time, the aggregation processor that triggered the event will reset its internal status and start a new aggregation period. The handled event still maintains its activation effect on subsequent rules because even though the threat has been handled, its indication role as a link in the attack chain is still valid. This processing mechanism ensures the integrity of the attack chain analysis.
[0147] This embodiment processes structured and semi-structured security knowledge using natural language processing, converts it into an alarm aggregation primitive set, and then maintains the meta-information of the aggregation rules by converting the primitive set into an alarm aggregation rule tree. The final aggregation alarm rules are maintained and generated through the tree similarity algorithm, the tree merging algorithm, the subtree query algorithm, as well as the constructed relationship graph and the corresponding graph search algorithm. Compared with the existing solutions, this solution can reduce the duplication between rules through the above technologies, obtain and maintain the logical relationship at the rule level, and reduce the cost of maintaining the existing rules when new security knowledge is added.
[0148] This embodiment compiles through SQL-Like syntax rules, parses and converts the alarm aggregation rules into an alarm aggregation processor through an alarm aggregation engine, and aggregates alarms in real time through the alarm aggregation processor. Compared with some alarm aggregation solutions using knowledge graph or artificial intelligence technologies, this solution has higher real-time performance, can detect and handle alarms in the early stage of the attack chain, and prevent security incidents from further threatening the organization.
[0149] Through a rule tree and a directed association graph that annotates the logical relationship between rules, this solution can conveniently apply technologies such as knowledge graphs to retrieve the handling methods corresponding to the alarm aggregation rules from the knowledge graph and bind them to the corresponding security events, facilitating security operation personnel to quickly handle and respond to threats.
[0150] The design of the alarm aggregation engine of this solution has better characteristics of real-time dynamic state maintenance compared with the implementation of existing solutions using technologies such as Flink (streaming framework). It can use the handling results of security personnel for security incidents to update the triggered security incidents, and use the logical relationship between security incidents to share the aggregation state of basic events, reducing the performance overhead of aggregation operations and further improving real-time performance.
[0151] Based on the above embodiment, the method of the present invention constructs a rule tree through alarm aggregation rules, merges similar rule trees, and associates the merged rule trees based on rule relationships, reducing the duplication between rules, obtaining and maintaining logical relationships at the rule level, and avoiding the occurrence of duplicate or suppressed triggering of security incidents.
[0152] The following combines Figure 4 , Figure 4 is a structural block diagram of a network security alarm device provided by an embodiment of the present invention. The device may include:
[0153] The first module 100 is used to extract an alarm primitive set from structured knowledge and semi-structured knowledge, and construct an alarm aggregation rule based on the alarm primitive set;
[0154] The second module 200 is used to construct a rule tree based on the alarm aggregation rule; the root node in the rule tree is the alarm primitive set, the child nodes are entities in the alarm primitive set, and the parent-child relationship between nodes is the pre-order of entities inside the alarm primitive set;
[0155] The third module 300 is used to determine the inter-tree similarity between each rule tree, cluster the rule trees through a clustering algorithm based on the inter-tree similarity, and merge the rule trees in each cluster;
[0156] The fourth module 400 is used to, after completion of the merging, generate a directed association graph by associating the rule trees based on the rule relationships; the vertices of the directed association graph are the rule trees, and the edges are the rule relationships;
[0157] The fifth module 500 is used to obtain alarm data and aggregate the alarm data based on the rule tree and the directed association graph to generate network security events.
[0158] Based on the above embodiments, the method of the present invention constructs a rule tree through an alarm aggregation rule, merges similar rule trees, and associates the merged rule trees based on the rule relationships, reducing the duplication degree between rules, obtaining and maintaining the logical relationships from the rule level, and avoiding the occurrence of repeated or suppressed triggering of security events.
[0159] Based on the above embodiments, the third module 300 may include:
[0160] The first unit is used to extract node features from the node attributes of the rule tree based on a feature extraction function;
[0161] The second unit is used to determine the hash feature value of the node features based on the SimHash algorithm, and determine the first similarity between the rule trees based on the hash feature value;
[0162] The third unit is used to determine a first similarity threshold, classify the rule trees with the first similarity between the rule trees greater than the first similarity threshold to obtain a preliminary screening similar tree set;
[0163] The fourth unit is used to, in the preliminary screening similar tree set, determine the tree edit distance between the rule trees based on the Zhang-Shasha algorithm, and determine the second similarity between the rule trees based on the tree edit distance;
[0164] The fifth unit is used to determine a second similarity threshold, and cluster the rule trees in each preliminary screening similar tree set based on the second similarity through the clustering algorithm until the similarity between clusters is lower than the second similarity threshold.
[0165] Based on the above embodiments, the third module 300 may include:
[0166] The sixth unit is used to determine a reference tree in each cluster, and align the remaining rule trees in each cluster with the reference tree based on a tree alignment algorithm;
[0167] The seventh unit is used to merge the aligned nodes and retain the feature branches and feature nodes of each rule tree.
[0168] Based on the above embodiments, the device may further include:
[0169] The sixth module is used to determine the hash feature values of the new rule tree and the node features in the rule tree based on the SimHash algorithm when generating a new rule tree;
[0170] The seventh module is used to determine the first inter-tree similarity between the new rule tree and the rule tree based on the hash feature values, and determine the maximum first inter-tree similarity from the first inter-tree similarities;
[0171] The eighth module is used to merge the new rule tree with the corresponding most similar rule tree and update the directed association graph when the maximum first inter-tree similarity is greater than the first similarity threshold;
[0172] The ninth module is used to determine the tree edit distance between the rule trees based on the Zhang-Shasha algorithm and determine the second inter-tree similarity between the new rule tree and the rule tree based on the tree edit distance when the maximum first inter-tree similarity is less than the first similarity threshold;
[0173] The tenth module is used to determine the maximum second inter-tree similarity from the second inter-tree similarities, and merge the new rule tree with the corresponding most similar rule tree and update the directed association graph when the second inter-tree similarity is greater than the second similarity threshold.
[0174] Based on the above embodiments, the fifth module 500 may include:
[0175] The eighth unit is used to convert the rule tree and the directed association graph into SQL-like rule statements;
[0176] The ninth unit is used to create an alarm aggregation engine and load the rule statements into the alarm aggregation engine;
[0177] The tenth unit is used to obtain the alarm data, input the alarm data into the normalizer and shunt of the alarm aggregation engine for processing, so as to shunt the normalized alarm data to the corresponding alarm aggregation processors in the alarm aggregation engine;
[0178] The eleventh unit is used to aggregate the normalized alarm data based on the alarm aggregation processors to generate the network security event.
[0179] Based on the above embodiments, the basic syntax structure of the rule statement includes: a WITH clause, which is used to define intermediate variables and aliases within the alarm aggregation rule; a FROM clause, which is used to define the alarm source; a WHERE clause, which is used to define the pre-filtering conditions for the alarm; a SELECT clause, which is used for field projection and conversion; a GROUP BY clause, which is used to group the alarms according to a preset dimension; and a HAVING clause, which is used to define the triggering conditions for the network security events.
[0180] Based on the above embodiments, the fifth module 500 may further include:
[0181] A twelfth unit, which is used to determine a handling solution for the network security event based on a predefined security knowledge graph and send the handling solution to the operation and maintenance end;
[0182] A thirteenth unit, which is used to update the aggregation period of the alarm aggregation engine when receiving the information indicating that the handling of the network security event is completed.
[0183] Based on the above embodiments, the present invention further provides a network security alarm device, which may include a memory and a processor. Among them, a computer program is stored in the memory. When the processor calls the computer program in the memory, the steps provided in the above embodiments can be implemented. Of course, the device may further include various necessary network interfaces, power supplies, and other components, etc.
[0184] The present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by an execution terminal or a processor, the network security alarm method provided by the embodiments of the present invention can be implemented; the storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store program codes.
[0185] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
Claims
1. A network security alarm method, characterized in that: include: Extracting alarm meta-words from structured knowledge and semi-structured knowledge, and constructing alarm aggregation rules based on the alarm meta-words; Constructing a rule tree based on the alarm aggregation rule; The root node in the rule tree is the alarm primitive set, the child node is the entity in the alarm primitive set, and the node parent-child relationship is the pre-order of the entity in the alarm primitive set; wherein the constructed rule trees are combined into a rule tree set {T1, T2, ..., T n }, a single rule tree is represented by T i =(V i ,E i ,R i ), the subscript is a number, V represents the node set of the rule tree, E represents the edge set, and R is the mapping function from the node set to the attribute set P; Determine the inter-tree similarity between the Ruled Trees, cluster the Ruled Trees using a clustering algorithm based on the inter-tree similarity, and merge the Ruled Trees in each cluster; After the merging is completed, the rule trees are associated based on the rule relationships to generate a directed association graph; the vertices of the directed association graph are the rule trees, and the edges are the rule relationships; Acquire alarm data, and generate a network security event by aggregating the alarm data based on the rule tree and the directed association graph; The step of determining the inter-tree similarity between the rule trees and clustering the rule trees by a clustering algorithm based on the inter-tree similarity comprises: Extracting node features from node attributes of the rule tree based on a feature extraction function; Determine a hash feature value of the node feature based on a SimHash algorithm, and determine a first inter-tree similarity between the rule trees based on the hash feature value; Determine a first similarity threshold, classify the rule trees whose first inter-tree similarities are greater than the first similarity threshold, and obtain a preliminarily screened similar tree set; In the initially screened similar tree set, a tree edit distance between the regular trees is determined based on the Zhang-Shasha algorithm, and a second inter-tree similarity between the regular trees is determined based on the tree edit distance; A second similarity threshold is determined, and the regular trees in each of the preliminarily screened similar tree sets are clustered by the clustering algorithm based on the second inter-tree similarity until the inter-cluster similarity is lower than the second similarity threshold.
2. The network security alarm method according to claim 1, characterized in that: Merging the rule trees in each of the clusters, comprising: Determine a reference tree in each of the clusters, and align the remaining regular trees in each of the clusters with the reference tree based on a tree alignment algorithm; The aligned nodes are merged, and the characteristic branches and characteristic nodes of each of the rule trees are retained.
3. The network security alarm method according to claim 1, characterized in that: Also includes: When generating a new Rule Tree, determining the hash feature values of the new Rule Tree and the node features in the Rule Tree based on the SimHash algorithm; Determine a first inter-tree similarity between the new Ruled Tree and the Ruled Tree based on the Hash feature value, and determine a maximum first inter-tree similarity from the first inter-tree similarities; When the maximum first inter-tree similarity is greater than a first similarity threshold, merging the new Rule-Tree with the corresponding most similar Rule-Tree, and updating the directed association graph; When the maximum first inter-tree similarity is less than the first similarity threshold, determining a tree edit distance between the regular trees based on the Zhang-Shasha algorithm, and determining a second inter-tree similarity between the new regular tree and the regular tree based on the tree edit distance; A maximum second inter-tree similarity is determined from the second inter-tree similarities, and when the second inter-tree similarity is greater than a second similarity threshold, the new Rule-Tree is merged with the corresponding most similar Rule-Tree, and the directed association graph is updated.
4. The network security alarm method according to claim 1, characterized in that: Aggregating the alarm data based on the rule tree and the directed association graph to generate a network security event includes: Converting the rule tree and the directed association graph into SQL-like rule statements; Creating an alarm aggregation engine, and loading the rule statement into the alarm aggregation engine; Acquire the alarm data, and input the alarm data into a normalizer and a diverter of the alarm aggregation engine for processing, so as to divert the normalized alarm data to a corresponding alarm aggregation processor in the alarm aggregation engine; The network security event is generated based on the alarm data aggregated and normalized by the alarm aggregation processor.
5. The network security alarm method according to claim 4, characterized in that: The basic grammatical structure of the rule statement includes: a WITH clause, which is used to define the intermediate variables and aliases within the alarm aggregation rule; a FROM clause, which is used to define the alarm source; a WHERE clause, which is used to define the pre-filtering conditions of the alarm; a SELECT clause, which is used to perform field projection and conversion; a GROUP BY clause, which is used to group the alarms according to preset dimensions; and a HAVING clause, which is used to define the triggering conditions of the network security event.
6. The network security alarm method according to claim 4, characterized in that: Also includes: Determine a disposal plan for the network security incident based on a predefined security knowledge graph, and send the disposal plan to the operation and maintenance end; After receiving the information that the handling of the network security event is completed, the aggregation period of the alarm aggregation engine is updated.
7. A network security alarm device, characterized in that: include: The first module is used to extract alarm meta-words from structured knowledge and semi-structured knowledge, and construct alarm aggregation rules based on the alarm meta-words; The second module is used to construct a rule tree based on the alarm aggregation rule; The root node in the rule tree is the alarm primitive set, the child node is the entity in the alarm primitive set, and the node parent-child relationship is the pre-order of the entity in the alarm primitive set; wherein the constructed rule trees are combined into a rule tree set {T1, T2, ..., T n }, a single rule tree is represented by T i =(V i ,E i ,R i ), the subscript is a number, V represents the node set of the rule tree, E represents the edge set, and R is the mapping function from the node set to the attribute set P; The third module is used to determine the inter-tree similarity between the Rule-Trees, cluster the Rule-Trees by a clustering algorithm based on the inter-tree similarity, and merge the Rule-Trees in each cluster; The fourth module is used for associating the rule trees based on the rule relationships to generate a directed association graph after the merging is completed; the vertices of the directed association graph are the rule trees, and the edges are the rule relationships; A fifth module is used to obtain alarm data, and aggregate the alarm data based on the rule tree and the directed association graph to generate a network security event; The step of determining the inter-tree similarity between the rule trees and clustering the rule trees by a clustering algorithm based on the inter-tree similarity comprises: Extracting node features from node attributes of the rule tree based on a feature extraction function; Determine a hash feature value of the node feature based on a SimHash algorithm, and determine a first inter-tree similarity between the rule trees based on the hash feature value; Determine a first similarity threshold, classify the rule trees whose first inter-tree similarities are greater than the first similarity threshold, and obtain a preliminarily screened similar tree set; In the initially screened similar tree set, a tree edit distance between the regular trees is determined based on the Zhang-Shasha algorithm, and a second inter-tree similarity between the regular trees is determined based on the tree edit distance; A second similarity threshold is determined, and the regular trees in each of the preliminarily screened similar tree sets are clustered by the clustering algorithm based on the second inter-tree similarity until the inter-cluster similarity is lower than the second similarity threshold.
8. A network security alarm device, characterized in that: include: Memory, for storing computer programs; A processor, used to implement the network security alarm method as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the network security alarm method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Log association analysis method for safety management center
CN105119945A
Key value storage method based on log-structured merged tree
CN105468298A