Alarm noise reduction-oriented interpretable correlation analysis method

Through an interpretable correlation analysis method for alarm noise reduction, combined with time, space and semantic similarity, an alarm correlation diagram is constructed and clustered analysis and isolated alarm mergers are carried out, which solves the problems of insufficient dimensions and insufficient interpretability in the traditional method, and efficient and accurate alarm analysis and complex attack detection are achieved.

CN120200888APending Publication Date: 2025-06-24HEILONGJIANG ELECTRIC POWER SCIENCE RESEARCH INSTITUTE +3
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510454020.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The traditional method of interpreting and analysis of alarm events has the problem of insufficient dimensions and insufficient interpretability, which leads to the lack of comprehensiveness of interpretation and analysis of alarm events and the inability to fully capture the complex relationships and potential patterns between events.

Method used

An interpretable correlation analysis method for alarm noise reduction is proposed. Through the multi-step process of alarm collection and preprocessing, key entity annotation, entity recognition model training, and inference stages, combining time, space and semantic similarity, an alarm correlation diagram is constructed, and cluster analysis and isolated alarm merging are performed through the Louvain algorithm.

Benefits of technology

Through multi-dimensional similarity analysis and the natural language processing capabilities of large language inference models, the false alarm rate is reduced, the detection ability of complex attacks is improved, the network security automation early warning capabilities are enhanced, and efficient, accurate and interpretable alarm analysis support is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200888A_ABST
    Figure CN120200888A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable correlation analysis method for alarm noise reduction, and relates to the field of network security and alarm analysis. The problems that a traditional alarm event interpretation and analysis method is insufficient in consideration of dimensionality and insufficient in interpretability are solved. According to the method, the space-time similarity and the semantic similarity are combined, the natural language processing capacity of large language reasoning is utilized, a novel alarm noise reduction and correlation analysis method is constructed, the false alarm rate can be reduced, efficient, accurate and interpretable alarm analysis support is provided, the complex attack detection capacity is improved, and the network security automatic early warning capacity is enhanced. The method is mainly applied to network security alarm analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of network security and alarm analysis. Background Art

[0002] With the deepening of digital transformation and the rapid development of information technology, the scale and complexity of modern network environments are continuously increasing, and network security threats are becoming increasingly diverse and complex. The wide deployment of various security devices and tools (such as intrusion detection systems, firewalls, vulnerability scanners, etc.) has generated a large amount of multi-source heterogeneous alarm logs. These log data have different formats, including structured, semi-structured, and unstructured content, making it extremely difficult to fuse, standardize, and uniformly process the data. Traditional alarm processing methods mainly rely on predefined rules, simple statistical analysis, and manual intervention. Facing the massive and changing alarm data, the following problems exist:

[0003] Insufficient rule dependence and adaptability. Traditional methods screen alarms by manually defining rules. However, new attack methods are constantly evolving, and static rules defined manually are difficult to cover complex attack patterns. For example, in the scenario of cross-device alarm correlation, rules that only rely on time windows or IP address matching are likely to ignore the semantic relevance, resulting in a high false alarm rate. In addition, rule adjustment relies on expert experience and has a long adaptation cycle, making it difficult to cope with the dynamic changes of network topologies.

[0004] Lack of multi-dimensional feature fusion. Existing technologies mostly adopt single-dimensional analysis (such as time series or spatial clustering), lacking comprehensive consideration of time, space, and semantics. For example, the merging of alarm sets based on time windows may mis-correlate irrelevant events, while the correlation analysis that only relies on network topologies may ignore the semantic relevance of cross-domain attacks.

[0005] Defects in handling isolated alarms. Isolated alarms are often misjudged as noise and lost, but new threats may be hidden in isolated alarms. Existing machine learning models (such as random forests, isolation forests) only judge isolation through statistical features, lacking the semantic reasoning ability for alarm texts, resulting in a risk of missed reports.

[0006] Lack of interpretability. Although existing deep learning models (such as CNN, LSTM) perform well in improving detection accuracy, they lack an interpretability mechanism. Security personnel are difficult to understand the decision-making basis of the models, resulting in low trust in handling high-risk alarms. For example, classification models based on neural networks cannot provide visual explanations of attack links, affecting the efficiency of emergency response.

[0007] In summary, traditional alarm event interpretation and analysis methods often rely on a single analysis dimension. For example, they only rely on rule matching, or only consider one or two aspects among time, space, and semantics. There are deficiencies in handling isolated events and interpretability. This limitation leads to the lack of comprehensiveness in the interpretation and analysis of alarm events, and it is unable to fully capture the complex relationships and potential patterns among events. The above problems urgently need to be solved. Summary of the Invention

[0008] The purpose of the present invention is to solve the problems of insufficient consideration dimensions and insufficient interpretability existing in traditional alarm event interpretation and analysis methods. The present invention provides an interpretable association analysis method for alarm noise reduction.

[0009] An interpretable association analysis method for alarm noise reduction, the method includes the following steps:

[0010] S1. Alarm collection and preprocessing: Collect different formats and types of alarm logs of various devices from within the grid area, and preprocess the different formats and types of alarm logs to obtain standardized logs;

[0011] S2. Key entity annotation: Annotate various key entities of each alarm event in the standardized logs to obtain an annotation dataset; the types of key entities include alarm content, and one or more of timestamp, source IP address, destination IP address, port, protocol, domain name, attack type, and threat level;

[0012] S3. Entity recognition model training: Use each alarm event in the annotation dataset as the input of the entity recognition model, and use the annotation corresponding to the alarm event as the true value to train the entity recognition model to obtain a trained entity recognition model;

[0013] S4. Inference stage: Preprocess the different formats and types of alarm logs of various devices in the current sampling period, and identify the key entities of each alarm event in the preprocessed alarm logs through the trained entity recognition model;

[0014] Take all the key entities of each alarm event as an alarm node, calculate the time, space, and semantic similarities between any two alarm nodes, and then calculate the comprehensive similarity between any two alarm nodes;

[0015] Construct an alarm association graph according to the comprehensive similarities between all any two alarm nodes;

[0016] Perform event association analysis on the alarm association graph to obtain the node importance values of each alarm node, sort the node importance values from high to low to obtain a sorted list, and the sorted list guides the Louvain algorithm to perform clustering analysis on all node importance values to obtain multiple alarm sets;

[0017] Determine whether each alarm set is isolated,

[0018] If the result is yes, use the pre-trained large language inference model to perform semantic relevance inference on the alarm contents of the isolated alarm nodes in the isolated alarm set and all alarm nodes in all non-isolated alarm sets, and obtain the association inference results between the isolated alarm nodes and all non-isolated alarm sets;

[0019] When the association inference result indicates that there is a connection between the current isolated alarm set and one of all non-isolated alarm sets, merge the current isolated alarm set into the non-isolated alarm set with which it is connected, and use the pre-trained alarm explanation model combined with the knowledge base to perform interpretation and analysis on the non-isolated alarm set;

[0020] When the association inference result indicates that there is no connection between the current isolated alarm set and all non-isolated alarm sets, use the pre-trained alarm explanation model combined with the knowledge base to perform interpretation and analysis on the current isolated alarm set;

[0021] If the result is no, directly use the pre-trained alarm explanation model combined with the knowledge base to perform interpretation and analysis on the non-isolated alarm set.

[0022] Preferably, the implementation methods for preprocessing alarm logs of different formats and types include:

[0023] Convert alarm logs of different formats and types into a unified format;

[0024] Perform data cleaning on the alarm logs in the unified format to obtain standardized logs.

[0025] Preferably, the implementation methods for data cleaning include removing noise data, duplicate data, and incomplete data.

[0026] Preferably, the pre-trained alarm explanation model performs interpretation and analysis on isolated and non-isolated alarm sets, and the output interpretation and analysis results include a text report composed of the overall threat level of the set, the overall background information of the set, the attack reason, and suggestions.

[0027] Preferably, the implementation methods for calculating the time similarity between any two alarm nodes include:

[0028]

[0029] Among them, a i is alarm node i, a j is alarm node j, is the time similarity between alarm node i and alarm node j, Indicates the total number of windows when the alarm content of alarm node i and the alarm content of alarm node j appear in the same time window during the sampling period. Indicates the total number of windows when the alarm content of alarm node i and the alarm content of alarm node j appear in different time windows, i≠j, i = 1, 2, 3... N, j = 1, 2, 3... N, and N is the total number of alarm nodes.

[0030] Preferably, the implementation method for calculating the spatial similarity between any two alarm nodes is as follows:

[0031] Use the Node2Vec algorithm to represent the alarm nodes as low-dimensional vectors to obtain node vectors, calculate the cosine similarity between any two node vectors, and obtain the spatial similarity between any two alarm nodes.

[0032] Preferably, the implementation method for calculating the semantic similarity between any two alarm nodes is as follows:

[0033] Map the alarm content of the alarm nodes to embedding vectors, calculate the cosine similarity between any two embedding vectors, and use this cosine similarity as the semantic similarity between the two alarm nodes.

[0034] Preferably, the expression for the comprehensive similarity between any two alarm nodes is:

[0035]

[0036] where a i is alarm node i, a j is alarm node j, is the comprehensive similarity between alarm node i and alarm node j, k time is the time similarity coefficient, is the time similarity between alarm node i and alarm node j, k spa is the spatial similarity coefficient, is the spatial similarity between alarm node i and alarm node j, k sem is the semantic similarity coefficient, is the semantic similarity between alarm node i and alarm node j.

[0037] Preferably, the implementation method for constructing the alarm association graph includes:

[0038] Connect two alarm nodes with a comprehensive similarity higher than the threshold to form an alarm association graph; and the connection between the two alarm nodes is used as an out-edge, and the weight of the out-edge is the comprehensive similarity between the two alarm nodes.

[0039] Preferably, the implementation method for obtaining the node importance value of each alarm node is as follows:

[0040]

[0041] Among them, a i is the alarm node i, a j is the alarm node j, PR(a i ) represents the node importance value of the alarm node i, PR(a j ) represents the node importance value of the alarm node i, d is the damping coefficient, N represents the total number of alarm nodes, M(a i ) is the set of all nodes connected to the alarm node i, w ji is the out-edge weight from the alarm node j to the alarm node k; L(a j ) is the set of out-edges corresponding to the alarm node j, is the sum of the weights of all out-edges corresponding to the alarm node j, a jk is the k-th alarm node among the alarm nodes adjacent to the alarm node j, w jk is the out-edge weight from the alarm node j to the alarm node a jk .

[0042] The beneficial effects brought by the present invention are as follows:

[0043] The interpretable association analysis method for alarm explanation proposed by the present invention combines spatio-temporal similarity and semantic similarity, integrates multi-dimensional similarity analysis, utilizes the natural language processing ability of large language reasoning, and makes breakthroughs in alarm noise reduction, association analysis, and interpretability enhancement. It can reduce the false alarm rate, improve the detection ability of complex attacks, and enhance the network security automation early warning ability. The specific technical effects are as follows:

[0044] 1. Multi-dimensional coverage of time, space, and semantics. Traditional methods mostly rely on single-dimensional analysis (such as time window or IP matching), with a high misjudgment rate. A weighted comprehensive similarity model is constructed to comprehensively analyze the spatio-temporal and semantic association characteristics of alarms.

[0045] Specifically, the present invention quantifies the co-occurrence relationship between different alarm contents in the time window through the similarity coefficient, combines the generation of network topology space embedding vectors to calculate the spatial similarity between alarms, and uses the semantic embedding technology of the large language model to calculate the semantic similarity between alarms.

[0046] 2. Construct an alarm association graph for alarm clustering. An alarm association graph is constructed by setting a comprehensive similarity threshold, key nodes (i.e., node importance values) are identified, and the Louvain algorithm is applied for community division to cluster associated alarms into event sets.

[0047] 3. The present invention adopts semantic reasoning and isolated alarm merging. The prior art relies on manual rules or simple classification models for alarm processing, resulting in a high risk of missed alarms. The present invention uses a large language reasoning model in combination with a knowledge base and the alarm content in the existing alarm set to perform semantic reasoning on the current isolated alarm content and explore its potential connection with the known alarm set.

[0048] 4. Natural language explanation and threat grading. When the alarm explanation model generates an alarm explanation, it combines the alarm content and the knowledge base in the alarm set to output a text report including the threat level (such as high / medium / low risk), the reason for the attack, suggestions, and background information.

[0049] In specific applications, the method of the present invention is expected to meet the accuracy and real-time requirements for network security alarm processing in a large-scale data environment, provide efficient, accurate, and interpretable alarm analysis support for security operation personnel, and enhance network security protection capabilities, which has important theoretical significance and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is an interpretable correlation analysis method for alarm noise reduction according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0053] The present invention aims to solve the problem of cross-grid alarm log correlation analysis and significantly improve the detection ability for complex attacks. The method combines time similarity, space similarity, and semantic similarity, and utilizes the natural language processing ability of the LLM (large language model) to perform correlation analysis and interpretation on multi-source heterogeneous network security alarm logs, thereby reducing the false alarm rate and enhancing the network security automation warning ability. To achieve the above object, refer to the following embodiments.

[0054] DETAILED DESCRIPTION OF THE INVENTION I. Refer to Figure 1 In this embodiment, the interpretable correlation analysis method for alarm noise reduction described in this embodiment includes the following steps:

[0055] S1. Alarm collection and preprocessing: Collect different - format and different - type alarm logs from various devices within the grid area, and preprocess the different - format and different - type alarm logs to obtain standardized logs;

[0056] In specific applications, the various devices are various security devices and scanning tools (such as intrusion detection systems, firewalls, vulnerability scanning tools, etc.). The different - format and different - type multi - source heterogeneous alarm log data collected from the various devices are preprocessed to obtain standardized logs for convenient subsequent processing;

[0057] S2. Key entity annotation: Annotate various key entities of each alarm event in the standardized logs to obtain an annotation dataset; The types of key entities include alarm content, and one or more of timestamp, source IP address, destination IP address, port, protocol, domain name, attack type, and threat level;

[0058] These annotation data provide rich features for subsequent model training;

[0059] S3. Entity recognition model training: Use each alarm event in the annotation dataset as the input of the entity recognition model, and use the annotation corresponding to the alarm event as the ground truth to train the entity recognition model to obtain a trained entity recognition model;

[0060] S4. Inference stage: Preprocess the different - format and different - type alarm logs of various devices in the current sampling period, and use the trained entity recognition model to recognize the key entities of each alarm event in the preprocessed alarm logs;

[0061] Take all the key entities of each alarm event as an alarm node, calculate the time, space, and semantic similarities between any two alarm nodes, and then calculate the comprehensive similarity between any two alarm nodes;

[0062] Construct an alarm association graph according to the comprehensive similarities between all pairs of alarm nodes;

[0063] Perform event association analysis on the alarm association graph to obtain the node importance values of each alarm node, sort the node importance values from high to low to obtain a sorted list, and the sorted list guides the Louvain algorithm to perform clustering analysis on all node importance values to obtain multiple alarm sets;

[0064] Judge whether each alarm set is isolated,

[0065] If the result is yes, use the pre - trained large - language inference model to perform semantic relevance inference on the alarm content of the isolated alarm nodes in the isolated alarm set and all alarm nodes in all non - isolated alarm sets to obtain the association inference results between the isolated alarm nodes and all non - isolated alarm sets;

[0066] When the correlation inference result indicates that there is a connection between the current set of isolated alarms and one of the sets of all non-isolated alarms, merge the current set of isolated alarms into the non-isolated alarm set with which it is connected, and use the pre-trained alarm explanation model combined with the knowledge base to perform explanatory analysis on this non-isolated alarm set;

[0067] When the correlation inference result indicates that there is no connection between the current set of isolated alarms and all non-isolated alarm sets, use the pre-trained alarm explanation model combined with the knowledge base to perform explanatory analysis on the current set of isolated alarms;

[0068] The result is no, and directly perform explanatory analysis on the non-isolated alarm set through the pre-trained alarm explanation model combined with the knowledge base.

[0069] In this embodiment, by combining temporal similarity, spatial similarity, and semantic similarity, the correlation relationship between alarm events is constructed.

[0070] Furthermore, conduct in-depth analysis of isolated alarms, mine their potential associations with the clusters of existing non-isolated alarm sets, and perform wandering merges, thereby improving the accuracy and integrity of alarm aggregation. Further utilize the natural language processing capabilities of large language inference models to generate clear alarm explanations to help security personnel understand the background and potential threats of alarms.

[0071] For isolated alarms that cannot be directly associated through similarity metrics, distill the large language model (LLM) to obtain a pre-trained large language inference model, and conduct in-depth analysis in combination with the knowledge base to determine the relevance. And the pre-trained alarm explanation model receives the feature information of isolated alarm nodes, the feature information of multiple known non-isolated alarm sets, as well as relevant rules and cases in the knowledge base. Based on these inputs, the pre-trained alarm explanation model directly determines whether an isolated alarm is connected to a certain known non-isolated alarm set through semantic understanding and reasoning, so as to determine whether it belongs to that certain known non-isolated alarm set. And both the pre-trained large language inference model and the alarm explanation model used can be implemented through existing technologies. Both can achieve knowledge distillation by fine-tuning the large language model (LLM), enabling them to have a customized large language inference model and a pre-trained alarm explanation model with alarm explanation and classification capabilities.

[0072] In terms of the construction and utilization of the security knowledge base, collect and organize security knowledge such as attack feature libraries and vulnerability information libraries to build a structured knowledge base. During the alarm analysis and explanation process, combine the information in the knowledge base to enhance the reasoning ability and explanation effect of the model. According to new security threat intelligence and attack patterns, update the content of the knowledge base in a timely manner to maintain the timeliness of knowledge and ensure that the model can adapt to the constantly changing security environment.

[0073] This embodiment constructs a novel alarm noise reduction and correlation analysis method, which can reduce the false alarm rate, improve the detection ability of complex attacks, and enhance the network security automation warning ability.

[0074] When specifically applied, it is expected to meet the accuracy and real-time requirements of network security alarm processing in a large-scale data environment, provide efficient, accurate, and interpretable alarm analysis support for security operation personnel, enhance the network security protection ability, and have important theoretical significance and practical application value.

[0075] Furthermore, the implementation methods for preprocessing alarm logs in different formats and types include:

[0076] Convert alarm logs in different formats and types into a unified format;

[0077] Clean the data of the alarm logs in the unified format to obtain standardized logs. When specifically applied, the implementation methods of data cleaning include removing noise data, duplicate data, and incomplete data.

[0078] In this preferred embodiment, multi-source heterogeneous data is first unified in format and then data-cleaned to improve data quality, and key features are extracted from the cleaned data to provide a basis for subsequent analysis.

[0079] Furthermore, the pre-trained alarm explanation model performs interpretation and analysis on the isolated and non-isolated alarm sets, and the output interpretation and analysis results include a text report composed of the overall threat level of the set, the overall background information of the set, the attack reason, and suggestions.

[0080] In this preferred embodiment, an alarm explanation and visualization text report is included. The pre-trained alarm explanation model obtained through a large language model (LLM) is used to generate natural language explanations for the alarms, describing the potential impacts, possible reasons, and recommended handling measures of the alarms. When specifically applied, visualization tools (such as Matplotlib, Plotly) can also be used to display alarm correlation graphs, importance scores, and classification results, providing intuitive information for security analysts.

[0081] Specifically, the implementation methods for calculating the time similarity between any two alarm nodes include:

[0082]

[0083] where a i is alarm node i, a j is alarm node j, is the time similarity between alarm node i and alarm node j, Denotes the total number of windows when the alarm content of alarm node i and the alarm content of alarm node j appear in the same time window within the sampling period. Denotes the total number of windows when the alarm content of alarm node i and the alarm content of alarm node j appear in different time windows, i≠j, i = 1, 2, 3... N, j = 1, 2, 3... N, and N is the total number of alarm nodes.

[0084] In this preferred embodiment, an implementation method for obtaining time similarity is given, which measures the overlapping degree of two alarms in multiple time windows. Through this method, the similarity between two alarms in time can be quantified.

[0085] Specifically, the implementation method for calculating the spatial similarity between any two alarm nodes is as follows:

[0086] Use the Node2Vec algorithm to represent alarm nodes as low-dimensional vectors to obtain node vectors, calculate the cosine similarity between any two node vectors, and obtain the spatial similarity between any two alarm nodes.

[0087] In this preferred embodiment, spatial similarity uses the Node2Vec algorithm to embed the network topology structure of alarm nodes, represents the key entities (such as IP addresses, etc.) of alarm nodes in the network as low-dimensional vectors, and obtains the spatial similarity between alarms by calculating the cosine similarity between node vectors.

[0088] Specifically, the implementation method for calculating the semantic similarity between any two alarm nodes is as follows:

[0089] Map the alarm content of alarm nodes to embedding vectors, calculate the cosine similarity between any two embedding vectors, and use this cosine similarity as the semantic similarity between two alarm nodes. In specific applications, a pre-trained word embedding model (such as bge-zh-1.5) can be used to map the alarm content to embedding vectors.

[0090] Next, set weight coefficients for time, space, and semantic similarities respectively, and perform weighted summation of the three similarities according to the set weights. The expression for the comprehensive similarity between any two alarm nodes is:

[0091]

[0092] Among them, a i is alarm node i, a j is alarm node j, is the comprehensive similarity between alarm node i and alarm node j, k time is the time similarity coefficient, is the time similarity between alarm node i and alarm node j, k spa is the spatial similarity coefficient, is the spatial similarity between alarm node i and alarm node j, and k sem is the semantic similarity coefficient, is the semantic similarity between alarm node i and alarm node j.

[0093] Furthermore, the implementation method for constructing the alarm association graph includes:

[0094] Connect two alarm nodes with a comprehensive similarity higher than the threshold to form an alarm association graph; and the connection between the two alarm nodes is used as an out-edge, and the weight of the out-edge is the comprehensive similarity between the two alarm nodes.

[0095] Apply graph algorithms to the alarm association graph for event correlation analysis. Specifically, the implementation method for obtaining the node importance value of each alarm node through event correlation analysis of the alarm association graph is:

[0096]

[0097] Among them, a i is alarm node i, a j is alarm node j, PR(a i ) represents the node importance value of alarm node i, PR(a j ) represents the node importance value of alarm node i, d is the damping coefficient, N represents the total number of alarm nodes, M(a i ) is the set of all nodes connected to alarm node i, w ji is the out-edge weight from alarm node j to alarm node k; L(a j ) is the set of out-edges corresponding to alarm node j, is the sum of the weights of all out-edges corresponding to alarm node j, a jk is the k-th alarm node among the alarm nodes adjacent to alarm node j, w jk is the out-edge weight from alarm node j to alarm node a jk .

[0098] In this preferred embodiment, the importance of each alarm node can be evaluated more accurately, thereby identifying key alarm events. Subsequently, the Louvain algorithm is used to partition the alarm association graph into communities, clustering the interrelated alarm nodes into the same community, representing potential security events or attacks.

[0099] Although the present invention has been described herein with reference to particular embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Accordingly, it should be understood that numerous modifications may be made to the exemplary embodiments, and other arrangements may be devised, without departing from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein may be combined in ways different from those described in the original claims. It should also be understood that the features described in connection with separate embodiments may be used in other described embodiments.

Claims

1. An interpretable correlation analysis method for alarm noise reduction, characterized in that: The method comprises the following steps: S1. Alarm collection and preprocessing: Collect alarm logs of different formats and types from various devices in the grid area, and preprocess the alarm logs of different formats and types to obtain standardized logs; S2. Key entity annotation: Annotate multiple key entities of each alarm event in the standardized log to obtain an annotation data set; key entity types include alarm content, and one or more of timestamp, source IP address, destination IP address, port, protocol, domain name, attack type and threat level; S3, entity recognition model training: using each alarm event in the labeled data set as the input of the entity recognition model, and using the label corresponding to the alarm event as the true value to train the entity recognition model, to obtain a trained entity recognition model; S4, reasoning stage: pre-process the alarm logs of different formats and types of various devices in the current sampling period, and identify the key entities of each alarm event in the pre-processed alarm logs through the trained entity recognition model; All key entities of each alarm event are taken as an alarm node. After calculating the time, space and semantic similarity between any two alarm nodes, the comprehensive similarity between the two alarm nodes is calculated. Construct an alarm association graph based on the comprehensive similarity between any two alarm nodes; Perform event association analysis on the alarm association graph to obtain the node importance value of each alarm node, and sort the node importance values ​​from high to low to obtain a sorting table. The sorting table guides the Louvain algorithm to perform cluster analysis on all node importance values ​​to obtain multiple alarm sets; Determine whether each alarm set is isolated. The result is yes. The pre-trained large language reasoning model is used to perform semantic relevance reasoning on the alarm contents of the isolated alarm node in the isolated alarm set and all the alarm nodes in all the non-isolated alarm sets, and the relevance reasoning results between the isolated alarm node and all the non-isolated alarm sets are obtained. When the result of association reasoning is that the current isolated alarm set is associated with one of all non-isolated alarm sets, the current isolated alarm set is merged into the non-isolated alarm set that is associated with it, and the non-isolated alarm set is interpreted and analyzed by combining the pre-trained alarm interpretation model with the knowledge base; When the result of association reasoning is that the current isolated alarm set has no connection with all non-isolated alarm sets, the current isolated alarm set is explained and analyzed by combining the pre-trained alarm explanation model with the knowledge base; If the result is no, the pre-trained alarm interpretation model is combined with the knowledge base to directly interpret and analyze the non-isolated alarm set.

2. The interpretable correlation analysis method for alarm noise reduction according to claim 1 is characterized in that: The implementation methods for preprocessing alarm logs of different formats and types include: Convert different formats and types of alarm logs into a unified format; Perform data cleaning on the alarm logs in a unified format to obtain standardized logs.

3. The interpretable correlation analysis method for alarm noise reduction according to claim 2 is characterized in that: Data cleaning can be achieved by removing noisy data, duplicate data, and incomplete data.

4. The interpretable correlation analysis method for alarm noise reduction according to claim 1 is characterized in that: The pre-trained alarm interpretation model interprets and analyzes isolated and non-isolated alarm sets, and the output interpretation and analysis results include a text report consisting of the overall threat level of the set, the overall background information of the set, the cause of the attack, and suggestions.

5. The interpretable correlation analysis method for alarm noise reduction according to claim 1 is characterized in that: The implementation methods of calculating the time similarity between any two alarm nodes include: Among them, a i is the alarm node i, a j is the alarm node j, is the time similarity between alarm node i and alarm node j, Indicates the total number of windows when the alarm content of alarm node i and the alarm content of alarm node j appear in the same time window during the sampling period. Indicates the total number of windows in which the alarm content of alarm node i and the alarm content of alarm node j appear in different time windows, i≠j, i=1,2,3……N, j=1,2,3……N, N is the total number of alarm nodes.

6. The interpretable correlation analysis method for alarm noise reduction according to claim 1 is characterized in that: The implementation method of calculating the spatial similarity between any two alarm nodes is: The Node2Vec algorithm is used to represent the alarm node as a low-dimensional vector to obtain the node vector. The cosine similarity between any two node vectors is calculated to obtain the spatial similarity between any two alarm nodes.

7. The interpretable correlation analysis method for alarm noise reduction according to claim 1 is characterized in that: The implementation method of calculating the semantic similarity between any two alarm nodes is: The alarm content of the alarm node is mapped into an embedding vector, the cosine similarity between any two embedding vectors is calculated, and the cosine similarity is used as the semantic similarity between the two alarm nodes.

8. The interpretable correlation analysis method for alarm noise reduction according to claim 1, characterized in that: The expression of the comprehensive similarity between any two alarm nodes is: Among them, a i is the alarm node i, a j is the alarm node j, is the comprehensive similarity between alarm node i and alarm node j, k time is the temporal similarity coefficient, is the time similarity between alarm node i and alarm node j, k spa is the spatial similarity coefficient, is the spatial similarity between alarm node i and alarm node j, k sem is the semantic similarity coefficient, is the semantic similarity between alarm node i and alarm node j.

9. The interpretable correlation analysis method for alarm noise reduction according to claim 1, characterized in that: The implementation methods of constructing an alarm correlation graph include: Two alarm nodes whose comprehensive similarity is higher than a threshold are connected to form an alarm association graph; and the connection between the two alarm nodes is used as an outgoing edge, and the weight of the outgoing edge is the comprehensive similarity between the two alarm nodes.

10. The interpretable correlation analysis method for alarm noise reduction according to claim 9, characterized in that: The implementation method of obtaining the node importance value of each alarm node is: Among them, a i is the alarm node i, a j is the alarm node j, PR(a i ) represents the node importance value of the alarm node i, PR(a j ) represents the node importance value of the alarm node i, d is the damping coefficient, N represents the total number of alarm nodes, M(a i ) is the set of all nodes connected to the alarm node i, w ji is the outgoing edge weight from alarm node j to alarm node k; L(a j ) is the outgoing edge set corresponding to the alarm node j, is the sum of all outgoing edge weights corresponding to the alarm node j, a jk is the kth alarm node among the alarm nodes adjacent to the alarm node j, w jk From alarm node j to alarm node a jk The outgoing edge weight.

Citation Information

Cited By

  • Security alarm information processing method and device based on multi-agent cooperation

    CN120378229A