Alarm information studying and judging noise reduction method

By building an alarm map and a multi-dimensional evaluation model, combined with whitelist management, the problem of existing technologies being unable to distinguish between real threats and irrelevant alarms is solved, and the accuracy and efficiency of alarm analysis are improved.

CN120743691APending Publication Date: 2025-10-03HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510646120.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing technologies cannot effectively distinguish between real threats and irrelevant alarms, resulting in low alarm accuracy and inability to accurately identify complex security threats.

Method used

By constructing an alarm graph, extracting local context subgraphs and feature vectors, and using information entropy analysis, custom rule evaluation, and large language semantic evaluation models to conduct multi-dimensional evaluation of alarm clusters, and combining whitelist dynamic management to identify and filter out irrelevant alarms.

Benefits of technology

It improves the accuracy and efficiency of alarm analysis, can more accurately identify high-threat alarms, and reduce the workload of operation and maintenance personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743691A_ABST
    Figure CN120743691A_ABST
Patent Text Reader

Abstract

The invention discloses an alarm information research and judgment noise reduction method. According to the method, multiple pieces of alarm information with internal relation are combined into a structured alarm cluster through deep mining and by utilizing context association information between alarm events, so that redundant information is greatly compressed, and the alarm signal-to-noise ratio is improved. And multi-dimensional and comprehensive evaluation is carried out on the alarm cluster by fusing three mechanisms of information entropy analysis, customization rule evaluation and semantic evaluation of a large model, so that high threat alarms can be identified more accurately. Wherein the information entropy analysis helps to identify potential unknown threats by quantifying the uncertainty of the alarm, and the large model semantic evaluation can deeply analyze the alarm content, so that the accuracy and efficiency of alarm analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of alarm information noise reduction, and in particular to an alarm information analysis and noise reduction method. Background Art

[0002] Traditional solutions for dealing with alert fatigue typically use static rules and simple filtering mechanisms to reduce the number of alerts. These solutions rely on setting thresholds, time windows, and frequency limits to block or merge duplicate, irrelevant, or low-priority alerts, thereby reducing the workload of operations personnel.

[0003] The limitations of traditional methods primarily lie in false positives and missed negatives. Because these methods rely on static rules and simple filtering mechanisms, they often fail to accurately identify complex security threats, resulting in low alert accuracy and an inability to effectively distinguish irrelevant alerts from real threats. Summary of the Invention

[0004] The present invention solves the technical problem in the prior art of being unable to effectively distinguish irrelevant alarms from real threats by providing a method for analyzing and reducing noise of alarm information, thereby achieving the technical effect of improving the accuracy and efficiency of alarm analysis.

[0005] The present invention provides a method for analyzing and reducing noise of alarm information, comprising: Obtain target alarm information; the characteristics of the target alarm information include at least: an alarm event, a source IP address, a destination IP address and a destination port number in the alarm event, an alarm payload, and an alarm time; Constructing an alarm graph based on the target alarm information; the nodes of the alarm graph are features in the target alarm information, and the edges of the nodes of the alarm graph are the relationships between the nodes; Extract all nodes and the edges between them within N hops around the node of the alarm graph to form a local context subgraph of each node; According to a predefined path pattern, biased or unbiased sampling is performed in the alarm graph to obtain a series of node path sequences representing specific context logic; Generate a feature vector of each node from the local context subgraph and each node path sequence; The local context density ρi of the feature vector vi is quantified by counting the number of other feature vectors falling into the neighborhood with a radius of εi centered on the feature vector vi; Calculate the relative distance δi between the feature vector vi and the feature vector with the closest local context density ρi; If the local context density ρi and the relative distance δi are both greater than or equal to the preset threshold, the feature vector vi corresponding to the local context density ρi and the relative distance δi is used as the cluster center, and an independent alarm cluster is maintained; For the feature vector vj that is not the cluster center, assign it to the alarm cluster maintained by the cluster center closest to the feature vector vj among all feature vectors whose local context density is greater than the local context density of the feature vector vj; Each alarm cluster is input into the constructed information entropy evaluation model, custom rule evaluation model and large language semantic evaluation model respectively, and the output scores of the constructed information entropy evaluation model, custom rule evaluation model and large language semantic evaluation model are weighted and fused to obtain the final alarm evaluation result.

[0006] Specifically, it also includes: Calculate the distance L between the feature vector vi and its Kth nearest neighbor feature vector, and use the distance L or a multiple of the distance L as the radius εi.

[0007] Specifically, obtaining target warning information includes: Get the original alarm information set; Matching the original alarm information in the original alarm information set with a whitelist template in a preset whitelist library; If the match is successful, the original alarm information that is successfully matched is removed from the original alarm information set, and the remaining original alarm information is retained as the target alarm information.

[0008] Specifically, after removing the successfully matched original alarm information from the original alarm information set, the method further includes: The remaining original alarm information is standardized, and multiple pieces of standardized alarm information within a preset time window are deduplicated, and the remaining alarm information after deduplication is used as the target alarm information.

[0009] Specifically, it also includes: If it is determined that the original alarm information X in the original alarm information set is a false alarm, extract the feature x of the original alarm information X, add the feature x to the preset whitelist library, and assign an initial weight and a survival time; If the feature x is matched once, the weight of the feature x increases, and the survival time of the feature x is extended accordingly; If the feature x is not matched within the preset time, the weight of the feature x is reduced and the survival time of the feature x is shortened accordingly.

[0010] Specifically, inputting each alarm cluster into the constructed information entropy evaluation model includes: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters in which the number of alarm information in each alarm cluster is less than the minimum size threshold, or the average local context density of the alarm information in the cluster is less than the comparison threshold, are regarded as outlier clusters; The cluster is input into the information entropy evaluation model, and the Shannon entropy H(X|C) of the feature X of the alarm information in the cluster is obtained by outputting the formula H(X|C) = -Σᵢ [P(xi) *log2(P(xi))]; wherein P(xi) is the probability that the value xi of the feature X of the alarm information appears in the cluster; the intra-cluster feature entropy score is output according to the Shannon entropy H(X|C) and the preset rules; and / or, the time distribution entropy H(T|C) of the cluster C is obtained by outputting the formula H(T|C) = -Σⱼ [P(tj) * log2(P(tj))]; wherein P(tj) is the probability that the alarm information falls into the preset time bucket; the intra-cluster time distribution entropy score is output according to the time distribution entropy H(T|C) and the preset rules; and / or, the time distribution entropy score of the cluster is output according to the formula H(Patterns|ΔT) = - Σ[P(pattern_k|ΔT) * log2(P(pattern_k|ΔT))] outputs the cluster evolution entropy H(Patterns|ΔT) of cluster C; where P(pattern_k) is the frequency of occurrence of clusters with the preset pattern k, and ΔT is the set time window; based on the cluster evolution entropy H(Patterns|ΔT) and preset rules, the cluster evolution entropy score is output; The outlier cluster is input into the information entropy evaluation model, and the rarity I(yi) of the characteristic value yi of the alarm information in the outlier cluster is output through the formula I(yi) = -log2(P_global(yi)); wherein P_global(yi) is the global occurrence probability of the characteristic value yi of the alarm information in the outlier cluster in the entire alarm data set or a related global historical alarm data range; the alarm entropy score of the outlier cluster is output based on the rarity I(yi) and preset rules.

[0011] Specifically, it also includes: The P(xi) is calculated by the formula P(xi) = count(xi) / N; wherein count(xi) is the frequency of the value xi of the feature X in the cluster C, and N is the number of alarm information in the cluster C; The P(tj) is calculated by the formula P(tj) = nj / N, where nj is the number of alarm information falling into the j-th time bucket Δtj.

[0012] Specifically, the step of inputting each alarm cluster into the constructed custom rule evaluation model includes: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters in which the number of alarm information in each alarm cluster is less than the minimum size threshold, or the average local context density of the alarm information in the cluster is less than the comparison threshold, are regarded as outlier clusters; The clusters and the outlier clusters are respectively input into the custom rule evaluation model, and the custom rule alarm scores of the clusters and the outlier clusters are respectively output according to their respective custom rules.

[0013] Specifically, the step of inputting each alarm cluster into the constructed large language semantic evaluation model includes: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters in which the number of alarm information in each alarm cluster is less than the minimum size threshold, or the average local context density of the alarm information in the cluster is less than the comparison threshold, are regarded as outlier clusters; The clusters and the outlier clusters are respectively input into the large language semantic evaluation model, and large model semantic alarm evaluation scores of the clusters and the outlier clusters are respectively output according to their respective prompt words.

[0014] Specifically, the output scores of the constructed information entropy evaluation model, the custom rule evaluation model, and the large language semantic evaluation model are weighted and integrated to obtain the final alarm evaluation result, including: Normalizing the output scores of the constructed information entropy evaluation model, the custom rule evaluation model, and the large language semantic evaluation model to obtain each normalized alarm score; A weighted sum is performed on the normalized alarm scores to obtain the final alarm evaluation result.

[0015] One or more technical solutions provided in the present invention have at least the following technical effects or advantages: 1. By deeply mining and leveraging the contextual associations between alarm events, multiple, intrinsically connected alarms are consolidated into structured alarm clusters, significantly reducing redundant information and improving the signal-to-noise ratio. Furthermore, by integrating three mechanisms—information entropy analysis, customized rule evaluation, and large-scale semantic evaluation—this system performs a multi-dimensional, comprehensive assessment of alarm clusters, enabling more accurate identification of high-threat alarms. Information entropy analysis quantifies the uncertainty of alarms, helping to identify potential unknown threats, while large-scale semantic evaluation provides in-depth analysis of alarm content, improving the accuracy and efficiency of alarm analysis.

[0016] 2. Whitelists effectively identify and exclude known harmless or common false alarms, significantly reducing the burden of alarm processing. The whitelist database is managed, dynamically updated, and flexibly configured, enabling precise filtering of irrelevant alarms based on actual conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flowchart of the method for analyzing and reducing noise of alarm information provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The embodiment of the present invention solves the technical problem in the prior art of being unable to effectively distinguish irrelevant alarms from real threats by providing a method for analyzing and reducing noise of alarm information, thereby achieving the technical effect of improving the accuracy and efficiency of alarm analysis.

[0019] The technical solution in the embodiment of the present invention is to solve the above technical problems, and the overall idea is as follows: This solution primarily consists of three phases: alarm whitelist matching, alarm filtering, and alarm analysis. The alarm whitelist matching phase uses an updated whitelist to identify and filter out known harmless or common false alarms, effectively reducing the burden of subsequent processing. This phase also receives feedback from the subsequent alarm analysis phase and dynamically updates the whitelist based on actual operational conditions to improve filtering accuracy and effectiveness. The alarm filtering phase begins with alarm standardization, converting alarms from different devices and categories into a unified format to ensure data consistency and lay the foundation for efficient processing. This phase also includes merging alarms with similar timeframes, eliminating invalid alarms of specific types, filtering alarms marked as successful attacks, and performing alarm aggregation to optimize alarm processing and improve accuracy. Identifying and filtering duplicate alarms within a short period of time significantly reduces the volume of data. By filtering and eliminating non-offensive or false alarms, the noise from irrelevant alarms is effectively reduced. Through aggregation strategies based on IP address, attack type, and alert payload, similar alerts are merged to form a more centralized alert information, facilitating rapid and accurate processing by the large-scale model. Finally, the filtered and aggregated alert information enters the alert analysis phase. This phase integrates three evaluation mechanisms: a large language semantic evaluation model, an information entropy evaluation model, and a custom rule evaluation model to comprehensively assess the potential risk and urgency of each alert. Alerts exceeding pre-set thresholds are marked as requiring further attention, ensuring that the security team prioritizes higher-risk incidents.

[0020] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0021] like Figure 1 As shown, the method for analyzing and reducing noise of alarm information provided by the embodiment of the present invention includes: Step S110: Obtain target alarm information; the characteristics of the target alarm information include at least: alarm event, source IP address, destination IP address and destination port number in the alarm event, alarm payload, and alarm time; In order to improve the efficiency of warning information analysis and noise reduction in the embodiment of the present invention, obtaining target warning information includes: Get the original alarm information set; Matching the original alarm information in the original alarm information set with the whitelist template in the preset whitelist library; If the match is successful, the original alarm information that is successfully matched is removed from the original alarm information set, and the remaining original alarm information is retained as the target alarm information.

[0022] Specifically, the whitelist library is developed using Python and Redis. A WhitelistManager class manages the whitelist library. The WhitelistManager class defines methods for adding, deleting, querying, and matching whitelist items. Because the order of alarm signatures is restricted, each identical whitelist item is uniquely keyed and stored as a Redis hash. Various methods of the class are used to maintain the whitelist library. When an alarm passes the whitelist, its signature is extracted and the whitelist matching method is called to determine whether the alarm is allowed to pass.

[0023] The whitelist library provides flexible configuration options for operators, allowing them to build whitelists that meet specific business needs based on different alarm factors. A variety of whitelisting conditions can be set, including but not limited to the following factors: Alarm source: Identify the alarm source. For example, alarms triggered by certain known and trusted monitoring systems or firewalls can be added to the whitelist.

[0024] IP address: You can set up a whitelist for specific internal or trusted IP addresses to prevent false positives.

[0025] Port number: Some known ports may not pose a threat in specific scenarios and can be configured based on this.

[0026] Alarm type: Based on specific alarm types (such as non-security behavioral alarms or frequent false alarms), operations and maintenance personnel can add whitelisting filters.

[0027] Attack results: For certain attack scenarios, you can set a whitelist based on the attack results, such as for unsuccessful attacks.

[0028] Specific instructions for maintaining the whitelist database include: If it is determined that the original alarm information X in the original alarm information set is a false alarm, extract the feature x of the original alarm information X, add the feature x to the preset whitelist library, and assign an initial weight and survival time; Specifically, if the model analyzes certain alarms and determines they are false positives or non-threatening alerts generated by normal business operations, these false positives are added to the whitelist database to prevent similar alarms from triggering security alerts in the future. Furthermore, all alarm features are extracted from the alarm information and a unique key is formed from the tuple. A JSON value is generated for each whitelist entry, which stores the number of whitelist hits and the time of the most recent hits. Finally, these whitelist entries are automatically stored in Redis.

[0029] To ensure the validity and timely updates of the whitelist, this embodiment of the present invention employs an automated maintenance mechanism that combines weighting and TTL (Time-to-Live) metrics. This mechanism ensures that entries in the whitelist can be adjusted based on actual conditions, preventing the accumulation of redundant data.

[0030] If feature x is matched once, its weight increases, and its lifetime is extended accordingly. Specifically, each whitelist entry is assigned to different weight groups based on its number of matches. A greater number of matches indicates a more enduring value for the whitelist entry, and its weight increases accordingly. A higher weighted whitelist entry has a longer time to live (TTL). Each time a whitelist entry is matched, its TTL is refreshed, extending its validity period.

[0031] If feature x is not matched within the preset time, the weight of feature x is reduced and the lifetime of feature x is shortened accordingly. If a whitelist entry is not matched again before the TTL expires, it is deleted to avoid the accumulation of long-term invalid entries in the whitelist and the resulting redundancy of the whitelist database.

[0032] In order to further improve the efficiency of alarm information analysis and noise reduction in the embodiment of the present invention, after removing the successfully matched original alarm information from the original alarm information set, the following steps are further included: The remaining original alarm information is standardized, and multiple pieces of standardized alarm information within a preset time window are deduplicated, and the remaining alarm information after deduplication is used as target alarm information.

[0033] Specifically, each received raw alarm message is first converted into a standardized alarm format. This conversion process includes assigning a standardized name to each raw alarm message and copying the alarm's attributes into the corresponding fields of the new standardized alarm according to the attribute mapping relationships predefined in the standardized database. The standardized alarm message will contain a series of unified attributes. The detailed list and structure of these attributes are shown in the following table: In addition, a fixed-size time window is maintained to continuously monitor and organize newly arriving alarm messages. Incoming alarm messages are sorted by timestamp to ensure that alarms are processed in chronological order. When processing each alarm message, a dictionary is used to store each unique alarm identifier (a five-tuple consisting of source IP address, destination IP address, attack type, source port, and destination port) and the time of its last appearance. The sorted alarm messages are then traversed, and for each alarm message, a check is performed to see if it has appeared repeatedly within a short period of time (e.g., within 3 seconds). If so, the alarm message is marked as a duplicate and removed from the processing queue. If not, it is considered a new or ongoing alarm message and retained.

[0034] Furthermore, informational alerts that do not involve actual security threats are screened out to improve the efficiency and accuracy of alert processing. Some specific behaviors, while marked as alerts, are actually normal business activities. For example, alerts such as "Remote access to the printer service discovered" and "Remote connection tool ToDesk activity discovered" may appear abnormal from a security system perspective, but are reasonable and expected behavior in daily business operations. Based on this information, a filtering mechanism has been implemented in the alert processing process to specifically exclude these business-related alerts that have been confirmed to be non-security threats.

[0035] Furthermore, alerts that truly represent security threats are identified and filtered out, allowing us to focus on incidents that truly require a response. This process is accomplished by analyzing alert characteristics and contextual information, ensuring that corresponding logs are only recorded when an attack is confirmed to be successful. For example, some alerts may indicate the occurrence of an attack attempt, but this does not necessarily mean that the attack was successful. By combining historical data, attack paths, and feedback from defense mechanisms, the screening mechanism can effectively distinguish potential false alarms from real security incidents, thereby improving the accuracy of alert processing. In this embodiment, a combination of regular expressions and alert fields is used to determine whether an attack was successful.

[0036] Step S120: constructing an alarm graph based on the target alarm information; the nodes of the alarm graph are features in the target alarm information, and the edges between the nodes of the alarm graph are the relationships between the nodes; In this embodiment, the nodes of the alarm map include: Alarm event node: Each node represents a unique alarm record.

[0037] Network asset node: Extracts the source and destination IP addresses from the alarm information, creates a corresponding node for each unique IP address (which may be further associated with asset information such as the host name), and clearly identifies the host entity in the network.

[0038] Network service node: Extracts the destination port number from the alarm information and creates a node for each port number (which can be associated with its standard service name, such as 80 / HTTP, 445 / SMB).

[0039] Payload feature node: Analyze the alarm payload (such as HTTP request body, response body, and data packet content) in the alarm information, extract key features, and create nodes. Specific methods may include: Calculate the hash value (such as SHA-256) of the key parts of the payload (for example, the URL path after removing variable parameters, the POST request body, and the malicious script content). Create a "payload summary node" for each unique hash value to quickly associate alerts with the same malicious payload.

[0040] Utilize the preset regular expression library to match and extract specific patterns in the alert payload, such as known malware family identifiers, common web attack vectors (such as SQL injection, XSS feature fragments), specific file names, sensitive system commands or API calls, etc., and create a "payload pattern node" for each successfully matched pattern.

[0041] Extracts and normalizes the URL path from an HTTP request, creating a "URL path node" for each unique path.

[0042] User-Agent node: Extracts and normalizes the User-Agent string in the HTTP request header and creates nodes for different user agents (which may represent browsers, scripts, scanning tools, etc.).

[0043] Based on the alarm information and predefined logical rules, various types of edges are established between the created nodes to represent the specific connections between them: 1. Basic association edges Trigger relationship: connects the source host asset node and the alarm event node initiated by it.

[0044] Target relationship: connects the alarm event node with the target host asset node and / or target service port node it accesses.

[0045] Containment relationship: connects the alarm event node with its parsed payload feature nodes (summary, pattern, URL, etc.) and user agent nodes.

[0046] 2. Contextual Edges Temporal proximity: Consider a pair of alert event nodes (Alert_i, Alert_j). If they occur within a set time window Δt (e.g., 5 minutes) and share at least one key associated node in the graph (e.g., the same source IP address, the same destination IP address, or the same payload summary node), a "temporal proximity" edge is established between the two alert event nodes.

[0047] Behavioral Similarity: For network asset nodes (especially host IP addresses), we periodically analyze their activity profiles over the past period (e.g., 24 hours). This profile can include the set of destination IP addresses accessed, the distribution of destination ports used, the sequence and frequency of triggered alarm types, and the characteristics of generated network traffic. We then use the Jaccard similarity coefficient to calculate the similarity scores between different host profiles. If the similarity score between any two hosts exceeds a preset threshold, θ_sim, a "Behavioral Similarity" relationship edge is established between the corresponding network asset nodes.

[0048] Step S130: extract all nodes and the edges between them within N hops around the node of the alarm graph to form a local context subgraph of each node; In this embodiment, N is 2 or 3.

[0049] Step S140: Based on a predefined path pattern (e.g., "alarm -> contains -> malicious payload summary -> contained in another alarm" or "alarm -> triggered from -> a host -> host behaves similarly -> another host -> triggers -> another alarm"), biased or unbiased sampling is performed in the alarm graph to obtain a series of node path sequences representing specific contextual logic; Step S150: Generate a feature vector of each node from the local context subgraph and each node path sequence; To explain this step in detail, the local context subgraph of each node and the path sequence of each node are input into the GraphSAGE (Graph Sample and Aggregate) model to generate a fixed-length (for example, set to 256 dimensions) context feature vector for each alarm event node. This vector condenses the core information of the alarm in its context graph environment.

[0050] Step S160: quantify the local context density ρi of the feature vector vi by counting the number of other feature vectors falling into the neighborhood with a radius of εi centered on the feature vector vi; specifically, calculate the distance L from the feature vector vi to its Kth nearest neighbor feature vector, and use the distance L or a multiple of the distance L as the radius εi.

[0051] Step S170: Calculate the relative distance δi between the feature vector vi and the nearest feature vector among all feature vectors whose local context density is greater than the local context density ρi; for the vector with the highest global density (there is no vector with a higher density than it), its δi is defined as the maximum distance from the point to all other vectors in the data set, which is treated as a boundary condition.

[0052] Step S180: If the local context density ρi and the relative distance δi are both greater than or equal to the preset threshold, the feature vector vi corresponding to the local context density ρi and the relative distance δi is used as the cluster center, and an independent alarm cluster is maintained; Step S190: For the feature vector vj that is not a cluster center, assign it to the alarm cluster maintained by the cluster center closest to the feature vector vj among all feature vectors whose local context density is greater than the local context density of the feature vector vj; Step S200: Input each alarm cluster into the constructed information entropy evaluation model, custom rule evaluation model and large language semantic evaluation model respectively, and perform weighted fusion on the output scores of the constructed information entropy evaluation model, custom rule evaluation model and large language semantic evaluation model to obtain the final alarm evaluation result.

[0053] In this embodiment, the information entropy evaluation model utilizes the concept of entropy from information theory to quantify the inherent complexity of clusters and the uncertainty of behavioral patterns, as well as assess the rarity and abnormality of outlier alerts. This model penetrates the aggregated alert set and, by evaluating its characteristic distribution, temporal dynamics, and evolutionary patterns, identifies alerts with complex structures, sudden behavioral changes, or rare features. These are often associated with more advanced, covert, or destructive attacks, thus providing a quantitative basis for risk assessment.

[0054] Specifically, the information entropy of clusters is evaluated from the following three dimensions: 1. Intra-cluster feature entropy: measuring the diversity of attack activities A cluster represents a group of internally related alert events. Intra-cluster feature entropy measures the degree of dispersion in the distribution of key attributes of these alerts, reflecting whether the attack activity is repetitive or complex and varied. A set of key features is selected, such as: source IP address set S = {s1, s2, ...}, destination IP address set D = {d1, d2, ...}, destination port set P = {p1, p2, ...}, attack type set A = {a1, a2, ...}, and payload feature node type set L = {l1, l2, ...}.

[0055] 2. Intra-cluster temporal distribution entropy: quantifying temporal clustering The temporal pattern of events within a cluster contains information about the rhythm and urgency of the attack. The temporal distribution entropy within a cluster is used to quantify the degree of concentration or dispersion of these alerts on the time axis.

[0056] 3. Cluster Evolution Entropy: Insight into Pattern Changes (Relying on Historical Data) By tracking the frequency and pattern changes of similar clusters over time, the dynamic evolution of the attack landscape can be assessed. This dimension requires maintaining data on historical alert clusters. Characteristic representations of alert clusters can be defined (such as cluster centroid vectors and key node combinations), and historically similar or similar clusters can be identified based on similarity (such as vector distance and shared node ratio).

[0057] Entropy analysis of outlier alarms evaluates the rarity of the key features that constitute the outlier alarm in its global context. An outlier alarm o may be identified by factors such as its unique payload feature node l_rare, rare source IP s_unusual, or access to atypical services d_port_special.

[0058] The custom rule-based evaluation model leverages clusters and outliers, along with the rich contextual information contained in these outputs (such as aggregate statistics, graph relationships, and asset associations), to conduct refined risk assessments. This model allows security teams to continuously optimize and adjust scoring strategies based on the evolving threat landscape and internal environment, ensuring the accuracy and relevance of their findings.

[0059] The Large Language Semantic Evaluation Model leverages advanced natural language processing (NLP) and AI reasoning capabilities to provide a deep understanding of the semantics and threat intent of alert clusters and outliers. This model utilizes Large Language Models (LLMs) such as Qwen2.5-7b-Instruct to simulate the analysis process of human security experts, uncovering the underlying attack logic, tactical phases, and overall risk landscape behind a collection of alert events.

[0060] Specific instructions for this step: Each alarm cluster is input into the constructed information entropy evaluation model, including: Alarm clusters whose number of alarm information is greater than or equal to the minimum scale threshold, or whose average local context density of alarm information within the cluster is greater than or equal to the comparison threshold, are regarded as clusters; each cluster contains a group of alarm information with similar contexts that are aggregated according to the adaptive density peak clustering criterion.

[0061] Alarm clusters in which the number of alarm information is less than the minimum scale threshold or the average local context density of the alarm information within the cluster is less than the comparison threshold are regarded as outlier clusters; outlier clusters contain alarm information identified as isolated points or belonging to low-density, small-scale groups.

[0062] Clusters are input into the information entropy evaluation model. The Shannon entropy H(X|C) of the alert feature X in cluster C is calculated using the formula H(X|C) = -Σᵢ [P(xi) * log2(P(xi))]. P(xi) is the probability that a value xi (e.g., a specific IP address) of the alert feature X (e.g., source IP address) occurs within the cluster. Based on the Shannon entropy H(X|C) and pre-defined rules, a cluster feature entropy score is output. A high H(X|C) value indicates that the feature is widely distributed and diverse within the cluster. For example, a cluster with a high source IP entropy H(S|C) strongly suggests that it involves a large number of different attack sources, possibly a distributed attack (e.g., DDoS) or large-scale scanning. Similarly, a high destination port entropy H(P|C) is often associated with port scanning, while a high load node type entropy H(L|C) may indicate that the attacker has attempted multiple vulnerability exploits or executed a complex attack chain involving reconnaissance, exploitation, and persistence. Conversely, low entropy across all features indicates a single attack pattern and a focused target. Alternatively, the temporal entropy H(T|C) of cluster C can be calculated using the formula H(T|C) = -Σⱼ [P(tj) * log2(P(tj))] , where P(tj) is the probability of an alert falling into a predefined time bucket. Based on the temporal entropy H(T|C) and predefined rules, a temporal entropy score within the cluster is generated. A low entropy value (close to 0) indicates that the majority of alerts are concentrated within a very small number of time buckets, demonstrating strong temporal clustering or bursts. This typically indicates high-intensity attack waves (such as short, intensive brute force attempts) or a critical phase of an attack, indicating a high degree of urgency. Conversely, a higher entropy value indicates a more even temporal distribution of alerts or a complex, multi-peaked pattern, potentially indicating persistent, slow-moving infiltration activity or automated attacks with specific temporal patterns. Alternatively, the cluster evolution entropy H(Patterns|ΔT) of cluster C is output using the formula H(Patterns|ΔT) = - Σ[P(pattern_k|ΔT) * log2(P(pattern_k|ΔT))] , where P(pattern_k) is the frequency of occurrence of clusters with the preset pattern k, and ΔT is the set time window. A cluster evolution entropy score is output based on H(Patterns|ΔT) and preset rules. Specifically, the score is calculated within a time window ΔT. If H(Patterns|ΔT) increases significantly within a certain time window, or a new, previously unseen pattern, pattern_new, appears (where P(pattern_new) changes from 0 to a positive value), this may indicate new attack activity, policy adjustments, or the emergence of a new threat, and thus has high information value and potential risks.In this embodiment, P(xi) is calculated using the formula P(xi) = count(xi) / N; where count(xi) is the frequency of the value xi of feature X appearing in cluster C, and N is the number of alarm information in cluster C; P(tj) is calculated using the formula P(tj) = nj / N; where nj is the number of alarm information falling into the j-th time bucket Δtj.

[0063] The outlier cluster is input into the information entropy evaluation model. The rarity I(yi) of the eigenvalue yi of the alarm information in the outlier cluster is output using the formula I(yi) = -log2(P_global(yi)). P_global(yi) is the global probability of occurrence of the eigenvalue yi of the alarm information in the outlier cluster within the entire alarm dataset or a related global historical alarm data set. Based on the rarity I(yi) and pre-set rules, the alarm entropy score of the outlier cluster is output. The smaller P_global(xi) (the rarer the eigenvalue), the greater its information content I(xi). The overall entropy score of the outlier alarm can be obtained by combining the information content of its key rare features (for example, taking the maximum value or weighted sum). An outlier alarm composed of multiple very rare features (i.e., high I(xi) values) will have a significantly higher entropy score. Such high-scoring outlier alarms are more likely to represent novel attack techniques, highly targeted attack (APT) attempts, or even early signs of zero-day vulnerability exploitation. Because they deviate from known, regular attack patterns (which have been absorbed by the cluster), they require priority attention.

[0064] Input each alarm cluster into the constructed custom rule evaluation model, including: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters whose number of alarm information is less than the minimum size threshold or whose average local context density of alarm information in the cluster is less than the comparison threshold are regarded as outlier clusters; The clusters and outlier clusters are respectively input into the custom rule evaluation model, and the custom rule alarm scores of the clusters and outlier clusters are respectively output according to their respective custom rules.

[0065] Specifically, the embodiment of the present invention defines two scoring rules for clustering, as follows: 1. Scoring rules based on aggregated indicators This rule directly uses the statistics within the cluster to quantify the score. They focus on factors such as the overall "size", "severity", "confidence" and "scope of impact" of the cluster, including: Severity and Confirmation Scoring Rules: If a cluster contains at least one alarm with a "high" severity level and the average confidence level of the alarms in the cluster exceeds the set value, its risk score will be significantly increased, indicating that there is a serious threat with a high degree of confirmation in the cluster.

[0066] Scoring rules for signs of successful attacks: If there are more than three alerts in a cluster that clearly indicate "successful attack" or are labeled "post-exploitation phase," the score is increased. This strongly suggests that the attacker has made progress.

[0067] Scoring rules for scope and direction: By analyzing the number and distribution of source / destination IP addresses and the direction of alerts (e.g., crossing network boundaries), we assess whether the attack is distributed, scanning, or targeted, and adjust the score accordingly, particularly giving high scores to attacks that cross critical security domains.

[0068] 2. Scoring rules based on context These rules leverage the graph relationships and contextual information built during the alert aggregation phase to provide a deeper level of risk assessment. They focus on how clusters relate to other entities (assets, other alert clusters, known threat patterns), specifically: Critical Asset Association Scoring Rule: If the activity in a cluster directly targets an asset marked as "critical" and belonging to the "Finance" business department, and the association indicates an "exploitation attempt," then the risk score is increased. This requires querying the node attributes associated with the cluster in the aggregate graph.

[0069] Attack chain pattern recognition scoring rules: By analyzing the alarm nodes in the cluster and their temporal / logical relationships (such as the edges in the graph), if a pattern that matches a specific attack chain (such as the MITRE ATT&CK Tactic sequence) is identified, a high score is given.

[0070] Behavior similarity association scoring rule: If the current cluster is connected to an alarm cluster that has been confirmed or marked as a known malicious activity (such as a specific APT organization) through a "behavior similarity" relationship, it inherits or is assigned a high risk score.

[0071] The embodiment of the present invention defines three scoring rules for outlier clusters, as follows: 1. Scoring rules based on the attributes of the alert itself and the associated assets: For an outlier alert with a high confidence level, a critical level, and a target of a high-criticality asset, a high score is assigned.

[0072] 2. Scoring rules based on rarity and novelty: If the payload features (such as hashes, behavior pattern nodes) carried by an outlier alert have never been seen in recent history, or if the combination of its source (such as an uncommon IP address from a sanctioned country) and the accessed service is unusual, its score will be increased to reflect its potential new threat value.

[0073] 3. Scoring rules based on external intelligence matching: If the key indicators (IP, domain name, file hash, etc.) in the outlier alert match with high-confidence external threat intelligence sources, its score is increased.

[0074] Each alarm cluster is input into the constructed large language semantic evaluation model, including: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters whose number of alarm information is less than the minimum size threshold or whose average local context density of alarm information in the cluster is less than the comparison threshold are regarded as outlier clusters; The clusters and outlier clusters are respectively input into the large language semantic evaluation model, and the large model semantic alarm evaluation scores of the clusters and outlier clusters are respectively output according to their respective prompt words.

[0075] Specifically, in order to fully utilize the understanding capabilities of the large language semantic evaluation model, in this embodiment, the input to the model is no longer isolated alarm information, but a comprehensive portrait of carefully organized clusters, including: Structured summary information: Key metadata, such as the main entities involved (Top N source / destination IP addresses, ports, and associated asset tags), alarm type distribution, total number of alarms in the cluster, time span (start, end, and duration), and preliminary aggregate scores given by other models (such as information entropy and rule models). This provides a macro overview for the large language semantic evaluation model.

[0076] Representative alert details: Select several representative alerts within the cluster (for example, the alerts with the highest severity, the earliest / latest alerts, or alerts near the cluster center vector) and provide their detailed information, especially fields containing semantic information, such as HTTP request header / body fragments, detected malware family names, rule trigger descriptions, and possible original payload samples (properly processed to control input length and security). This provides concrete "evidence" for the large language semantic evaluation model.

[0077] Context vector or explanation: If the alarm aggregation stage generates a feature vector representing the cluster context, this vector (or its interpretable transformation, such as a description of key feature contributions) is also input to provide more abstract pattern information for the large language semantic evaluation model.

[0078] Without fine-tuning the model, a pre-trained large-language semantic evaluation model can be guided to perform zero-shot or few-shot reasoning by designing precise prompts. The prompts must clearly define the task and provide the necessary context and output format requirements. For example, the prompt: "Please analyze the comprehensive information of the following cluster. This cluster [provides structured summary information]. It contains representative alerts: [List 1-3 alert details]. Based on cybersecurity knowledge, particularly the MITRE ATT&CK framework, determine the most likely attack intent (e.g., initial access, execution, persistence, privilege escalation, defense evasion, credential access, discovery, lateral movement, collection, command and control, impact) and the current attack phase for this cluster. Assess its overall threat level (low / medium / high / critical) and briefly explain the reasoning." This leverages the generalization capabilities of the large-language semantic evaluation model to connect discrete alert data points into a logically connected attack narrative. The model is required to go beyond simply identifying keywords and instead perform context-based reasoning.

[0079] To achieve more accurate, efficient, and environmentally-specific analysis capabilities, embodiments of the present invention also employ supervised fine-tuning. This requires constructing a high-quality annotated dataset containing a large number of (cluster representation, expert analysis results) data pairs. The cluster representation is constructed in the same way as the input described above, and the expert analysis results include standardized labels, such as the precise attack type (e.g., "SQL injection attempt," "Log4j vulnerability exploit," "Cobalt Strike C&C communication"), attack phase (Tactic ID), risk score (numerical or graded), and recommended action category. A fine-tuned model (such as a fine-tuned version of Qwen2.5-7b-Instruct) can more directly and accurately output structured analysis conclusions for specific alert clusters, reducing reliance on complex prompt words and potentially capturing subtle patterns that would be difficult to detect using a general model alone. The output can be directly used to drive subsequent automated responses or work order systems.

[0080] For outlier alerts within an outlier cluster, the large language semantic evaluation model assesses their uniqueness and potential high-risk indicators, particularly their potential as emerging threats or highly targeted attacks (such as early signs of APTs and zero-day exploits). The input primarily consists of detailed information about the outlier alert, but the prompt word specifically emphasizes its "outlier" status. For example: Prompt: "Analyze the following independent alert: [Provide alert details, including time, source, destination, protocol, port, payload summary, rule information, etc.] This alert cannot be associated with any known activity pattern (alert cluster) in our system. Please assess the possibility of it being one of the following: (a) a new or unknown attack technique; (b) a tentative behavior for a targeted attack; (c) a zero-day vulnerability exploit; (d) a false positive or a benign anomaly. Please provide your risk judgment (low / medium / high / critical) and reasoning basis, paying special attention to its characteristics that are different from the normal pattern.", thereby utilizing the extensive knowledge base and reasoning capabilities of the large language semantic evaluation model to deeply interpret a single seemingly isolated but abnormally behaving alert to determine whether it indicates a special threat that requires immediate attention.

[0081] The output scores of the constructed information entropy evaluation model, custom rule evaluation model, and large language semantic evaluation model are weighted and integrated to obtain the final alarm evaluation results, including: Normalize the output scores of the constructed information entropy evaluation model, custom rule evaluation model, and large language semantic evaluation model to obtain the normalized alarm scores; Specifically, the output scores of the information entropy evaluation model, the custom rule evaluation model, and the large language semantic evaluation model are normalized to a unified numerical range (e.g., 0 to 100). For example, the "critical," "high," "medium," and "low" levels output by the large model can be mapped to preset numerical values ​​(e.g., 100, 80, 50, 20).

[0082] The normalized alarm scores are weighted and summed to obtain the final alarm evaluation result.

[0083] Specifically, a weight (W_entropy, W_rules, W_llm) is assigned to the normalized score of each evaluation model. These weights reflect the organization's assessment of the reliability, accuracy, and importance of each model in a specific scenario. The sum of the weights is 1. The formula for calculating the comprehensive risk score is as follows: FinalScore = W_entropy * NormalizedScore_entropy + W_rules * NormalizedScore_rules + W_llm * NormalizedScore_llm. The weights can be adjusted based on security policies, current threat focus, or historical model performance. For example, when dealing with new or unknown threats, the weight of the semantic evaluation of the large model may be increased; when focusing on specific compliance violations, the weight of the custom rule model may be higher. Mapping the comprehensive risk score to a clear risk level enables effective screening and prioritization of alert events, and provides clear action guidelines for subsequent security response activities.

[0084] It should be noted that in order to handle extreme cases or high-confidence judgments, special rules can be set to override the weighted scoring. For example: If the custom rule model directly determines it as a "confirmed true positive" based on manual annotation or specific high-priority rules, the FinalScore can be directly set to the highest score (such as 100).

[0085] If the large language semantic evaluation model outputs a "critical" risk level with high confidence and is supported by the rule model (such as involving key assets or signs of successful attack), the FinalScore can also be forcibly upgraded to a high segment.

[0086] On the contrary, the FinalScore of an alarm cluster / outlier alarm that is clearly marked as a "false positive" or "benign scan" can be directly set to zero.

[0087] After obtaining the comprehensive risk score, thresholds need to be set to divide it into different risk levels (for example: critical, high, medium, and low).

[0088] In summary, the embodiment of the present invention can first filter out alarms that do not pose a security threat through customized rules and result feedback mechanisms, thereby reducing the number of invalid alarms and optimizing the quality of the alarm stream input into the large model; then, through methods such as standardization, time deduplication, and special type screening, it can reduce the total number of alarms and aggregate similar alarms, thereby improving processing efficiency; finally, it uses information entropy analysis, rule evaluation, and large model semantic analysis to deeply analyze alarm information, identify high-threat alarms, and assist in decision-making. The embodiment of the present invention can be widely used in network security operation and maintenance scenarios, and is particularly suitable for processing high-frequency, high-volume alarm data, significantly improving the efficiency and accuracy of alarm analysis and judgment.

[0089] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0090] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0091] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0092] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0093] Any details not described in the embodiments of the present invention are well-known to those skilled in the art. Finally, it should be noted that the above embodiments are only intended to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or equivalents should be included in the scope of the claims of the present invention.

Claims

1. A method for analyzing and reducing noise of alarm information, characterized in that: include: Obtain target warning information; The characteristics of the target alarm information include at least: the alarm event, the source IP address, destination IP address and destination port number in the alarm event, the alarm payload, and the alarm time; Constructing an alarm graph based on the target alarm information; the nodes of the alarm graph are features in the target alarm information, and the edges of the nodes of the alarm graph are the relationships between the nodes; Extract all nodes and the edges between them within N hops around the node of the alarm graph to form a local context subgraph of each node; According to a predefined path pattern, biased or unbiased sampling is performed in the alarm graph to obtain a series of node path sequences representing specific context logic; Generate a feature vector of each node from the local context subgraph and each node path sequence; The local context density ρi of the feature vector vi is quantified by counting the number of other feature vectors falling into the neighborhood with a radius of εi centered on the feature vector vi; Calculate the relative distance δi between the feature vector vi and the feature vector with the closest local context density ρi; If the local context density ρi and the relative distance δi are both greater than or equal to the preset threshold, the feature vector vi corresponding to the local context density ρi and the relative distance δi is used as the cluster center, and an independent alarm cluster is maintained; For the feature vector vj that is not the cluster center, assign it to the alarm cluster maintained by the cluster center closest to the feature vector vj among all feature vectors whose local context density is greater than the local context density of the feature vector vj; Each alarm cluster is input into the constructed information entropy evaluation model, custom rule evaluation model and large language semantic evaluation model respectively, and the output scores of the constructed information entropy evaluation model, custom rule evaluation model and large language semantic evaluation model are weighted and fused to obtain the final alarm evaluation result.

2. The method for analyzing and reducing noise of alarm information according to claim 1, characterized in that: Also includes: Calculate the distance L between the feature vector vi and its Kth nearest neighbor feature vector, and use the distance L or a multiple of the distance L as the radius εi.

3. The method for analyzing and reducing noise of warning information according to claim 1, characterized in that: The obtaining of target warning information includes: Get the original alarm information set; Matching the original alarm information in the original alarm information set with a whitelist template in a preset whitelist library; If the match is successful, the original alarm information that is successfully matched is removed from the original alarm information set, and the remaining original alarm information is retained as the target alarm information.

4. The method for analyzing and reducing noise of warning information according to claim 3, characterized in that: After removing the successfully matched original alarm information from the original alarm information set, the method further includes: The remaining original alarm information is standardized, and multiple pieces of standardized alarm information within a preset time window are deduplicated, and the remaining alarm information after deduplication is used as the target alarm information.

5. The method for analyzing and reducing noise of warning information according to claim 3, characterized in that: Also includes: If it is determined that the original alarm information X in the original alarm information set is a false alarm, extract the feature x of the original alarm information X, add the feature x to the preset whitelist library, and assign an initial weight and a survival time; If the feature x is matched once, the weight of the feature x increases, and the survival time of the feature x is extended accordingly; If the feature x is not matched within the preset time, the weight of the feature x is reduced and the survival time of the feature x is shortened accordingly.

6. The method for analyzing and reducing noise of warning information according to claim 1, characterized in that: The step of inputting each alarm cluster into the constructed information entropy evaluation model includes: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters in which the number of alarm information in each alarm cluster is less than the minimum size threshold, or the average local context density of the alarm information in the cluster is less than the comparison threshold, are regarded as outlier clusters; The cluster is input into the information entropy evaluation model, and the Shannon entropy H(X|C) of the feature X of the alarm information in the cluster is obtained by outputting the formula H(X|C) = -Σᵢ [P(xi) * log2(P(xi))]; wherein P(xi) is the probability that the value xi of the feature X of the alarm information appears in the cluster; the intra-cluster feature entropy score is output according to the Shannon entropy H(X|C) and the preset rules; and / or, the time distribution entropy H(T|C) of the cluster C is obtained by outputting the formula H(T|C) = -Σⱼ [P(tj) * log2(P(tj))]; wherein P(tj) is the probability that the alarm information falls into the preset time bucket; the intra-cluster time distribution entropy score is output according to the time distribution entropy H(T|C) and the preset rules; and / or, the time distribution entropy score of the cluster is output according to the formula H(Patterns|ΔT) = -Σ[P(pattern_k|ΔT) * log2(P(pattern_k|ΔT))] outputs the cluster evolution entropy H(Patterns|ΔT) of cluster C; where P(pattern_k) is the frequency of occurrence of clusters with the preset pattern k, and ΔT is the set time window; based on the cluster evolution entropy H(Patterns|ΔT) and preset rules, the cluster evolution entropy score is output; The outlier cluster is input into the information entropy evaluation model, and the rarity I(yi) of the characteristic value yi of the alarm information in the outlier cluster is output through the formula I(yi) = -log2(P_global(yi)); wherein P_global(yi) is the global occurrence probability of the characteristic value yi of the alarm information in the outlier cluster in the entire alarm data set or a related global historical alarm data range; the alarm entropy score of the outlier cluster is output according to the rarity I(yi) and preset rules.

7. The method for analyzing and reducing noise of warning information according to claim 6, characterized in that: Also includes: The P(xi) is calculated by the formula P(xi) = count(xi) / N; wherein count(xi) is the frequency of the value xi of the feature X in the cluster C, and N is the number of alarm information in the cluster C; The P(tj) is calculated by the formula P(tj) = nj / N, where nj is the number of alarm information falling into the j-th time bucket Δtj.

8. The method for analyzing and reducing noise of alarm information according to claim 1, wherein: The step of inputting each alarm cluster into the constructed custom rule evaluation model includes: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters in which the number of alarm information in each alarm cluster is less than the minimum size threshold, or the average local context density of the alarm information in the cluster is less than the comparison threshold, are regarded as outlier clusters; The clusters and the outlier clusters are respectively input into the custom rule evaluation model, and the custom rule alarm scores of the clusters and the outlier clusters are respectively output according to their respective custom rules.

9. The method for analyzing and reducing noise of warning information according to claim 1, wherein: The step of inputting each alarm cluster into the constructed large language semantic evaluation model includes: The alarm clusters whose number of alarm information in each alarm cluster is greater than or equal to the minimum size threshold, or whose average local context density of alarm information in the cluster is greater than or equal to the comparison threshold, are regarded as clusters; The alarm clusters in which the number of alarm information in each alarm cluster is less than the minimum size threshold, or the average local context density of the alarm information in the cluster is less than the comparison threshold, are regarded as outlier clusters; The clusters and the outlier clusters are respectively input into the large language semantic evaluation model, and large model semantic alarm evaluation scores of the clusters and the outlier clusters are respectively output according to their respective prompt words.

10. The method for analyzing and reducing noise of alarm information according to any one of claims 1 to 9, characterized in that: The output scores of the constructed information entropy evaluation model, the custom rule evaluation model and the large language semantic evaluation model are weighted and integrated to obtain the final alarm evaluation result, including: Normalizing the output scores of the constructed information entropy evaluation model, the custom rule evaluation model, and the large language semantic evaluation model to obtain each normalized alarm score; A weighted sum is performed on the normalized alarm scores to obtain the final alarm evaluation result.

Citation Information

Cited By

  • Task alarm processing method and system based on intelligent grading

    CN121455650A

  • A task alarm handling method and system based on intelligent hierarchical classification

    CN121455650B

  • Alarm noise reduction method and system based on large model of power system

    CN121998403A

  • Intelligent research and judgment method and system for false alarm of multi-source static analysis alarm

    CN122364046A