Network security alarm data research and judgment method, electronic equipment, storage medium and program product

By building a network security alarm data analysis method that combines an expert experience rule base and a large model, the problem of high false alarm rate is solved, efficient and accurate alarm data analysis is achieved, and it adapts to the dynamic changes in the network security environment.

CN120639447AActive Publication Date: 2025-09-12BEIJING TOPSEC NETWORK SECURITY TECH +2

Patent Information

Application Number
CN202510960424.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-12
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

The high false alarm rate of network security alarm data in existing technologies leads to a heavy and difficult workload for security operations personnel in research and analysis, and it is difficult to identify real attacks from the massive amount of false alarms.

Method used

Build a knowledge base based on expert experience rules, combine big models and RAG technology, obtain real-time alarm data and match it with the knowledge base, and use the analysis model to perform automated and intelligent alarm data analysis, including a rule base of structured and unstructured knowledge, and vectorize processing to improve matching efficiency.

Benefits of technology

It improves the accuracy and efficiency of analyzing network security alarm data, can quickly identify real attacks and reduce false alarm rates, and adapt to changes in the network security environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639447A_ABST
    Figure CN120639447A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a network security alarm data research and judgment method, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining real-time alarm data and a pre-stored knowledge base which comprises expert experience rules, the expert empirical rule is used for judging whether the alarm data is data generated by a real attack or data generated by an unreal attack; determining a target expert experience rule corresponding to the real-time alarm data from the knowledge base; and inputting the real-time alarm data and the target expert experience rule into a research and judgment model to obtain a research and judgment result for the real-time alarm data, the research and judgment result being used for indicating that the real-time alarm data is data generated due to a real attack or data generated due to an unreal attack. Expert experience rules are used for providing guidance for research and judgment, automatic and intelligent analysis is achieved in combination with the research and judgment model, and the efficiency and accuracy of research and judgment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of analysis and judgment technology, and specifically to a network security alarm data analysis and judgment method, electronic equipment, storage medium and program product. Background Art

[0002] With the gradual development of network security, more and more equipment and tools are being deployed in security detection and protection efforts within enterprise security operations. While firewalls, traffic probes, and security analysis tools provide network security protection, they also generate a large number of false positives. Due to the time differences between the attacker and the defender, as well as the timeliness of detection rules, false positives are inevitable. These false positives complicate security operations, often drowning out real attacks amidst a flood of false positives, creating a significant workload and challenges for security operations personnel. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide a network security alarm data analysis method, electronic device, storage medium and program product to achieve the technical effect of improving the accuracy and efficiency of analysis.

[0004] A first aspect of an embodiment of the present application provides a method for analyzing and judging network security alarm data, the method comprising: Acquiring real-time alarm data and a pre-stored knowledge base, wherein the knowledge base includes expert experience rules, and the expert experience rules are used to determine whether the alarm data is generated by a real attack or by a false attack; Determining target expert experience rules corresponding to the real-time alarm data from the knowledge base; The real-time alarm data and the target expert experience rules are input into the analysis and judgment model to obtain an analysis and judgment result for the real-time alarm data, and the analysis and judgment result is used to indicate whether the real-time alarm data is data generated by a real attack or data generated by a non-real attack.

[0005] In this implementation, real-time alert data is collected from a knowledge base containing expert experience rules, from which corresponding target rules are determined. The analysis model then uses this to determine whether the alert data represents a genuine attack. This complete analysis process, guided by expert experience rules and combined with the analysis model for automated and intelligent analysis, enables rapid and accurate determination of the nature of network security alert data, improving both efficiency and accuracy.

[0006] Furthermore, the expert experience rules include structured language rules and / or unstructured knowledge, and the structured language rules include non-real attack rules and real attack rules defined from the network security domain dimension, attack feature dimension, and asset attribute dimension; the unstructured knowledge includes security analysis reports, threat intelligence documents, and historical research and judgment records.

[0007] In this implementation, expert experience rules comprise both structured language rules and unstructured knowledge. The structured language rules define the rules for both fake and real attacks from multiple perspectives, while the unstructured knowledge encompasses a wide range of security-related documents. This comprehensive rule-building approach enriches the knowledge base, enabling expert experience rules to describe network security characteristics from diverse perspectives and levels. This enhances the rules' coverage of various network security scenarios, thereby improving the comprehensiveness and accuracy of network security alert data analysis.

[0008] Furthermore, the expert experience rules are stored in the knowledge base in the form of vectors; and before obtaining the real-time alarm data and the pre-stored knowledge base, the method further includes: The initial expert experience rules are vectorized to obtain the knowledge base consisting of the vectorized expert experience rules.

[0009] In the above implementation process, vectorizing the rules facilitates the computer to perform efficient similarity calculation and data understanding analysis, and can quickly and accurately match the target expert experience rules corresponding to the real-time alarm data from the knowledge base.

[0010] Furthermore, the expert experience rules include a plurality of rules; and determining the target expert experience rules corresponding to the real-time alarm data from the knowledge base includes: Vectorizing the real-time alarm data to obtain vectorized real-time alarm data; The similarity between the vectorized real-time alarm data and each of the expert experience rules is calculated, and the expert experience rule whose similarity meets the preset requirements is determined as the target expert experience rule.

[0011] In the above implementation process, the target rule is determined by vectorizing the real-time alarm data and calculating its similarity with each expert experience rule in the knowledge base. This allows the analysis of real-time alarm data to leverage the computer's ability to efficiently process vector data, quickly and accurately screening out matching rules from numerous rules in the knowledge base.

[0012] Furthermore, the real-time alarm data and the target expert experience rules are input into the analysis and judgment model to obtain the analysis and judgment results for the real-time alarm data, including: Inputting the real-time alarm data and the target expert experience rule into the analysis and judgment model, and if the similarity between the vectorized real-time alarm data and the target expert experience rule exceeds a first threshold, determining the analysis and judgment conclusion in the target expert experience rule as the analysis and judgment result; If the similarity between the vectorized real-time alarm data and the target expert experience rule is lower than the first threshold, the analysis result and the confidence level of the analysis result are determined based on the real-time alarm data.

[0013] In the above implementation, when the similarity exceeds the threshold, the target expert's empirical rules are directly used to determine the conclusions, leveraging the reliability of expert experience to quickly and accurately produce results. When the similarity does not exceed the threshold, the data analysis and processing capabilities of the large model are utilized to provide the judgment results and confidence levels. This ensures that judgments can be made even in complex or unknown situations, and the confidence level reflects the credibility of the results.

[0014] Furthermore, the real-time alarm data includes an attacker identifier, a victim identifier, an alarm type, and attack characteristics.

[0015] In the above implementation process, it is clear that the real-time alarm data includes the attacker identifier, the attacked identifier, the alarm type, and the attack characteristics, etc., providing comprehensive and detailed basic information for research and judgment.

[0016] Furthermore, the method further comprises: When it is determined that the knowledge base does not have the target expert experience rules corresponding to the real-time alarm data, the knowledge base is adjusted to obtain an updated knowledge base, and the step of determining the target expert experience rules corresponding to the real-time alarm data from the knowledge base is returned to be executed.

[0017] In the above implementation process, the knowledge base is decoupled from the large model. When the knowledge base does not contain corresponding target expert experience rules, the knowledge base is adjusted to overcome the limitations of the fixed knowledge base, enabling the model to adapt to the ever-changing network security environment and ensuring the model's ability to analyze and judge various types of alarm data.

[0018] According to a second aspect of an embodiment of the present application, an electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instruction, any method described in the first aspect is implemented.

[0019] A third aspect of an embodiment of the present application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of any of the methods described in the first aspect.

[0020] A fourth aspect of the embodiments of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements any method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 A schematic diagram of an overall process provided in an embodiment of the present application; Figure 2 A flowchart of a method for analyzing network security alarm data provided in an embodiment of the present application; Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0024] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0025] Among related technologies, large-scale modeling has some applications in reducing noise during alarm analysis and assessment. However, due to the limited security capabilities of general-purpose large-scale models, this has hindered rapid and effective improvement in noise reduction. Many noise reduction strategies and practices are often based on the manual experience of security personnel and cannot be rapidly applied with the capabilities of large-scale modeling technology.

[0026] To this end, this application proposes a network security alarm data analysis method. Starting from expert experience, combined with large models and RAG technology, a process for alarm noise reduction processing is constructed, which can quickly apply the expert experience of security service personnel and can be freely adjusted to improve alarm analysis and security operation efficiency.

[0027] Among them, Retrieval-augmented Generation (RAG) is one of the large-scale model technologies. The RAG model combines language modeling and information retrieval techniques. Specifically, when the model needs to generate text or answer a question, it first retrieves relevant information from a large collection of documents. It then uses this retrieved information to guide text generation, thereby improving the quality and accuracy of predictions.

[0028] Based on the large model and RAG technology, an intelligent alarm analysis agent can be designed. This agent can automatically draw conclusions such as attack success or false alarm based on the input alarm context and the accumulated experience and knowledge of security service and other business experts (for example, the experience provided by the security service → "Scan attacks from security domain A are all false alarms"). For example, the analysis process is as follows: (1) Input: Extract key information from alarms + analysis and judgment rule knowledge base; (2) Business processing: The prompt word is combined with the above input and the LLM (Large Language Model) is used to give the judgment result; (3) Output conclusion: real attack / false alarm / needs manual review.

[0029] In the specific implementation, the detailed analysis and judgment steps are as follows: 1. Collect expert experience and store knowledge base: The core concept of intelligent alert analysis and judgment based on a rule base based on expert experience is to automatically output conclusions through rule matching and context analysis. Its advantages are that the rules are based on expert experience, easy to understand and verify, can be quickly implemented without complex model training, and can be dynamically updated to quickly adapt to new threat changes. The implementation steps include: ① Rule sources include two parts: one is structured language rules, such as existing expert rules (e.g., "The scan of security domain A is a false positive"), and the other is unstructured knowledge, such as security analysis reports, threat intelligence documents, and historical research records. As an example, the structured language rules are shown in Table 1 below: Table 1

[0030] ② The rules compiled by business experts can be vectorized through text embedding models such as bge-small-zh and then stored in a vector database (such as milvus), which is the knowledge base.

[0031] 2. Alarm context extraction. Key content is extracted from raw alarm data (i.e., alarm log data) and associated with external data (such as threat intelligence and assets) as context. The extracted key content includes the attacking IP address, attacked IP address, alarm type, attack characteristics and response content, attack phase, and time. Associated external data includes attacker information from threat intelligence data and security domains and asset types from asset data, which are combined to form the alarm context. For example: "Event description: At 2025-02-24 21:49:25, the system detected a [brute force attack] from the public IP 9.7.20.28. This behavior may be an attacker trying to crack the system login password. Source IP: 9.7.20.28, A security domain. Destination IP: 192.168.23.5, from the DT security domain in Haidian District, Beijing. Kill chain stage: reconnaissance and tracking. Occurrence time: from 2025-02-24 21:49:25 to 2025-02-24 22:02:31".

[0032] 3. Knowledge retrieval and large model application: See also Figure 1 , Figure 1 This is a schematic diagram of an overall process provided by an embodiment of the present application. Based on a large model, the alarm context is used as input for knowledge retrieval, and the best matching expert experience is used as the output of the research and judgment conclusion.

[0033] Among them, the prompt can be designed according to Table 2: Table 2

[0034] Through the above processing steps, after analysis and processing by the large model, a mark will be output to indicate whether the alarm is a confirmed attack or a false alarm, etc., and the corresponding expert experience rule information will be given.

[0035] In summary, by building a knowledge base based on expert experience (by storing expert experience in the knowledge base through built-in or added experience), performing analysis based on large model technology, extracting results based on prompts, and giving clear research and judgment conclusions, research and judgment can be automated, effectively improving the efficiency of research and judgment. In addition, expert experience can be quickly accumulated and applied, and it is easy to adjust or configure according to the actual project situation or on-site network environment, which can quickly produce research and judgment results.

[0036] In view of this, the embodiment of the present application provides a network security alarm data analysis method, which applies the expert analysis experience of security service personnel to the alarm analysis and noise reduction process, and does not rely on the on-site environment for adjustment as much as possible, and is preset as much as possible to improve the analysis effect and efficiency. Figure 2 , Figure 2 A flowchart of a network security alarm data analysis method provided in an embodiment of the present application.

[0037] In this embodiment, the method includes: Step S10: Acquire real-time alarm data and a pre-stored knowledge base, wherein the knowledge base includes expert experience rules, and the expert experience rules are used to determine whether the alarm data is generated by a real attack or a fake attack; It should be noted that this embodiment can be applied to the process of analyzing and judging network security alarms.

[0038] Real-time alarm data refers to data generated in real time by various security monitoring devices (such as firewalls, intrusion detection systems (IDS), and intrusion prevention systems (IPS)) during network system operation, indicating potential security threats or anomalies. For example, if a firewall detects a large number of connection requests from an external IP address to multiple ports of an internal server within a short period of time, it will generate an alarm data item, which may include information such as the attacking IP address, the attacked IP address, the attack time, and the attack port.

[0039] A pre-existing knowledge base is a pre-built database that stores knowledge related to cybersecurity alert analysis. This knowledge base includes expert experience rules, developed by cybersecurity experts based on practical experience and research into various attack patterns and abnormal behaviors. These rules serve as the logic for determining whether alert data indicates a real attack or a false one.

[0040] Step S20: determining target expert experience rules corresponding to the real-time alarm data from the knowledge base; It is understandable that the target expert experience rule refers to the expert experience rule in the pre-stored knowledge base that best matches the currently acquired real-time alarm data and is suitable for analyzing and judging the real-time alarm data.

[0041] For example, the real-time alarm data characteristics (such as attack type, attack source, and attack frequency) are compared and screened against the applicable conditions of the expert experience rules in the knowledge base. For example, if the real-time alarm data indicates the characteristics of a SQL injection attack, then the expert experience rules specifically for SQL injection attack judgment are found in the knowledge base and used as the target expert experience rules.

[0042] Suppose there's an expert rule in the knowledge base that states, "When a large number of requests containing special characters (such as single quotes and semicolons) are submitted to a specific parameter of a web application within a short period of time, a SQL injection attack may be detected." If real-time alert data shows that a web server has received a large number of parameter requests containing special characters within a short period of time, then this rule is the target expert rule.

[0043] Step S30: Input the real-time alarm data and the target expert experience rules into the analysis and judgment model to obtain an analysis and judgment result for the real-time alarm data, wherein the analysis and judgment result is used to indicate whether the real-time alarm data is data generated by a real attack or data generated by a non-real attack.

[0044] It should be noted that the analysis and judgment model can be a large language model LLM.

[0045] Real attack: This means that the situation reflected by the real-time alarm data is an actual security threat.

[0046] Non-real attack: This indicates that the real-time alarm data may be caused by system false alarms, misjudgments due to normal business operations, etc., and is not an actual security attack. For example, the alarm may be generated by packet retransmission due to network congestion, which is mistakenly interpreted by the security device as an attack.

[0047] In this embodiment, by obtaining real-time alarm data and a knowledge base containing expert experience rules, the corresponding target rules are determined, and then the judgment result of whether the alarm data is a real attack is obtained with the help of the judgment model. A complete judgment process based on RAG technology is constructed. On the one hand, expert experience rules are used to provide guidance for judgment, and combined with the judgment model to achieve automated and intelligent analysis, it can quickly and accurately determine the nature of network security alarm data, thereby improving the efficiency and accuracy of judgment. On the other hand, this process realizes the retrieval of relevant rule information from a large document collection (i.e., a knowledge base composed of expert experience rules), and uses this retrieved rule information to guide the generation of text, thereby improving the effectiveness of large model technology in noise reduction in alarm judgment.

[0048] Based on any of the above embodiments, the expert experience rules include structured language rules and / or unstructured knowledge, and the structured language rules include non-real attack rules and real attack rules defined from the network security domain dimension, attack feature dimension, and asset attribute dimension; the unstructured knowledge includes security analysis reports, threat intelligence documents, and historical research and judgment records.

[0049] It should be noted that structured language rules are rules organized and expressed in a structured manner, with a clear format and logical structure. Structured language rules define both non-real attack rules and real attack rules based on multiple dimensions, such as network security domains, attack characteristics, and asset attributes.

[0050] As an example, a structured language rule uses a structured statement such as "If (network security domain is internal network) and (attack characteristics are normal business traffic patterns) and (asset attributes are ordinary office terminals), then it is determined to be a non-real attack."

[0051] It should be understood that the network security domain dimension: network security domain is a logical area divided according to network security requirements, business functions and other factors. Different domains have different security policies and access control requirements.

[0052] As an example, consider the rule for a non-authentic attack: If the alarm data indicates that the attack source and target are both in the same internal security domain, and normal communication is allowed between devices within the domain, and there are no other abnormal features, it can be determined to be a non-authentic attack.

[0053] True attack rule: When an alarm indicates that traffic from an external untrusted network security domain attempts to access key assets in the internal core security domain, and this access does not comply with the preset security policy, it is determined to be a true attack.

[0054] It should be understood that the attack feature dimension: attack features refer to the unique patterns, behaviors or data characteristics exhibited by the attack behavior, such as specific protocol anomalies, data packet content patterns, etc.

[0055] As an example, consider the non-real attack rule: if the attack signature in an alert is a normal scanning behavior generated by a common network scanning tool (such as a port scanning frequency within a reasonable range and no subsequent malicious operations), it can be determined to be a non-real attack.

[0056] Real attack rule: When an alarm detects an attack with malware propagation characteristics, such as specific vulnerability exploit code, malicious file transfer mode, etc., it is determined to be a real attack.

[0057] It should be understood that the asset attribute dimension: asset attributes are characteristic descriptions of various assets in the network (such as servers, terminal devices, data, etc.), which may include the importance, purpose, operating system type, etc. of the assets.

[0058] As an example, consider the rule for a non-real attack: if the alarm is for an abandoned and no longer used asset, and there is no other associated abnormal activity, it can be determined to be a non-real attack.

[0059] True attack rule: When an alarm involves a core server of a critical business system, and the asset attributes of the server indicate that it stores sensitive data, and the attack characteristics indicate an attempt to steal data, it is determined to be a true attack.

[0060] Unstructured knowledge is, understandably, information without a predefined data model or fixed format. This includes security analysis reports, threat intelligence documents, and historical research records, which contain a wealth of cybersecurity experience and case studies.

[0061] Unstructured knowledge is used to supplement and expand structured language rules. By analyzing this unstructured knowledge, we can discover some potential attack patterns, emerging threat trends, and research and judgment experience in special situations, further improving the accuracy and comprehensiveness of research and judgment.

[0062] Security Analysis Report: A report written by cybersecurity experts after an in-depth analysis of a specific cybersecurity incident, system vulnerability, attack trend, etc. It may include the incident background, attack process analysis, impact assessment, and countermeasure recommendations.

[0063] Specifically, a security analysis report details a new ransomware attack method, including its propagation methods, infection characteristics, and encryption algorithm. When similar alert data appears, you can refer to the information in this report to determine whether it is a new type of ransomware attack.

[0064] Threat intelligence documents: A collection of information about potential or existing cyber threats, including threat sources, attack targets, attack methods, malware samples, etc. Threat intelligence documents can come from various sources, such as security vendors, open source communities, and government agencies.

[0065] Specifically, if a threat intelligence document indicates that a hacker group is conducting targeted attacks against businesses in a specific industry, using specific attack tools and techniques, then when the enterprise's network security system detects an alert that matches these characteristics, it can analyze and judge, combined with the threat intelligence document, whether it is an attack by the hacker group.

[0066] Historical analysis records: Information recorded after analyzing past network security alert data, including the basic situation of the alert, the analysis process, and the results.

[0067] Historical analysis records can serve as an experience library to provide reference for new alarm analysis.

[0068] Specifically, when handling a similar historical alarm, the analysis record at the time indicated that the alarm was a false alarm caused by a system configuration error, and detailed the specific configuration error and solution. When a similar alarm appears again, you can refer to the historical analysis record to quickly determine whether it is the same configuration problem, improving analysis efficiency.

[0069] In this embodiment, expert experience rules comprise structured language rules and unstructured knowledge. The structured language rules define the rules for both fake and real attacks from multiple dimensions, while the unstructured knowledge encompasses a wide range of security-related documents. This comprehensive rule-building approach enriches the knowledge base, enabling expert experience rules to describe network security characteristics from diverse perspectives and levels. This enhances the rules' coverage of various network security scenarios, thereby improving the comprehensiveness and accuracy of network security alert data analysis.

[0070] Based on any of the above embodiments, the expert experience rules are stored in the knowledge base in the form of vectors; before obtaining the real-time alarm data and the pre-stored knowledge base, the method further includes: The initial expert experience rules are vectorized to obtain the knowledge base consisting of the vectorized expert experience rules.

[0071] It should be noted that initial expert experience rules: expert experience rules that have not been vectorized, including structured language rules and unstructured knowledge. These rules exist in the form of natural language text, which is difficult for computers to directly understand and process.

[0072] Converting expert experience rules into vector form enables computers to process and analyze these rules more efficiently, facilitates mathematical operations and similarity comparisons, and thus quickly determines the degree of match between real-time alarm data and expert experience rules.

[0073] Specifically, the initial expert experience rules may be vectorized using a text embedding model to obtain the knowledge base consisting of the vectorized expert experience rules.

[0074] A text embedding model is a machine learning model that maps text data into a low-dimensional vector space. It captures semantic information in text and maps text with similar semantics to similar locations in the vector space. Examples of text embedding models include Word2Vec, GloVe, BERT, and their variants (such as the bge-small-zh model, a small version of the BAAI General Embedding series designed specifically for Chinese semantic vector representation. It offers high precision and efficiency, making it suitable for tasks such as Chinese search, classification, and clustering, and supports fast inference on CPUs and NPUs. Combined with frameworks like LangChain, bge-small-zh can be used to build knowledge bases, supporting sparse and multi-vector search without the high cost of training large vertical models). Text embedding models learn vector representations of words or sentences in text by training on large text corpora.

[0075] It is understandable that converting the initial expert experience rules (whether structured rule text or unstructured knowledge text) into vector form enables these rules to be represented and calculated in vector space.

[0076] For example, for structured language rules: Although structured language rules have a certain format and logic, they are still essentially text descriptions. The text embedding model will use the entire text as input. Through the model's internal encoding mechanism, it converts each word or phrase in the rule into a vector, and then combines these vectors to obtain a vector representation of the entire rule. For example, for the rule "If (network security domain is internal network) and (attack characteristics are normal business traffic patterns) and (asset attributes are ordinary office terminals), then determine it as a non-genuine attack," the model will analyze the keywords and semantic relationships within it to generate the corresponding vector.

[0077] For unstructured knowledge, such as security analysis reports and threat intelligence documents, text embedding models can perform preprocessing operations such as word segmentation and word vector initialization. Then, using a multi-layer neural network structure (such as the Transformer architecture), they deeply encode the text, extracting its semantic features and ultimately generating a fixed-dimensional vector to represent the entire document or report. For example, for a security analysis report, the model will analyze the report's event description, attack method analysis, impact assessment, and other content and convert them into vectors.

[0078] It should be understood that the knowledge base is a database that stores vectorized expert experience rules. This knowledge base can be updated and maintained as needed, such as adding new expert experience rules or deleting outdated rules.

[0079] Vectorized rules can better capture semantic information, making the analysis process more accurate.

[0080] In this embodiment, vectorizing the rules facilitates efficient similarity calculation and data understanding analysis by the computer, and can quickly and accurately match target expert experience rules corresponding to real-time alarm data from the knowledge base.

[0081] Based on any of the above embodiments, the expert experience rules include a plurality of rules; and determining the target expert experience rules corresponding to the real-time alarm data from the knowledge base includes: Vectorizing the real-time alarm data to obtain vectorized real-time alarm data; The similarity between the vectorized real-time alarm data and each of the expert experience rules is calculated, and the expert experience rule whose similarity meets the preset requirements is determined as the target expert experience rule.

[0082] It's important to note that network security scenarios are complex and diverse, requiring distinct judgment rules for different attack types, system environments, and business scenarios. Multiple expert experience rules can cover a wider range of network security situations, improving the accuracy and comprehensiveness of real-time alert data analysis. For example, different types of network attacks (such as DDoS attacks, SQL injection attacks, and Trojan infections) as well as assets of varying importance (such as core servers and standard terminal devices) all require corresponding rules for accurate judgment.

[0083] Because the expert experience rules in the knowledge base are stored in vector form, real-time alarm data must also be converted to vector form to calculate similarity with these rules. Vectorized real-time alarm data can be mathematically operated and compared with the expert experience rule vectors in vector space to determine the degree of similarity between them. Converting both expert experience rules and real-time alarm data to vector form for similarity calculations is important because similarity calculations can consider semantic connections between the two, rather than just superficial textual matching, thereby improving matching accuracy.

[0084] For example, if the real-time alert data contains natural language descriptions (such as text descriptions of attack types, event descriptions, etc.), text preprocessing is required, which may include word segmentation, stop word removal, stemming, and other operations. For example, for the description "Possible SQL injection attack detected" in the alert data, word segmentation may produce words such as "detected," "possible," "of," "SQL," "injection," and "attack." After removing the stop words "of" and "possible," the key information "SQL injection attack detected" is retained.

[0085] Furthermore, features can be extracted from processed text or directly from the structured information in the alarm data. Structured information such as the time of occurrence, source IP address, and destination port can be directly used as features; text information can be converted into numerical feature vectors.

[0086] It's understandable that by calculating similarity, we can quantify the degree of match between real-time alarm data and each expert experience rule. The higher the similarity, the more consistent the real-time alarm data is with the rule, and the more likely it is that the rule will be applied for analysis and judgment.

[0087] Optionally, the preset requirement may be a pre-set second threshold value, which is lower than the first threshold value. For example, the similarity threshold value is 0.8. In this case, only when the similarity between the real-time alarm data and a certain expert experience rule is greater than or equal to 0.8, the rule is determined as the target expert experience rule.

[0088] In this embodiment, the target rule is determined by vectorizing the real-time alarm data and calculating its similarity with each expert experience rule in the knowledge base. This allows the analysis of the real-time alarm data to leverage the computer's ability to efficiently process vector data, quickly and accurately screening out matching rules from numerous rules in the knowledge base.

[0089] Based on any of the above embodiments, the step of inputting the real-time alarm data and the target expert experience rules into the analysis and judgment model to obtain the analysis and judgment results for the real-time alarm data includes: Inputting the real-time alarm data and the target expert experience rule into the analysis and judgment model, and if the similarity between the vectorized real-time alarm data and the target expert experience rule exceeds a first threshold, determining the analysis and judgment conclusion in the target expert experience rule as the analysis and judgment result; If the similarity between the vectorized real-time alarm data and the target expert experience rule is lower than the first threshold, the analysis result and the confidence level of the analysis result are determined based on the real-time alarm data.

[0090] It should be noted that the first threshold value may be determined according to actual conditions. For example, the first threshold value may be 0.9.

[0091] In specific implementations, if the target expert experience rules clearly match (i.e., the similarity between the alarm data and the target expert experience rules exceeds a first threshold), the model can directly output a conclusion (which can be a true attack, a false alarm, or one requiring manual review). This conclusion is extracted from the target expert experience rules. If the identified target expert experience rules include multiple and conflicting rules, or if the identified target expert experience rules are ambiguous and do not point to a clear conclusion, the model needs to analyze the context of the alarm data and provide a confidence level (e.g., "80% chance of a false alarm"). In addition to the judgment result and confidence level, the model output can also include a brief explanation of the judgment.

[0092] In this embodiment, when the similarity exceeds a threshold, the conclusions of the target expert's experience rules are directly adopted, leveraging the reliability of expert experience to quickly and accurately produce results. When the similarity does not exceed the threshold, the data analysis and processing capabilities of the large model are utilized to produce a judgment result and confidence level. This ensures that judgments can be made even in complex or unknown situations, while also reflecting the credibility of the results through the confidence level.

[0093] Based on any of the above embodiments, the real-time alarm data includes an attacker identifier, a victim identifier, an alarm type, and attack characteristics.

[0094] It should be noted that an attacker identifier is information used to uniquely identify the entity that initiated a cyberattack. The purpose of determining the attacker identifier is to accurately track and locate the source of the attack. The attacker identifier can be the attacker's IP address. By analyzing the IP address's geographic location, network, and other information, we can initially determine the possible source and nature of the attack.

[0095] Victim Identifier: Information used to uniquely identify the entity that suffered a cyberattack. It helps determine the target of the attack and assess the impact on systems and services. The victim identifier can be the victim's IP address.

[0096] Alert type: A classification description of network security events that summarizes the nature and characteristics of the attack. Different alert types correspond to different types of network security threats. Alert types may include, but are not limited to: ① Intrusion detection alert: When a network intrusion detection system (IDS) or host intrusion detection system (HIDS) detects abnormal network activity or host behavior, an intrusion detection alert is triggered. For example, detecting behaviors such as port scans and buffer overflow attacks will generate corresponding intrusion detection alerts. ② Vulnerability exploitation alert: If an attacker exploits a known vulnerability in a system or application, the security system will issue a vulnerability exploitation alert. For example, detecting an attack request targeting an unpatched web application vulnerability will generate a vulnerability exploitation alert. ③ Malware alert: When security software detects the presence of malware (such as viruses, Trojans, worms, etc.) in the system, it will issue a malware alert. For example, antivirus software will trigger a malware alert if it discovers suspicious malicious code while scanning system files.

[0097] Attack characteristics: Information that describes the specific characteristics and patterns of attack behavior, which can provide a more detailed understanding of the attack methods and means. Attack characteristics may include, but are not limited to: ① Packet characteristics: including the packet's source port, destination port, protocol type, packet size, and packet content pattern. For example, some DDoS attacks send a large number of UDP packets with specific source and destination ports. By analyzing these packet characteristics, DDoS attacks can be identified; ② Behavioral characteristics: such as the attacker's operation sequence, access frequency, and access time. For example, when conducting a password cracking attack, a hacker may frequently try different username and password combinations. By analyzing this frequent attempt behavior, password cracking attacks can be detected; ③ File characteristics: If the attack involves the dissemination of malicious files, file characteristics may include the file's hash value, file size, file type, and code snippets within the file. For example, by calculating a file's hash value and comparing it to a database of known malicious file hash values, it is possible to quickly determine whether a file is malicious.

[0098] In a specific implementation, the real-time alarm data includes attacker identifier, victim identifier, alarm type, attack characteristics, attack response content, attack stage, alarm time, attacker information in threat intelligence data, security domain of the level in asset data, and asset type.

[0099] In this embodiment, the real-time alarm data clearly includes the attacker identifier, the victim identifier, the alarm type, and the attack characteristics, etc., providing comprehensive and detailed basic information for research and judgment.

[0100] Based on any of the above embodiments, the method further includes: When it is determined that the knowledge base does not have the target expert experience rules corresponding to the real-time alarm data, the knowledge base is adjusted to obtain an updated knowledge base, and the step of determining the target expert experience rules corresponding to the real-time alarm data from the knowledge base is returned to be executed.

[0101] It's important to note that the actual network environments, business systems, and security requirements vary significantly across different sites. A fixed knowledge base struggles to adapt to these unique environments, and to improve its comprehensiveness, it requires continuous adjustment and optimization based on site realities. New business launches, system upgrades, and evolving security threat landscapes can all lead to rule changes. To keep pace with these dynamics, the knowledge base also requires continuous optimization.

[0102] When it is determined that the target expert experience rules do not exist in the current knowledge base, it means that the rules in the current knowledge base cannot cover the network security situation represented by the real-time alarm data. Therefore, it is necessary to further adjust the knowledge base to enhance its analysis and judgment capabilities.

[0103] Optionally, the knowledge base can be adjusted by adding new rules. For example, new expert experience rules can be obtained from multiple sources, such as security researchers' analysis reports on new attacks, threat intelligence shared within the industry, and experience accumulated from historical research and analysis that is not yet included in the knowledge base.

[0104] After obtaining new expert experience rules, they need to be vectorized and stored in the knowledge base to complete the expansion of the knowledge base, so that they can be represented and calculated in the same vector space as the existing rule vectors in the knowledge base.

[0105] Optionally, the knowledge base can be adjusted by modifying existing rules. For example, if it is found that it is impossible to determine expert experience rules that meet the preset requirements from the knowledge base, or if it is found that the target expert experience rules determined during the analysis process are inaccurate or incomplete, the existing rules in the knowledge base need to be modified. For example, as the network environment changes, certain attack characteristics have changed, and the original rules may not accurately match the new attack situation. In this case, the structured language description or unstructured knowledge content of the rules can be modified according to the actual situation, and then re-vectorized and the corresponding rule vectors in the knowledge base can be updated.

[0106] Optionally, the knowledge base can be adjusted by deleting outdated existing rules. For example, over time, certain cybersecurity threats may no longer exist or become less significant, and the corresponding expert experience rules lose their value. For example, attack rules corresponding to widely patched vulnerabilities are no longer necessary to be retained in the knowledge base. In this case, these outdated rule vectors can be deleted from the knowledge base to reduce its size and improve retrieval and matching efficiency.

[0107] It should be understood that after the knowledge base is updated, the updated knowledge base is deployed and applied. Alarm data is acquired in real time and analyzed using the updated knowledge base until the target expert experience rules corresponding to the alarm data are determined from the updated knowledge base. Otherwise, the knowledge base is updated again until the conditions for loop termination are met (i.e., the target expert experience rules are determined from the updated knowledge base). This helps ensure that for any real-time alarm data, the appropriate rules can be found for accurate analysis by continuously adjusting the knowledge base. Optionally, real-time alarm data that could not match the target expert experience rules before can also be re-analyzed. This is because the updated knowledge base may contain new rules or modified rules that can match the real-time alarm data.

[0108] In this embodiment, the knowledge base is decoupled from the large model. When the knowledge base does not have corresponding target expert experience rules, the knowledge base is adjusted to overcome the limitation of the fixed knowledge base, so that the model can adapt to the ever-changing network security environment and ensure the model's ability to analyze and judge various types of alarm data.

[0109] Based on the method described in any of the above embodiments, the present application also provides Figure 3 A schematic diagram of the structure of an electronic device is shown in FIG. Figure 3 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its services. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the method described in any of the above embodiments.

[0110] Based on the method described in any of the above embodiments, the present application also provides a computer storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the method described in any of the above embodiments.

[0111] Based on the method described in any of the above embodiments, the present application further provides a computer program product comprising one or more computer programs or instructions. The computer program or instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. When the computer program is executed by a processor, the method described in any of the above embodiments is implemented.

[0112] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0113] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0114] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.

[0115] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.

[0116] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0117] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

Claims

1. A method for analyzing and judging network security alarm data, characterized in that: The method comprises: Acquiring real-time alarm data and a pre-stored knowledge base, wherein the knowledge base includes expert experience rules, and the expert experience rules are used to determine whether the alarm data is generated by a real attack or by a false attack; Determining target expert experience rules corresponding to the real-time alarm data from the knowledge base; The real-time alarm data and the target expert experience rules are input into the analysis and judgment model to obtain an analysis and judgment result for the real-time alarm data, and the analysis and judgment result is used to indicate whether the real-time alarm data is data generated by a real attack or data generated by a non-real attack.

2. The method according to claim 1, characterized in that The expert experience rules include structured language rules and / or unstructured knowledge, wherein the structured language rules include non-real attack rules and real attack rules defined from the network security domain dimension, attack feature dimension, and asset attribute dimension; the unstructured knowledge includes security analysis reports, threat intelligence documents, and historical research and judgment records.

3. The method according to claim 1 or 2, characterized in that The expert experience rules are stored in the knowledge base in the form of vectors; Before acquiring the real-time alarm data and the pre-stored knowledge base, the method further includes: The initial expert experience rules are vectorized to obtain the knowledge base consisting of the vectorized expert experience rules.

4. The method according to claim 3, characterized in that The expert experience rules include a plurality of rules; the target expert experience rules corresponding to the real-time alarm data are determined from the knowledge base, including: Vectorizing the real-time alarm data to obtain vectorized real-time alarm data; The similarity between the vectorized real-time alarm data and each of the expert experience rules is calculated, and the expert experience rule whose similarity meets the preset requirements is determined as the target expert experience rule.

5. The method according to claim 4, characterized in that The step of inputting the real-time alarm data and the target expert experience rules into the analysis and judgment model to obtain the analysis and judgment results for the real-time alarm data includes: Inputting the real-time alarm data and the target expert experience rule into the analysis and judgment model, and if the similarity between the vectorized real-time alarm data and the target expert experience rule exceeds a first threshold, determining the analysis and judgment conclusion in the target expert experience rule as the analysis and judgment result; If the similarity between the vectorized real-time alarm data and the target expert experience rule is lower than the first threshold, the analysis result and the confidence level of the analysis result are determined based on the real-time alarm data.

6. The method according to claim 1, characterized in that The real-time alarm data includes an attacker identifier, a victim identifier, an alarm type, and attack characteristics.

7. The method according to claim 1, characterized in that The method further comprises: When it is determined that the knowledge base does not have the target expert experience rules corresponding to the real-time alarm data, the knowledge base is adjusted to obtain an updated knowledge base, and the step of determining the target expert experience rules corresponding to the real-time alarm data from the knowledge base is returned to be executed.

8. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instruction, the method according to any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium, characterized in that Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Security information processing method and device, equipment and storage medium

    CN113452700A

  • Network alarm confirmation method, device and equipment

    CN117014287A

  • Operation and maintenance alarm analysis method, device and equipment based on large language model

    CN119065917A

  • Network security alarm automatic studying and judging method, device, equipment and medium

    CN119402282A

  • Network security alarm processing method, device, equipment, medium and program product

    CN119598023A

Cited By

  • Safety alarm noise reduction method based on rule and large model

    CN121350016A

  • A rule and large model-based security alarm noise reduction method

    CN121350016B

  • Multi-step attack detection method and device, medium and program product

    CN121509094A

  • Alarm log analysis method, electronic equipment, storage medium and program product

    CN121814393A

  • Alarm log analysis method, electronic device, storage medium and program product

    CN121814393B