Security alarm white list generation method and related device
By obtaining false alarm information and using natural language description to generate and optimize the alarm whitelist, the problem of high technical threshold and low applicability in generating security alarm whitelists in the existing technology is solved, and the effect of lowering the technical threshold and improving applicability is achieved.
Patent Information
- Application Number
- CN202510128188.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-27
AI Technical Summary
In the prior art, there are problems with high technical thresholds and low applicability in generating whitelists for security alarms. Users need to master complex technologies such as regular expressions, which are difficult to adapt to the needs of multiple fields and multiple scenarios.
By obtaining alarm information marked as false alarms, finding a matching alarm whitelist collection, using natural language description to generate alarm whitelists when there is no match, and using language models to optimize the whitelist, lowering the technical threshold and improving applicability.
It lowers the technical threshold for generating alarm whitelists, improves operational efficiency, facilitates understanding and maintenance, and improves the applicability of alarm whitelists, and can cover the needs of multiple fields and multiple scenarios.
Smart Images

Figure CN119938887A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for generating a security alert whitelist. Background Art
[0002] With the continuous development of computer technology, security protection products for security testing have emerged. Security protection products can detect computing devices such as computers and hosts or virtual operating environments such as containers to ensure operational safety.
[0003] Generally, security products support the configuration of alarm whitelists, which can be used to identify and exclude unimportant alarms. In other words, when the generated alarm information matches the alarm whitelist, the alarm information will not trigger an alarm event.
[0004] In the related art, users (such as security operators) usually configure an alarm whitelist in the form of a regular expression based on specific scenarios. However, the generation method of the above alarm whitelist has certain technical barriers and is only applicable to specific scenarios, so its applicability is low. Summary of the invention
[0005] The present application provides a method for generating a security alarm whitelist. The method can reduce the technical threshold for generating an alarm whitelist and improve the applicability of the alarm whitelist. The present application also provides a device, an electronic device, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0006] In a first aspect, the present application provides a method for generating a security alert whitelist, the method comprising:
[0007] Obtaining the first alarm information marked as a false alarm;
[0008] Searching the target alarm whitelist that matches the first alarm information from the alarm whitelist set;
[0009] In response to not finding a target alarm whitelist matching the first alarm information, obtaining a first alarm whitelist based on the first alarm information and described in a natural language;
[0010] Optimizing the first alarm whitelist using the first language model to obtain an optimized alarm whitelist;
[0011] The optimized alarm whitelist is added to the alarm whitelist set.
[0012] In a second aspect, the present application provides a device for generating a security alarm whitelist, the device comprising:
[0013] An acquisition module, used for acquiring first alarm information marked as a false alarm;
[0014] A search module, used to search for a target alarm whitelist matching the first alarm information from the alarm whitelist set;
[0015] A generating module, configured to obtain, in response to failing to find a target alarm whitelist matching the first alarm information, a first alarm whitelist described in a natural language and based on the first alarm information;
[0016] An optimization module, configured to optimize the first alarm whitelist by using a first language model to obtain an optimized alarm whitelist;
[0017] An adding module is used to add the optimized alarm whitelist to the alarm whitelist set.
[0018] In a third aspect, the present application provides an electronic device, the electronic device comprising a processor and a memory. The processor and the memory communicate with each other. The processor is used to execute instructions stored in the memory so that the electronic device executes the method for generating a security warning whitelist as in the first aspect or any implementation of the first aspect.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions instruct an electronic device to execute the method for generating a security alert whitelist described in the above-mentioned first aspect or any one of the implementations of the first aspect.
[0020] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method for generating a security alert whitelist as described in the first aspect or any one of the implementations of the first aspect.
[0021] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations.
[0022] It can be seen from the above technical solutions that this application has the following advantages:
[0023] The present application provides a method for generating a security alarm whitelist. The method first obtains first alarm information marked as a false alarm, then searches for a target alarm whitelist matching the first alarm information from an alarm whitelist set, and in response to failing to find a target alarm whitelist matching the first alarm information, obtains a first alarm whitelist based on the first alarm information and described in a natural language, optimizes the first alarm whitelist using a first language model to obtain an optimized alarm whitelist, and adds the optimized alarm whitelist to the alarm whitelist set.
[0024] In this method, for the first alarm information that is a false alarm, when the existing alarm whitelist cannot match it, it supports directly describing the alarm whitelist used to match the first alarm information in natural language, without the need to master regular expressions and other technologies, which can reduce the technical threshold for generating alarm whitelists, improve operational efficiency, and facilitate understanding and maintenance. At the same time, the alarm whitelist described in natural language can cover multiple fields and scenarios, improving the applicability of the alarm whitelist. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical method of the embodiments of the present application, the drawings required for use in the embodiments are briefly introduced below.
[0026] Figure 1 A schematic diagram of a process for generating a security warning whitelist provided in an embodiment of the present application;
[0027] Figure 2 A schematic diagram of a process for generating a security warning whitelist provided in an embodiment of the present application;
[0028] Figure 3 A schematic diagram of the structure of a device for generating a security alarm whitelist provided in an embodiment of the present application;
[0029] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The terms "first" and "second" in the embodiments of the present application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.
[0031] First, some technical terms and application scenarios involved in the embodiments of this application are introduced.
[0032] With the continuous development of computer technology, security protection products for security detection and ensuring the safe operation of computers, hosts and other computing devices or containers and other virtual operating environments have emerged. Security protection products can perform multi-faceted security detection for a variety of operating scenarios. For example, the security protection product can be a cloud workload protection platform (CWPP), which can detect host security and network security. For another example, the security protection product can be a host-based intrusion detection system (HIDS), which can perform security detection on the behavior and status of the computer system. For another example, the security protection product can be cloud security posture management (CSPM), which can assess and manage cloud security risks and identify configuration errors and security vulnerabilities in cloud environments.
[0033] In security protection products, false alarms are difficult to avoid. Efficient and accurate filtering of false alarms can improve the stability and security of the system. Generally, security protection products support the configuration of alarm whitelists. Alarm whitelists, also known as security alarm whitelists, can be used to identify and exclude unimportant alarms. Specifically, the alarm whitelist includes alarms that are configured as not requiring any operation. When the generated alarm information matches the alarm whitelist, the alarm information will be ignored and will not trigger an alarm event. The above process can also be referred to as "whitening the alarm information", and the alarm information can also be referred to as "whitened alarm information". In this way, by configuring the alarm whitelist, it is ensured that the focus of security protection is placed on important alarm information, thereby improving the efficiency of security operations and maintenance.
[0034] In the related art, the alarm whitelist is usually manually compiled by a user (eg, a security operator). Specifically, the user compiles static rules for whitening alarm information using regular expressions according to specific scenarios to form an alarm whitelist.
[0035] However, the above method has the following problems: First, the regular expression form of the alarm whitelist has certain technical barriers, low readability, and complex syntax increases the cost of managing and maintaining the alarm whitelist. In addition, static rules are usually designed for specific scenarios, which are difficult to adapt to the needs of multiple fields and scenarios, and cannot dynamically adapt to environmental changes, and have a limited scope of application.
[0036] In view of this, the present application provides a method for generating a security alarm whitelist, which method first obtains a first alarm information marked as a false alarm, then searches for a target alarm whitelist that matches the first alarm information from an alarm whitelist set, and in response to not finding a target alarm whitelist that matches the first alarm information, obtains a first alarm whitelist based on the first alarm information and described in a natural language, optimizes the first alarm whitelist using a first language model to obtain an optimized alarm whitelist, and adds the optimized alarm whitelist to the alarm whitelist set.
[0037] In this method, for the first alarm information that is a false alarm, when the existing alarm whitelist cannot match it, it supports directly describing the alarm whitelist used to match the first alarm information in natural language, without the need to master regular expressions and other technologies, which can reduce the technical threshold for generating alarm whitelists, improve operational efficiency, and facilitate understanding and maintenance. At the same time, the alarm whitelist described in natural language can cover multiple fields and scenarios, improving the applicability of the alarm whitelist.
[0038] To facilitate understanding of the technical solution provided by the embodiments of the present application, the following will be described in conjunction with the accompanying drawings. Figure 1 A flowchart of a method for generating an alarm whitelist is shown, and the method specifically includes:
[0039] S101: Acquire first alarm information marked as a false alarm.
[0040] Among them, the first alarm information can be understood as any alarm information received by the security protection product. For the first alarm information, the user (for example, a security operator) or the security protection product can determine whether it is a false alarm and generate an alarm tag associated with the first alarm information. In an embodiment of the present application, the alarm tag associated with the first alarm information indicates that the first alarm information is a false alarm, that is, the first alarm information is not a real alarm.
[0041] In an embodiment of the present application, the first alarm information may be composed of multiple alarm fields. For example, the first alarm information may include a command line field, a parent process command line field, a process group command line field, a process tree information field, a runtime link field, an execution directory field, etc.
[0042] S102: Searching for a target alarm whitelist that matches the first alarm information from the alarm whitelist set.
[0043] The alarm whitelist set stores multiple existing alarm whitelists. In other words, the alarm whitelist set can be understood as a set for storing existing alarm whitelists in security protection products. The target alarm whitelist can be understood as an alarm whitelist in the alarm whitelist set that can whitelist the first alarm information.
[0044] Combine the following Figure 2 For a detailed description of the matching process between the first alarm information and the alarm whitelist set, see Figure 2 A flow chart of a method for generating a security alarm whitelist is shown. In some possible implementations, the target alarm whitelist is searched by "first recalling a set of candidate alarm whitelists, and then matching from the set of candidate alarm whitelists". In this way, the alarm whitelist set is initially filtered, and then precise matching is performed to improve matching efficiency and matching accuracy.
[0045] The candidate alarm whitelist set stores a candidate alarm whitelist, which can be understood as an alarm whitelist used to match the first alarm information. That is, according to the first alarm information, screening is performed from the alarm whitelist set, and candidate alarm whitelists similar to the first alarm information are dynamically recalled to form a candidate alarm whitelist set, and the candidate alarm whitelist set is used to perform deep matching on the first alarm information.
[0046] In some embodiments, the alarm whitelist set may be a vector library, in which case, an embedded vector related to a key alarm field in the first alarm information is determined, as well as an embedded vector of each alarm whitelist in the alarm whitelist set, and the similarity between the embedded vector related to the key alarm field and the embedded vector of each alarm whitelist in the alarm whitelist set is determined. Then, the alarm whitelist in the alarm whitelist set whose similarity is greater than a similarity threshold is determined as an alarm whitelist in the candidate alarm whitelist set, and a target alarm whitelist matching the first alarm information is searched from the candidate alarm whitelist set.
[0047] Among them, the key alarm field can be understood as an alarm field that can be used to represent the characteristics of the first alarm information. For example, the key alarm field can be a rule field that triggers the first alarm information, a summary field of the first alarm information, etc.
[0048] The embodiment of the present application does not limit the method of determining the embedding vector associated with the key alarm field and the embedding vector of each alarm whitelist in the alarm whitelist set. For example, the field value of the key alarm field in the first alarm information and each alarm whitelist in the alarm whitelist set can be processed using word embedding technology to obtain the embedding vector associated with the key alarm field and the embedding vector of each alarm whitelist in the alarm whitelist set. For another example, the field value of the key alarm field in the first alarm information and each alarm whitelist in the alarm whitelist set can be processed using a vector model to obtain the embedding vector associated with the key alarm field and the embedding vector of each alarm whitelist in the alarm whitelist set.
[0049] The embodiments of the present application do not limit the method for determining the similarity between the embedding vector associated with the key alarm field and the embedding vector of each alarm whitelist in the alarm whitelist set. For example, the cosine similarity between the embedding vector associated with the key alarm field and the embedding vector of each alarm whitelist in the alarm whitelist set can be calculated.
[0050] By configuring different similarity thresholds and based on different recall rules, candidate alarm whitelists with high similarity to the first alarm information are dynamically recalled from the alarm whitelist set, thereby forming a candidate alarm whitelist set and completing the recall of the candidate alarm whitelist set.
[0051] In this way, through the above method of determining the embedding vector and calculating the similarity to recall the candidate alarm whitelist set, the generalization and matching efficiency are improved, and at the same time, different security protection products can be flexibly adapted.
[0052] S103: In response to not finding a target alarm whitelist matching the first alarm information, obtaining a first alarm whitelist based on the first alarm information and described in a natural language.
[0053] Failure to find a target alarm whitelist matching the first alarm information indicates that there is no alarm whitelist in the alarm whitelist that can whitelist the first alarm information.
[0054] In the embodiment of the present application, failure to find a target alarm whitelist matching the first alarm information can be divided into two situations, which are described below respectively:
[0055] Case 1: Not recalled from the alarm whitelist set to the candidate alarm whitelist set.
[0056] Specifically, in response to the candidate alarm whitelist set being empty, a first alarm whitelist based on the first alarm information and described in a natural language is obtained.
[0057] In other words, there is no alarm whitelist similar to the first alarm information in the alarm whitelist set, for example, there is no alarm whitelist with a similarity greater than a similarity threshold. In this case, the first alarm information cannot be matched with the alarm whitelist in the alarm whitelist set.
[0058] Case 2: the candidate alarm whitelist in the candidate alarm whitelist set cannot match the first alarm information.
[0059] Specifically, in response to the candidate whitelist set including at least one second alarm whitelist, a second language model is used to match the first alarm information based on at least one second alarm whitelist to obtain a third matching result. In response to the third matching result indicating that the first alarm information does not match each alarm whitelist in at least one second alarm whitelist, a first alarm whitelist based on the first alarm information and described in natural language is obtained.
[0060] In other words, in the alarm whitelist set, there is an alarm whitelist that is similar to the first alarm information (i.e., the second alarm whitelist), but after matching the first alarm information with the second alarm whitelist, the first alarm information does not meet the whitelisting conditions of the second alarm whitelist, and the first alarm information does not match the second alarm whitelist. In this case, the first alarm information cannot be matched with the alarm whitelist in the alarm whitelist set.
[0061] In an embodiment of the present application, the first alarm information is matched with at least one second alarm whitelist by means of a second language model. The second language model can be understood as a language model used to match the first alarm information and the second alarm whitelist, and the second language model can be a language model with natural language processing capabilities, capable of understanding the meaning of natural language, and processing different types of natural language tasks. For example, the second language model can be a deep learning model trained using text data.
[0062] In some possible implementations, the second language model matches the first alarm information with the second alarm whitelist based on prompt learning. Among them, prompt words can be used to guide the language model to perform specific outputs in generative tasks (such as text generation tasks, question-answering tasks, and dialogue tasks). By configuring prompt words, the language model is helped to understand the background and requirements of the task, so that the language model can handle different types of natural language processing tasks without retraining the language model, thereby increasing the scalability and flexibility of the language model.
[0063] In a specific implementation, in response to the candidate whitelist set including at least one second warning whitelist, a second prompt word is generated, the second prompt word is sent to the second language model, and a third matching result returned by the second language model is received.
[0064] The second prompt word includes: a key alarm field in the first alarm information, at least one second alarm whitelist, and information for indicating whether the key alarm field in the first alarm information matches the at least one second alarm whitelist.
[0065] In some embodiments, the key alarm fields in the first alarm information may include an alarm field representing command line parameters, an alarm field representing command line parameters of a parent process, an alarm field representing a process group, an alarm field representing the operation that generates the alarm information, and the like.
[0066] By configuring the above information in the second prompt word, the second language model can compare the alarm content of the first alarm information with each second alarm whitelist one by one based on the prompt capability of the second prompt word, determine whether the alarm content of the first alarm information meets the whitelisting logic of the second alarm whitelist, and determine the third matching result.
[0067] By using the second language model to match the first alarm information with the second alarm whitelist, and relying on the natural language processing capabilities of the second language model, we can understand the complex whitelisting logic of the second alarm and the different scenarios and fields involved in the first alarm information, making the third matching result more accurate.
[0068] Further, in response to the third matching result indicating that the first alarm information matches a target alarm whitelist in at least one second alarm whitelist, whitening processing is performed on the first alarm information.
[0069] That is to say, when there is a target alarm whitelist in the candidate alarm whitelist set that matches the first alarm information, it indicates that the target alarm whitelist is found in the alarm whitelist set and matches the first alarm information, and the security protection product marks the first alarm information as whitened. In this way, subsequent false alarm interference is reduced and the operating efficiency of the security protection product is improved.
[0070] It is understandable that the alarm whitelist in the security protection product should whiten the alarm information that is a false alarm to reduce the false alarm situation. When the existing alarm whitelist of the security protection product cannot whiten the first alarm information, but the first alarm information is a false alarm, it indicates that there is a new demand for the alarm whitelist, and a new alarm whitelist should be added to whiten the first alarm information.
[0071] In the embodiment of the present application, it is supported to add a first alarm whitelist for matching the first alarm information by inputting natural language. That is to say, the whitelist logic or matching rule associated with the first alarm information is described in natural language, and the first alarm whitelist is described in natural language to represent the natural language content. In this way, there is no need to write an alarm whitelist in the form of a regular expression, which facilitates the rapid expansion of the alarm whitelist and reduces the complexity of alarm whitelist generation.
[0072] The embodiment of the present application does not limit the method of obtaining the first alarm whitelist described in natural language. For example, the first alarm whitelist can be input by a user (such as a security operator). In this way, the user does not need to learn complex regular expression syntax, and can conveniently add new alarm whitelists by simply describing the whitelisting scenarios and whitelisting logic in natural language. For another example, the first alarm whitelist can also be automatically generated by a model (such as a language model) that has the ability to analyze alarm information, that is, the first alarm information is analyzed by a model that has the ability to analyze alarm information, and the first alarm whitelist described in natural language for matching the first alarm information is automatically output, thereby improving the automation level of security protection products.
[0073] S104: Optimizing the first alarm whitelist using the first language model to obtain an optimized alarm whitelist.
[0074] Considering that the accuracy of the first alarm whitelist described in natural language may be weak, the first alarm whitelist is optimized with the help of the first language model, so that the optimized alarm whitelist is improved in terms of logic and fluency.
[0075] In some possible implementations, the first language model optimizes the first warning whitelist based on prompt learning. Specifically, a first prompt word is generated, the first prompt word is sent to the first language model, and the optimized warning whitelist returned by the first language model is received.
[0076] The first prompt word includes the first alarm whitelist and rule logic for indicating extraction of the first alarm whitelist, and information for optimizing the first alarm whitelist based on the rule logic.
[0077] By configuring the above information in the first prompt word, the first language model can extract the core rule logic (such as whitelisting logic, matching rules, etc.) in the first alarm whitelist based on the prompt capability of the first prompt word, and optimize the natural language expression of the first alarm whitelist while ensuring that the core rule logic remains unchanged, thereby improving the readability and logic of the first alarm whitelist and generating an optimized alarm whitelist.
[0078] For example, the optimized alarm whitelist can be: "Determine whether the access IP is from the internal network segment and the accessed file does not belong to the sensitive directory. When the access IP is from the internal network segment and the accessed file does not belong to the sensitive directory, add it to the whitelist", "If the downloaded file is from an educational website and the size is less than 1GB, add it to the whitelist", "Ignore login alarms outside working hours, except on weekdays", "Ignore all alarms for visiting well-known websites".
[0079] In this way, with the help of the natural language processing capabilities of the first language model, the first alarm whitelist is automatically and quickly optimized to improve the accuracy of the first alarm whitelist, so that the first alarm whitelist can be adapted to alarm whitelisting in different fields, different scenarios, and different security protection products.
[0080] Furthermore, in order to ensure that the optimized first alarm whitelist has a good matching effect, in the embodiment of the present application, the optimized first alarm whitelist may also be verified. Figure 2 As shown, the optimized alarm whitelist is matched with the first alarm information to obtain a first matching result, and second alarm information similar to the first alarm information and marked as a non-false alarm (i.e., a real alarm) is obtained, and the optimized alarm whitelist is matched with the second alarm information to obtain a second matching result.
[0081] The second warning information can be obtained from a historical risk warning library, which stores warning information marked as non-false alarms. Specifically, the similarity between the first warning information and the warning information in the historical risk warning library is calculated, and the warning information whose similarity meets the set conditions (for example, the similarity is greater than the second similarity threshold) is determined as the second warning information.
[0082] That is to say, two aspects of verification are performed on the optimized alarm whitelist: on the one hand, since the first alarm whitelist is used to match the first alarm information, it is verified whether the optimized alarm whitelist can match the first alarm information, that is, whether the optimized alarm whitelist can add the first alarm information to whitelist and obtain the first matching result. On the other hand, since the first alarm whitelist should not add the real alarm, it is verified whether the optimized alarm whitelist can match the second alarm information that is similar to the first alarm information and marked as a real alarm, that is, whether the optimized alarm whitelist can add the second alarm information to whitelist and obtain the second matching result.
[0083] In this way, it is verified whether the optimized alarm whitelist can correctly match the alarm information, and whether there are any mistaken whitelistings in the optimized alarm whitelist. The optimized alarm whitelist is verified from two aspects to ensure that the rule logic of the optimized alarm whitelist is unambiguous, and can accurately match the alarm information that should be whitelisted, and at the same time, real alarms will not be mistakenly whitelisted.
[0084] Further, in response to the first matching result indicating that the optimized alarm whitelist does not match the first alarm information, or the second matching result indicating that the optimized alarm whitelist matches the second alarm information, a prompt message is sent, wherein the prompt message is used to indicate that the optimized alarm whitelist has not passed the verification.
[0085] That is, when the optimized alarm whitelist cannot add the first alarm information, or the optimized alarm whitelist can add the second alarm information, it indicates that the optimized alarm whitelist has not passed the verification and needs to be revised. In this case, a prompt message is presented so that the first alarm whitelist can be modified in time.
[0086] In some embodiments, the prompt message may also include optimization suggestions, such as modification suggestions generated by combining historical alarm information, the first matching result, and the second matching result, to assist in quickly improving the alarm whitelist.
[0087] For example, the optimized alarm whitelist is "When the alarm argv contains fields such as curl / wget, determine whether the command line downloads data from a well-known platform. If so, whitelist it." The prompt message is "After querying the historical risk alarm information, there is alarm_id: xxxx argv:wget www.ABCD.com / xxxx. After verification, the alarm whitelist can match the historical risk alarm information and needs to be re-entered. Modification suggestion: The attacker often remotely implants code through code repositories such as ABCD. It is recommended to remove ABCD from well-known platforms." In this way, the optimized alarm whitelist is modified according to the modification suggestions in the prompt message to obtain the final alarm whitelist "When the alarm information argv contains curl / wget fields, determine whether the command line downloads data from a well-known platform that cannot be a code / mirror repository. If so, whitelist it."
[0088] S105: Add the optimized alarm whitelist to the alarm whitelist set.
[0089] By storing the optimized alarm whitelist in the alarm whitelist set, the number of alarm whitelists in the alarm whitelist set is enriched, and the whitelisting capability of the security protection product is improved.
[0090] In some embodiments, the optimized alarm whitelist is verified. In this case, in response to the first matching result indicating that the optimized alarm whitelist matches the first alarm information, and the second matching result indicating that the optimized alarm whitelist does not match the second alarm information, the optimized alarm whitelist is added to the alarm whitelist set.
[0091] That is to say, when the optimized alarm whitelist can whitelist the first alarm information, and the optimized alarm whitelist does not whitelist the second alarm information, it indicates that the optimized alarm whitelist has passed the verification, ensuring that the alarm whitelists in the alarm whitelist set are all verified high-quality alarm whitelists.
[0092] In some possible implementations, the alarm whitelist set exists in the form of a vector library. In this case, the embedded vector corresponding to the optimized alarm whitelist is determined, and the embedded vector corresponding to the optimized alarm whitelist is added to the alarm whitelist set.
[0093] In other words, the alarm whitelist set stores a natural language alarm whitelist represented in the form of an embedded vector. In this way, the alarm whitelist set can be flexibly adapted to different security protection products to improve the scope of application.
[0094] In addition, the alarm whitelist set can also be updated iteratively. For example, based on one or more of the newly added whitelisted alarm information, the alarm information belonging to the real alarm, and user feedback, the alarm whitelist in the alarm whitelist set is updated, and it is continuously optimized iteratively to improve the accuracy, generalization, and adaptability of the alarm whitelist to new alarm scenarios.
[0095] In this method, for the first alarm information that is a false alarm, when the existing alarm whitelist cannot match it, it supports directly describing the alarm whitelist used to match the first alarm information in natural language, without the need to master regular expressions and other technologies, which can reduce the technical threshold for generating alarm whitelists, improve operational efficiency, and facilitate understanding and maintenance. At the same time, the alarm whitelist described in natural language can cover multiple fields and scenarios, improving the applicability of the alarm whitelist.
[0096] Combination of the above Figure 1 and Figure 2 The method for generating the security alarm whitelist provided in the embodiment of the present application is introduced in detail. The apparatus and device provided in the embodiment of the present application will be introduced in conjunction with the accompanying drawings.
[0097] See also Figure 3 The schematic diagram of the structure of the device for generating the security alarm whitelist shown in FIG. 30 includes:
[0098] An acquisition module 301 is used to acquire first alarm information marked as a false alarm;
[0099] A search module 302 is used to search a target alarm whitelist that matches the first alarm information from the alarm whitelist set;
[0100] A generating module 303 is configured to obtain, in response to failing to find a target alarm whitelist matching the first alarm information, a first alarm whitelist described in a natural language and based on the first alarm information;
[0101] An optimization module 304, configured to optimize the first alarm whitelist by using a first language model to obtain an optimized alarm whitelist;
[0102] The adding module 305 is used to add the optimized alarm whitelist to the alarm whitelist set.
[0103] In some possible implementations, the optimization module 304 is specifically configured to:
[0104] Generate a first prompt word; wherein the first prompt word includes the first alarm whitelist and a rule logic for indicating extraction of the first alarm whitelist, and information for optimizing the first alarm whitelist based on the rule logic;
[0105] The first prompt word is sent to a first language model, and an optimized alarm whitelist returned by the first language model is received.
[0106] In some possible implementations, the apparatus 30 further includes a verification module, and the verification module is configured to:
[0107] Matching the optimized alarm whitelist with the first alarm information to obtain a first matching result; and,
[0108] Acquire second alarm information that is similar to the first alarm information and marked as a non-false alarm, and match the optimized alarm whitelist with the second alarm information to obtain a second matching result;
[0109] The adding module 305 is specifically used for:
[0110] In response to the first matching result indicating that the optimized alarm whitelist matches the first alarm information, and the second matching result indicating that the optimized alarm whitelist does not match the second alarm information, the optimized alarm whitelist is added to the alarm whitelist set.
[0111] In some possible implementations, the apparatus 30 further includes a prompt module, and the prompt module is configured to:
[0112] In response to the first matching result indicating that the optimized alarm whitelist does not match the first alarm information, or the second matching result indicating that the optimized alarm whitelist matches the second alarm information, a prompt message is sent; wherein the prompt message is used to prompt that the optimized alarm whitelist has not passed verification.
[0113] In some possible implementations, the adding module 305 is specifically used to:
[0114] Determine the embedding vector corresponding to the optimized alarm whitelist;
[0115] The embedding vector corresponding to the optimized alarm whitelist is added to the alarm whitelist set.
[0116] In some possible implementations, the search module 302 is specifically configured to:
[0117] Determining an embedding vector associated with a key alarm field in the first alarm information, and determining an embedding vector for each alarm whitelist in the alarm whitelist set;
[0118] Determining the similarity between the embedding vector associated with the key alarm field and the embedding vectors of each alarm whitelist in the alarm whitelist set;
[0119] Determine the alarm whitelists in the alarm whitelist set whose similarity is greater than the similarity threshold as the alarm whitelists in the candidate alarm whitelist set;
[0120] A target alarm whitelist matching the first alarm information is searched from the candidate alarm whitelist set.
[0121] In some possible implementations, the generating module 303 is specifically used to:
[0122] In response to the candidate alarm whitelist set being empty, obtaining a first alarm whitelist described in natural language and based on the first alarm information; or,
[0123] In response to the candidate whitelist set including at least one second alarm whitelist, using a second language model, based on the at least one second alarm whitelist, the first alarm information is matched to obtain a third matching result; in response to the third matching result indicating that the first alarm information does not match each alarm whitelist in the at least one second alarm whitelist, a first alarm whitelist based on the first alarm information and described in a natural language is obtained.
[0124] In some possible implementations, the apparatus 30 further includes a whitening module, and the whitening module is configured to:
[0125] In response to the third matching result indicating that the first alarm information matches the target alarm whitelist in the at least one second alarm whitelist, whitening processing is performed on the first alarm information.
[0126] In some possible implementations, the verification module is specifically used to:
[0127] In response to the candidate whitelist set including at least one second alarm whitelist, generating a second prompt word; wherein the second prompt word includes: a key alarm field in the first alarm information, the at least one second alarm whitelist, and information for indicating whether the key alarm field in the first alarm information matches the at least one second alarm whitelist;
[0128] The second prompt word is sent to a second language model, and a third matching result returned by the second language model is received.
[0129] The device 30 for generating the security warning whitelist according to the embodiment of the present application may correspond to the method described in the embodiment of the present application, and the above and other operations and / or functions of each module / unit of the device 30 for generating the security warning whitelist are respectively to realize Figure 1 For the sake of brevity, the corresponding processes of each method in the illustrated embodiment are not described in detail here.
[0130] The present application also provides an electronic device. The electronic device is specifically used to implement Figure 3 The function of the device 30 for generating the security warning whitelist in the illustrated embodiment.
[0131] Figure 4 A schematic diagram of the structure of an electronic device 400 is provided. Figure 4 As shown, the electronic device 400 includes a bus 401, a processor 402, a communication interface 403 and a memory 404. The processor 402, the memory 404 and the communication interface 403 communicate with each other via the bus 401.
[0132] The bus 401 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0133] The processor 402 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0134] The communication interface 403 is used for communicating with the outside. For example, the communication interface 403 can be used for communicating with a terminal.
[0135] The memory 404 may include a volatile memory, such as a random access memory (RAM). The memory 404 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0136] The memory 404 stores executable codes, and the processor 402 executes the executable codes to perform the aforementioned method for generating the security warning whitelist.
[0137] Specifically, in implementing Figure 3 In the case of the embodiment shown, and Figure 3 When each module or unit of the security warning whitelist generation device 30 described in the embodiment is implemented by software, the execution Figure 3 The software or program code required for the functions of each module / unit in the system may be partially or completely stored in the memory 404. The processor 402 executes the program code corresponding to each unit stored in the memory 404 to execute the aforementioned method for generating the security warning whitelist.
[0138] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned method for generating a security alarm whitelist of the device 30 for generating a security alarm whitelist.
[0139] The embodiment of the present application further provides a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the process or function described in the embodiment of the present application is generated in whole or in part.
[0140] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer or data center to another website, computer or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0141] When the computer program product is executed by a computer, the computer executes any of the aforementioned methods for generating a security warning whitelist. The computer program product may be a software installation package, and when any of the aforementioned methods for generating a security warning whitelist is needed, the computer program product may be downloaded and executed on a computer.
[0142] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to each embodiment of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0143] The units involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the unit / module does not, in some cases, constitute a limitation on the unit itself.
[0144] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0145] In the context of the present application embodiment, machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include but is not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the above. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0146] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0147] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0148] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0149] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0150] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating a security warning whitelist, characterized in that: The method comprises: Obtaining the first alarm information marked as a false alarm; Searching the target alarm whitelist that matches the first alarm information from the alarm whitelist set; In response to not finding a target alarm whitelist matching the first alarm information, obtaining a first alarm whitelist based on the first alarm information and described in a natural language; Optimizing the first alarm whitelist using the first language model to obtain an optimized alarm whitelist; The optimized alarm whitelist is added to the alarm whitelist set.
2. The method according to claim 1, characterized in that The optimizing the first alarm whitelist by using the first language model to obtain an optimized alarm whitelist includes: Generate a first prompt word; wherein the first prompt word includes the first alarm whitelist and a rule logic for indicating extraction of the first alarm whitelist, and information for optimizing the first alarm whitelist based on the rule logic; The first prompt word is sent to a first language model, and an optimized alarm whitelist returned by the first language model is received.
3. The method according to claim 1, characterized in that After optimizing the first alarm whitelist by using the first language model to obtain an optimized alarm whitelist, the method further includes: Matching the optimized alarm whitelist with the first alarm information to obtain a first matching result; and, Acquire second alarm information that is similar to the first alarm information and marked as a non-false alarm, and match the optimized alarm whitelist with the second alarm information to obtain a second matching result; The adding the optimized alarm whitelist to the alarm whitelist set includes: In response to the first matching result indicating that the optimized alarm whitelist matches the first alarm information, and the second matching result indicating that the optimized alarm whitelist does not match the second alarm information, the optimized alarm whitelist is added to the alarm whitelist set.
4. The method according to claim 3, characterized in that The method further comprises: In response to the first matching result indicating that the optimized alarm whitelist does not match the first alarm information, or the second matching result indicating that the optimized alarm whitelist matches the second alarm information, a prompt message is sent; wherein the prompt message is used to prompt that the optimized alarm whitelist has not passed verification.
5. The method according to claim 1, characterized in that The adding the optimized alarm whitelist to the alarm whitelist set includes: Determine the embedding vector corresponding to the optimized alarm whitelist; The embedding vector corresponding to the optimized alarm whitelist is added to the alarm whitelist set.
6. The method according to any one of claims 1 to 5, characterized in that: The searching, from the alarm whitelist set, for a target alarm whitelist matching the first alarm information comprises: Determining an embedding vector associated with a key alarm field in the first alarm information, and determining an embedding vector for each alarm whitelist in the alarm whitelist set; Determining the similarity between the embedding vector associated with the key alarm field and the embedding vectors of each alarm whitelist in the alarm whitelist set; Determine the alarm whitelists in the alarm whitelist set whose similarity is greater than the similarity threshold as the alarm whitelists in the candidate alarm whitelist set; A target alarm whitelist matching the first alarm information is searched from the candidate alarm whitelist set.
7. The method according to claim 6, characterized in that In response to not finding a target alarm whitelist matching the first alarm information, obtaining a first alarm whitelist based on the first alarm information and described in a natural language, includes: In response to the candidate alarm whitelist set being empty, obtaining a first alarm whitelist described in natural language and based on the first alarm information; or, In response to the candidate whitelist set including at least one second alarm whitelist, using a second language model, based on the at least one second alarm whitelist, the first alarm information is matched to obtain a third matching result; in response to the third matching result indicating that the first alarm information does not match each alarm whitelist in the at least one second alarm whitelist, a first alarm whitelist based on the first alarm information and described in a natural language is obtained.
8. The method according to claim 7, characterized in that The method further comprises: In response to the third matching result indicating that the first alarm information matches the target alarm whitelist in the at least one second alarm whitelist, whitening processing is performed on the first alarm information.
9. The method according to claim 7, characterized in that: In response to the candidate whitelist set including at least one second alarm whitelist, using a second language model and based on the at least one second alarm whitelist, matching the first alarm information to obtain a third matching result includes: In response to the candidate whitelist set including at least one second alarm whitelist, generating a second prompt word; wherein the second prompt word includes: a key alarm field in the first alarm information, the at least one second alarm whitelist, and information for indicating whether the key alarm field in the first alarm information matches the at least one second alarm whitelist; The second prompt word is sent to a second language model, and a third matching result returned by the second language model is received.
10. A device for generating a security warning whitelist, characterized in that: The device comprises: An acquisition module, used for acquiring first alarm information marked as a false alarm; A search module, used to search for a target alarm whitelist matching the first alarm information from the alarm whitelist set; A generating module, configured to obtain, in response to failing to find a target alarm whitelist matching the first alarm information, a first alarm whitelist described in a natural language and based on the first alarm information; An optimization module, configured to optimize the first alarm whitelist by using a first language model to obtain an optimized alarm whitelist; An adding module is used to add the optimized alarm whitelist to the alarm whitelist set.
11. An electronic device, characterized in that: The electronic device comprises a processor and a memory; The processor is configured to execute instructions stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The method comprises instructions, wherein the instructions instruct an electronic device to execute the method as claimed in any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product comprises computer readable instructions for implementing the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Alarm information processing method, system, equipment and medium
CN118337422A
File protection method and device, equipment, storage medium and program product
CN119312320A
System and method for providing whitelist functionality for use with a cloud computing environment
US20140075520A1