Log-based Event Acquisition Method, Device, Equipment and Medium

A rule tree structure optimizes security log analysis by addressing inefficiencies in existing methods, enhancing event detection through structured rule tree construction and matching.

CN114297046BActive Publication Date: 2025-07-15CHINA TELECOM CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111658789.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-15
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art has problems in the security log matching with high time complexity, inability to handle rule dependencies and multiple log matching results, resulting in inefficiency.

Method used

Build a rule tree, optimize the rule tree structure through node relationships, use rule tree files to match log data, and correspond to the node type and rule dependencies, and optimize the matching path.

Benefits of technology

It improves the efficiency of secure log matching, reduces unnecessary node matching, and supports efficient matching in complex functions and multi-rule scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114297046B_ABST
    Figure CN114297046B_ABST
Patent Text Reader

Abstract

In an embodiment of the present application, a method, apparatus, device, and medium for obtaining events based on logs are provided. By obtaining a rule tree file and constructing a rule tree based on the rule tree file; wherein, each node of the rule tree corresponds to the construction of a rule; the node relationships between the nodes correspond to the dependency relationships between the rules; inputting log data into the root node of the rule tree and performing matching of corresponding rules at each node along the matching path; obtaining events according to the matching results of the nodes. By establishing a rule-related rule tree, log data can pass through each node along the matching path. Since the node relationships correspond to the dependency relationships between the rules, the efficiency is effectively improved compared to the previous method of traversing the rule set. In addition, when constructing the rule tree, the rule tree structure can be optimized by forming multiple nodes through, for example, node balancing and regular text splitting, which is also more conducive to improving the matching efficiency; or the node types can be set according to the rule requirements to call corresponding methods during matching to meet the rule requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular, to a method, device, equipment and medium for obtaining events based on logs. Background Art

[0002] Security event recognition is the basis for prediction and alert in the security field. Security researchers extract rules into regular expressions to match security logs and identify and output alert events. Currently, the general method for security log matching is as follows: For a single log, all rules are traversed, and the results are respectively matched. The event information of the corresponding rules is output. This method cannot efficiently solve the problem of rule matching for security logs, which is reflected in the following aspects:

[0003] 1. Each time the entire rule set is traversed for matching, the time complexity is high;

[0004] 2. If there are dependencies between rules, ordered matching is required, and the result of the subsequent rule depends on the result of the preceding rule and its own matching result. Traversing one by one cannot handle this dependency;

[0005] 3. For rules that determine events based on the matching results of multiple logs, the current method cannot meet the requirements.

[0006] Summary of the Invention

[0007] In view of the above-mentioned disadvantages of the prior art, the objective of this application is to provide a method, device, equipment and medium for obtaining events based on logs, which constructs a rule tree with optimized paths according to each rule for matching to solve the problems of the prior art.

[0008] The first aspect of this application provides a method for obtaining events based on logs, including: obtaining a rule tree file, and constructing a rule tree based on the rule tree file; wherein, each node of the rule tree corresponds to a rule; the node relationship between each node corresponds to the dependency relationship between rules; inputting log data into the root node of the rule tree and performing corresponding rule matching at each node along the matching path; obtaining events according to the matching results of the nodes.

[0009] In some embodiments, the method for generating the rule tree file includes: traversing each rule, parsing the content of the rule to obtain rule information, and determining the node relationship based on the dependency relationship between rules; the rule information includes: rule attributes and methods, and the rule attributes include the regularized text for matching; constructing a first rule tree according to the rule information and dependency relationship of each rule; wherein, each node stores the rule information of a single rule; equalizing the dependency relationship of each node in the first rule tree to obtain a second rule tree; forming a rule tree file according to the second rule tree.

[0010] In some embodiments, before determining the corresponding node relationships in the rule tree based on the dependencies between the rules, rule preprocessing is further included; the rule preprocessing includes at least one of the following: 1) in response to a downstream rule depending on a first number of upstream rules, creating a downstream rule with the same content as the downstream rule so that the number of downstream rules reaches the first number; 2) splitting the regular text included in a rule into multiple sub-regular texts, and forming a rule according to each sub-regular text; 3) splitting the regular text included in a downstream rule into multiple sub-regular texts, and forming a rule according to each sub-regular text to form a rule set corresponding to the downstream rule; in response to the downstream rule depending on a first number of upstream rules, creating a rule set with the same content as the rule set so that the rule set reaches the first number.

[0011] In some embodiments, equalizing the dependencies of each node in the first rule tree includes: generating a corresponding text feature matrix based on the regularized text of each node; determining a preset number of target feature dimensions with the largest amount of information based on the text feature matrix; determining each text segment corresponding in the regularized text according to the preset number of target feature dimensions; constructing an intermediate layer between the current layer and the upper layer according to each text segment; wherein, the intermediate layer includes: a preset number of first nodes, each first node is respectively constructed corresponding to one of the text segments; and a second node; based on the matching relationship between the regularized text of each current node in the current layer and each text segment, determining the dependency relationships between each current node and the first node and the second node respectively; wherein, the current node that obtains a match forms a dependency relationship with the first node, and the current node that does not obtain a match forms a dependency relationship with the second node.

[0012] In some embodiments, the text segment is a word; generating a corresponding text feature matrix based on the regularized text of each node includes: mapping the regularized text corresponding to each node into a text feature vector through a bag-of-words model, and stacking to form a text feature matrix.

[0013] In some embodiments, each node has a node type, and the node type is related to the requirements of the corresponding rule.

[0014] In some embodiments, the node type includes: a specific type and a general type; the specific type includes: a frequency-related type, a keyword grouping-related type, and a frequency and keyword grouping-related type.

[0015] In some embodiments, the node also has a method of the corresponding rule, and the method of the rule is used to be called to obtain an event according to the matching result.

[0016] In some embodiments, each of the nodes caches the matching results generated by itself.

[0017] A second aspect of the present application provides a log-based event acquisition device, including: a rule tree construction module, configured to obtain a rule tree file and construct a rule tree based on the rule tree file; wherein, each node of the rule tree corresponds to a rule construction; the node relationships between the nodes correspond to the dependency relationships between the rules; a rule tree matching module, configured to input log data into the root node of the rule tree and perform matching of corresponding rules at each node along the matching path; an event acquisition module, configured to acquire events according to the matching results of the nodes.

[0018] A third aspect of the present application provides a computer device, including: a communicator, a memory, and a processor; the communicator is configured to communicate with the outside; the memory is configured to store program instructions; the processor is configured to run the program instructions to execute the log-based event acquisition method according to any one of the first aspect.

[0019] A fourth aspect of the present application provides a computer-readable storage medium, storing program instructions, and the program instructions, when running, execute the log-based event acquisition method according to any one of the first aspect.

[0020] As described above, the embodiments of the present application provide a log-based event acquisition method, device, equipment, and medium. By obtaining a rule tree file and constructing a rule tree based on the rule tree file; wherein, each node of the rule tree corresponds to a rule construction; the node relationships between the nodes correspond to the dependency relationships between the rules; inputting log data into the root node of the rule tree and performing matching of corresponding rules at each node along the matching path; acquiring events according to the matching results of the nodes. By establishing a rule-related rule tree, log data can pass through each node along the matching path. Since the node relationships correspond to the dependency relationships between the rules, the efficiency is effectively improved compared with the previous method of traversing the rule set. In addition, when constructing the rule tree, the rule tree structure can be optimized by means of node balancing and regular text splitting to form multiple nodes, which is also more conducive to improving the matching efficiency; the node type can also be set according to the rule requirements to call the corresponding method during matching to meet the rule requirements. Description of the Drawings

[0021] Figure 1 A flowchart showing the log-based event acquisition method in an embodiment of the present application.

[0022] Figure 2 A flowchart showing the generation method of the rule tree file in an embodiment of the present application.

[0023] Figure 3 A diagram showing the class diagram and inheritance relationship of each node type in an embodiment of the present application.

[0024] Figure 4a Show a schematic structural diagram of a rule tree in an embodiment of the present application.

[0025] Figure 4b Show a schematic structural diagram of the rule tree after equalization in an embodiment of the present application.

[0026] Figure 5 Show a schematic flowchart of an event acquisition method in an application example of the present application.

[0027] Figure 6 Show a schematic diagram of the schematic code of 4 rules in an application example of the present application.

[0028] Figure 7 Show a schematic diagram of the regularized text of multiple rules in an application example of the present application.

[0029] Figure 8 Show a schematic diagram of the modules of a log-based event acquisition device in an embodiment of the present application.

[0030] Figure 9 Show a schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners

[0031] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the information disclosed in the present application. The present application can also be implemented or applied through other different specific implementation manners. Various details in the present application can also be modified or changed according to different viewpoints and application systems without departing from the spirit of the present application. It should be noted that, without conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0032] The following takes the accompanying drawings as a reference and details the embodiments of the present application so that those skilled in the technical field to which the present application belongs can easily implement it. The present application can be embodied in many different forms and is not limited to the embodiments described herein.

[0033] In the description of the present application, the reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics represented in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics represented can be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine the different embodiments or examples represented in the present application and the features of different embodiments or examples.

[0034] In addition, the terms "first" and "second" are used only to indicate an object and shall not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the representation of this application, "a plurality of" means two or more, unless otherwise specifically defined.

[0035] To clearly illustrate this application, devices irrelevant to the description are omitted, and the same or similar constituent elements throughout the specification are given the same reference signs.

[0036] Throughout the specification, when it is said that a device is "connected" to another device, this includes not only the case of "direct connection", but also the case of "indirect connection" with other elements placed therebetween. In addition, when it is said that a certain device "includes" a certain constituent element, unless there is a particularly contrary record, it does not exclude other constituent elements, but means that other constituent elements may also be included.

[0037] Although in some examples the terms first, second, etc. are used herein to denote various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, the first interface and the second interface, etc. are indicated. Furthermore, as used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms, unless the context indicates otherwise. It should be further understood that the terms "comprise", "include" indicate the presence of the described features, steps, operations, elements, modules, items, kinds, and / or groups, but do not exclude the presence, occurrence or addition of one or more other features, steps, operations, elements, modules, items, kinds, and / or groups. The terms "or" and "and / or" used herein are interpreted as inclusive, or meaning any one or any combination. Thus, "A, B or C" or "A, B and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B and C". The exception to this definition occurs only when the combination of elements, functions, steps or operations is inherently mutually exclusive in some way.

[0038] The technical terms used herein are only for referring to specific embodiments and are not intended to limit this application. The singular forms used herein also include the plural forms as long as the statement does not explicitly indicate the contrary meaning. The meaning of "include" used in the specification is to embody specific characteristics, regions, integers, steps, operations, elements and / or components, and does not exclude the existence or addition of other characteristics, regions, integers, steps, operations, elements and / or components.

[0039] Although not defined differently, including technical terms and scientific terms used herein, all terms have the same meaning as generally understood by those skilled in the technical field to which this application belongs. Terms defined in commonly used dictionaries are additionally interpreted to have meanings consistent with relevant technical literature and the currently presented messages. As long as they are not defined, they shall not be over-interpreted as ideal or overly formulaic meanings.

[0040] In the related art, the method of traversing regular expressions through each log is generally adopted, and the efficiency is extremely low.

[0041] In view of this, in the embodiments of the present application, a method for obtaining events based on logs can be provided, which performs matching through a rule tree generated according to corresponding rules, effectively improving the matching efficiency.

[0042] As Figure 1 shown, a schematic flowchart of a method for obtaining events based on logs in an embodiment of the present application is shown. The method for obtaining events based on logs includes:

[0043] Step S101: Obtain a rule tree file and construct a rule tree based on the rule tree file.

[0044] Among them, each node of the rule tree corresponds to the construction of a rule, and the node relationship between the nodes corresponds to the dependency relationship between the rules.

[0045] In some embodiments, the rule tree file can be pre-established. The rule tree file is established according to the rule tree, and the rule tree is constructed according to the rules. The dependency relationship between the rules is related to the matching path of the log data in each rule.

[0046] Exemplarily, when the rule tree is used for the first time, it can be generated, compiled and then run corresponding to the rule tree. After use, it is serialized and stored as a rule tree file. Among them, compilation is the process of program language translation; serialization is the process of converting the state information of an object into a form that can be stored or transmitted. After that, as long as the rules do not change, the content of the generated rule tree file can be directly loaded to restore the rule tree, which is beneficial to improving the matching efficiency.

[0047] As Figure 2 shown, a flowchart of a method for generating the rule tree file in an embodiment of the present application is shown. Optionally, Figure 2 the process in can occur when the rule tree is used for the first time to perform matching. The process includes:

[0048] Step S201: Traverse each rule, parse the content of the rule to obtain rule information, and determine the node relationship based on the dependency relationship between the rules.

[0049] In some embodiments, there may be a preset rule set, such as a rule set for network security, and each rule may be from the rule set.

[0050] In some embodiments, the rule information includes: rule attributes and methods. The rule attributes include regularized text for matching. In some embodiments, the rule attributes may include, for example: rule id, regularized text (i.e., regular expression string), upstream rule id (i.e., the id of the upstream rule in the matching path), etc. If it is a rule with statistical information frequency requirements, the rule attributes may further include frequency, time window, etc. After parsing the rule attributes of each rule, they can be organized into a list. For example, each row is for each rule: Rule 1, Rule 2...; the fields in each column are: rule id, regularized text, upstream rule id.... etc.

[0051] In some embodiments, some node relationships can be constructed through preprocessing of the rules, including at least one of the following:

[0052] 1) In response to a downstream rule depending on a first number of upstream rules, create a downstream rule with the same content as the downstream rule so that the number of downstream rules reaches the first number. For example, if a downstream rule has a upstream rules, then repeat the corresponding downstream rule to generate a sub-nodes for each of the a upstream nodes corresponding to the a upstream rules, and each sub-node contains an upstream rule id.

[0053] 2) Split the regular text contained in a rule into multiple sub-regular texts, and form a rule according to each sub-regular text. For example, a rule contains regular text "A|B", where | represents "or". Then process the regular text, remove the redundant symbols, etc., and take out A and B to form nodes separately.

[0054] Another way is to combine 1) and 2) as 3):

[0055] 3) Split the regular text contained in a downstream rule into multiple sub-regular texts, and form a rule according to each sub-regular text to form a rule set corresponding to the downstream rule; in response to the downstream rule depending on a first number of upstream rules, create a rule set with the same content as the rule set so that the rule set reaches the first number.

[0056] For example, if the regular text of a downstream rule is split into b pieces and it has a upstream rules, then the downstream rule can form a*b rule subsets with the same id, form a total of a*b sub-nodes for the a upstream nodes, and b sub-nodes under each upstream node.

[0057] Step S202: Construct a first rule tree according to the rule information and dependency relationship of each rule.

[0058] Among them, each node correspondingly stores the rule information of one rule.

[0059] In some embodiments, each node has a node type, and the node type is related to the requirements of the corresponding rule. For example, requirements for statistical information frequency, rules for classification according to keywords, etc.

[0060] In some embodiments, a rule tree object can be created, the rule list can be traversed, the rule information can be retrieved, and a node can be established for each rule.

[0061] In some embodiments, the node types include: specific types and general types (Base); the specific types include: types related to frequency (Frequency), types related to keyword (Key) grouping, and types related to frequency and keyword (KeyFrequency) grouping, etc. The class diagram and inheritance relationship are as follows Figure 3 shown. Figure 3 In each box in the figure, the class name is at the top, the parameters of the class are separated by a dotted line in the middle, and the methods of the class are below the dotted line.

[0062] It should be specifically noted that new node types can also be expanded according to actual needs, not limited to the 4 examples.

[0063] For rules with frequency requirements, a FrequencyNode instance is created; for a Key that needs to extract a certain substring as a result group, the position of the Key in the regular text needs to be specified, and a KeyNode instance is created; for rules with both of the above two characteristics, a KeyFrequencyNode instance is created; for general rules, a BaseNode instance is created. Optionally, since a rule may be split into multiple nodes due to its regularized text, and to facilitate identifying these nodes and knowing their sources, these nodes can have the same node id (node_id), but different indexes.

[0064] According to the upstream rule information of the downstream rules, the nodes in the rule tree are connected to complete the node relationship modeling of the rule dependency relationship. The structure of the rule tree is as follows Figure 4a shown, each node corresponds to a rule and stores the corresponding rule attributes. Optionally, the node also has the method of the corresponding rule, and the method of the rule is used to be called to obtain an event according to the matching result. For example, the Frequency method of FrequencyNode is used to be called to count the frequency of information appearance according to the matching result, etc.

[0065] Step S203: Equalize the dependency relationships of the nodes in the first rule tree to obtain a second rule tree.

[0066] In some embodiments, equalizing the dependency relationships of the nodes in the first rule tree includes: generating a corresponding text feature matrix based on the regularized text of each node; determining a preset number of target feature dimensions with the largest information content based on the text feature matrix; determining each text segment corresponding to the preset number of target feature dimensions in the regularized text; constructing an intermediate layer between the current layer and the upper layer according to each text segment; wherein, the intermediate layer includes: a preset number of first nodes, each first node is respectively constructed corresponding to a text segment; and a second node; based on the matching relationship between the regularized text of each current node in the current layer and each text segment, determining the dependency relationships between each current node and the first node and the second node respectively; wherein, the current node that obtains a match forms a dependency relationship with the first node, and the current node that does not obtain a match forms a dependency relationship with the second node.

[0067] In some embodiments, the text segment can be a word, or can also be a phrase or a longer text, etc.; generating a corresponding text feature matrix based on the regularized text of each node includes: mapping the regularized text corresponding to each node into a text feature vector through a bag-of-words model, and stacking to form a text feature matrix.

[0068] The following gives a specific example of equalization. If a certain layer in the rule tree has a large number of nodes, such as Figure 4a the second layer of the rule tree in, it may lead to a relatively large number of nodes that need to be passed through in the matching path during matching. Exemplarily, a natural language processing method can be used to analyze the text segments (words) in the regular text, extract the nodes of an intermediate layer, and then break up the large number of nodes and construct dependency relationships with the intermediate layer nodes.

[0069] In a possible implementation example, after preprocessing all the regularized texts of the rules (which can be splitting the regularized texts), a bag-of-words model can be used to represent the regularized text of each node (non-word characters can be removed). Thus, the regularized text of each rule can be quantified into a d-dimensional text feature vector, and the vectorization algorithm can adopt word frequency or TF-IDF, etc. Each feature dimension corresponds to a text segment, which is a word in this example. Suppose there are n regularized texts in total,

[0070] In a possible example, when forming the nodes of the intermediate layer, in order to disperse these nodes as much as possible, according to the principle of maximum entropy, the information content vector of each feature dimension can be calculated according to formula (1):

[0071]

[0072] wherein, H represents the vector of a preset number of target feature dimensions, They are the eigenvalues of the text feature matrix on the feature dimension (i.e., the corresponding vocabulary) of the d-th column, and p(x) is the distribution of the eigenvalues of this feature dimension x.

[0073] The greater the amount of information, the more balanced the allocation of the lower-layer nodes that depend on the reconstruction of the middle-layer nodes can be. As shown in Equation (2), the m dimensions with the largest amount of information are selected, and the corresponding vocabulary is used as the middle-layer nodes. The lower-layer nodes are linked to the first node in the middle layer associated with them (i.e., containing the vocabulary corresponding to the target feature dimension). To handle the case where some regular texts only contain non-word strings, a second node is added to the middle layer, and the nodes that do not contain the above m vocabulary are linked to the default node.

[0074]

[0075] Among them is in descending order.

[0076] Reference can be made to Figures 4a to 4b the changes of Figure 4b where each node in the middle layer is added and the dependency relationships of the middle-layer nodes are reallocated.

[0077] Step S204: Generate a rule tree file according to the second rule tree.

[0078] In some embodiments, the rule tree can be serialized and stored in the rule tree file in binary form for subsequent reading.

[0079] Return to Figure 2 in, step S102: Input the log data into the root node of the rule tree and perform corresponding rule matching at each node along the matching path;

[0080] In some embodiments, each of the nodes caches the matching results generated by itself.

[0081] In a specific example, step S102 can be implemented as reading log data from, for example, a distributed log storage system and inputting it item by item into the root node of the rule tree. If the regular text of the current node matches the log, the index of the log is recorded in the node cache instead of directly storing the log content, thus reducing the occupation of storage resources. Optionally, the node type may also affect the change of the matching path, so the setting of the node type can be changed according to actual needs. For example, if the node type is KeyNode, the key value is extracted according to the position specified by the key for grouping, the index of the log is recorded in the key value list of the corresponding group, and then the log is passed to the child node of the current node for continued matching; if the current node does not match the log, the matching task of the current matching path can be terminated, and subsequent nodes do not need to perform the matching operation on the log.

[0082] Step S103: Obtain an event according to the matching result of the node.

[0083] In some embodiments, the cache results of each node of the rule tree are read to analyze and obtain an event.

[0084] Nodes of different node types obtain events in different ways.

[0085] For example, for a node of the FrequencyNode type, the number of logs with matching hits can be counted, and the difference between the maximum time and the minimum time within the window can be calculated in a sliding window manner to determine whether a group of log indexes within the time range meets the quantity requirement. If such a group of results can be found, an event that meets the frequency is output; otherwise, no event is output. Another example is that for other types of nodes, according to specific business requirements, the node information and the cached log indexes can be integrated, and an event can be output for each log. Further optionally, the output event can be persisted to the storage end.

[0086] After the rule tree is used for the first time, the compiled rule tree can also be serialized and stored in a rule tree file in binary form for subsequent reading.

[0087] Reference can be made to Figure 5 , and an application example is given to illustrate the principle of the above event acquisition method. Taking a certain security log matching system as an example.

[0088] In Figure 5 , the specific process may include:

[0089] Step S501: Determine whether the rule tree file is used for the first time. If so, go to step S502; if not, go to step S509: Read the rule tree file.

[0090] Step S502: Rule content parsing and complex regular expression splitting.

[0091] In some embodiments, the content of each rule in the rule definition file can be parsed and the complex regular text can be split.

[0092] For example, the attributes read from the rule definition file include id, regular expression string, upstream rule id, frequency, time window, rule description, alarm level, whether to alarm, etc.

[0093] In Figure 6 , (a) to (d) respectively schematically show the schematic codes of 4 rules. (a) is the 5710 rule, (b) is the 5719 rule, (c) is the 5716 rule, and (d) is the 5720 rule, each with different requirements.

[0094] Among them, for the rule shown in (a), the outermost or symbol in the regular expression is separated into two parts: "illegal user" and "invalid user". If the upstream rule is 5710, then this rule can occupy 2 entries in the rule list, corresponding to the nodes to be generated for "illegal user" and "invalid user" respectively.

[0095] Step S503: Construct a rule tree according to the dependency relationship of the rules:

[0096] In a specific example, a rule tree object can be created, the rule list can be traversed, the rule information can be retrieved, and a node can be established for each rule. Four types of nodes are defined in the present invention, and the class diagram and inheritance relationship are as shown in the appendix Figure 3 shown. Figure 6 The rule in (a) is instantiated into two BaseNodes; Figure 6 The rule in (b) is instantiated into one FrequencyNode; Figure 6 The rule in (c) is instantiated into one KeyNode; Figure 6 The rule in (d) is instantiated into one KeyFrequencyNode. Although Figure 6 the rule in (d) does not define the position of the key, it can inherit from the upstream node 5716. The regularized text of the rule is pre-compiled for the pattern to improve the efficiency during matching.

[0097] According to the upstream rule information of the rule, the nodes in the rule tree are connected into a tree structure as shown in Figure 4a shown.

[0098] Step S504: Equalize the number of nodes: After that, it can also enter step S508: Serialize the rule tree and store it in the rule tree file for use by step S509 next time.

[0099] Figure 4aThere are a large number of nodes in the second layer of the rule tree. To further reduce the number of nodes to be traversed during matching, natural language processing is used to analyze the regular expressions, extract an intermediate layer, break up the numerous nodes and reconstruct the dependency relationships, as Figure 4b shown.

[0100] Describe the acquisition of the nodes in the intermediate layer:

[0101] All regularized texts of the rules can be preprocessed to obtain Figure 7 the regularized texts of each rule (corresponding to nodes) in the examples: The bag-of-words model is used to represent the regularized text of each node, and non-word characters can be removed from the feature dimension. Each text can be quantified as a 169-dimensional term frequency vector. To disperse these nodes as much as possible, according to the principle of maximum entropy, if there are 64 rules, the text matrix is 64 * 169. Calculate the information content of each dimension according to Equation (1), and take the 8 feature dimensions with the largest information content as the intermediate layer nodes according to Equation (2). Link the nodes to the associated intermediate layer nodes.

[0102] The words corresponding to these 8 feature dimensions are: "error", "not", "failed", "user", "fatal", "Connection", "to", "for". Move these 8 words to the intermediate layer to form the first node; and add a second node.

[0103] Step S505: Use the rule tree to perform log and rule matching:

[0104] Log data can be read from the distributed log storage system and input into the root node of the rule tree one by one.

[0105] As shown in the following table, 4 logs 1 to 4 are provided exemplarily:

[0106] 1 Invalid user abc from 192.168.2.1 2021-07-02 16:41:55 2 Failed password for invalid user abc from 192.168.2.1 port 26605 ssh2 2021-07-02 16:41:55 3 Failed password for invalid user abc from 192.168.2.1 port 26605 ssh2 2021-07-02 16:41:55 4 Failed password for invalid user abc from 192.168.2.1 port 26605 ssh2 2021-07-02 16:41:56 … … …

[0107] Step S506: Obtain the matching result and cache it in the rule tree nodes;

[0108] Assume that the lower-level nodes corresponding to rule 5700 are 5701 to 5709. After log 1 passes through 5700, it is matched with the child nodes of 5700 in sequence. Since none of the child nodes 5701 - 5709 match, it does not match the next lower level. When reaching 5710, log 1 matches the regular text, so the index of log 1 is recorded in the cache of node 5710. Assume that logs 2 to 4 are matched at node 5716, so they are all recorded in the cache of node 5716. Since node 5716 is defined as the key type, the third value, that is, the ip value, is obtained from the group where the regular text match hits as the key of the result. The indexes of these 3 logs are saved in the hash table with the key "192.168.2.1", and then these 3 logs are also recorded in the hash table of 5720 with the key "192.168.2.1".

[0109] Step S507: Organize the content and format of the output data according to specific business requirements.

[0110] Read the results of each node of the rule tree. For the node of 5710, output one event; for the node of 5716, output three events; for the node of 5720, it is statistically found that there are three matches within 120 seconds, so output one event.

[0111] The output events can be stored in, for example, the output event storage database.

[0112] As Figure 8 shown, a module schematic diagram of the log-based event acquisition device in an embodiment of the present application is presented. The implementation of the event acquisition device can refer to the event acquisition method in the previous embodiment, so the technical features will not be repeated in this embodiment.

[0113] The event acquisition device 800 includes:

[0114] A rule tree construction module 801, configured to obtain a rule tree file and construct a rule tree based on the rule tree file; wherein, each node of the rule tree corresponds to a rule construction; the node relationship between each node corresponds to the dependency relationship between the rules;

[0115] A rule tree matching module 802, configured to input log data into the root node of the rule tree and perform corresponding rule matching on each node along the matching path;

[0116] An event acquisition module 803, configured to obtain events according to the matching results of the nodes.

[0117] In some embodiments, the method for generating the rule tree file includes: traversing each rule, parsing the content of the rule to obtain rule information, and determining node relationships based on the dependency relationships between the rules; the rule information includes: rule attributes and methods, and the rule attributes include regularized text for matching; constructing a first rule tree according to the rule information and dependency relationships of each rule; wherein each node stores the rule information of one rule correspondingly; equalizing the dependency relationships of the nodes in the first rule tree to obtain a second rule tree; and forming a rule tree file according to the second rule tree.

[0118] In some embodiments, before determining the corresponding node relationships in the rule tree based on the dependency relationships between the rules, rule preprocessing is further included; the rule preprocessing includes at least one of the following: 1) in response to a downstream rule depending on a first number of upstream rules, creating a downstream rule with the same content as the downstream rule so that the number of downstream rules reaches the first number; 2) splitting the regular text included in a rule into multiple sub-regular texts, and forming a rule according to each sub-regular text; 3) splitting the regular text included in a downstream rule into multiple sub-regular texts, and forming a rule according to each sub-regular text to form a rule set corresponding to the downstream rule; in response to the downstream rule depending on a first number of upstream rules, creating a rule set with the same content as the rule set so that the rule set reaches the first number.

[0119] In some embodiments, equalizing the dependency relationships of the nodes in the first rule tree includes: generating a corresponding text feature matrix based on the regularized text of each node; determining a preset number of target feature dimensions with the largest amount of information based on the text feature matrix; determining each text segment in the regularized text according to the preset number of target feature dimensions; constructing an intermediate layer between the current layer and the upper layer according to each text segment; wherein the intermediate layer includes: a preset number of first nodes, each first node corresponding to one of the text segments constructed; and a second node; based on the matching relationship between the regularized text of each current node in the current layer and each text segment, determining the dependency relationships between each current node and the first node and the second node respectively; wherein the current node that obtains a match forms a dependency relationship with the first node, and the current node that does not obtain a match forms a dependency relationship with the second node.

[0120] In some embodiments, the text segment is a word; generating a corresponding text feature matrix based on the regularized text of each node includes: mapping the regularized text corresponding to each node into a text feature vector through a bag-of-words model and stacking them to form a text feature matrix.

[0121] In some embodiments, each node has a node type, and the node type is related to the requirements of the corresponding rule.

[0122] In some embodiments, the node types include: specific types and general types; the specific types include: frequency-related types, keyword grouping-related types, and frequency and keyword grouping-related types.

[0123] In some embodiments, the node also has a corresponding rule method, which is used to be called to obtain an event according to the matching result.

[0124] In some embodiments, each of the nodes caches the matching results generated by itself.

[0125] It should be specifically noted that in Figure 8 Each functional module in the embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a program instruction product. The program instruction product includes one or more program instructions. When the program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The program instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.

[0126] And, Figure 8 The devices disclosed in the embodiments can be implemented by other module division methods. The device embodiments shown above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or modules can be combined or can be dynamically moved to another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical or other forms.

[0127] In addition, Figure 8 Each functional module and sub-module in the embodiments can be dynamically located in a processing component, or each module can exist physically alone, or two or more modules can be dynamically located in a component. The above-mentioned dynamic component can be implemented in the form of hardware or in the form of a software function module. When the above-mentioned dynamic component is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0128] It should be particularly noted that the flowcharts of the above embodiments of the present application represent processes or methods that can be understood as representing modules, segments, or portions of code including one or more executable instructions for implementing specific logical functions or processes. And the scope of the preferred embodiments of the present application includes additional implementations, where functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed.

[0129] For example, Figure 1 , Figure 2 , Figure 5 etc. The order of each step in the embodiments may be changed in a specific scenario and is not limited to the above representation.

[0130] As Figure 9 shown, a schematic diagram of the circuit structure of a network device in an embodiment of the present application is shown.

[0131] In some embodiments, the computer device 900 may be implemented in a server, a server group, etc.

[0132] The computer device 900 includes a bus 901, a processor 902, a memory 903, and a communicator 904. The processor 902 and the memory 903 can communicate through the bus 901. Program instructions (such as system or application software) can be stored in the memory 903. The processor 902 implements the steps of the log-based event acquisition method in the embodiments of the present application by running the program instructions in the memory 903.

[0133] The bus 901 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, although Figure 1 only a thick line is used to represent it in the figure, it does not mean that there is only one bus or one type of bus.

[0134] In some embodiments, the processor 902 may be implemented as a Central Processing Unit (CPU), a Microcontroller Unit (MCU), a System On Chip, or a Field Programmable Gate Array (FPGA), etc. The memory 903 may include volatile memory for temporarily storing data when running programs, such as Random Access Memory (RAM).

[0135] The memory 903 may also include non-volatile memory for data storage, such as Read-Only Memory (ROM), flash memory, Hard Disk Drive (HDD), or Solid-State Disk (SSD).

[0136] The communicator 904 is used for external communication. In a specific example, the communicator 904 may include one or more wired and / or wireless communication circuit modules. For example, the wired communication circuit module may include one or more of, for example, a wired network card, a USB module, a serial interface module, etc. Again, for example, the wireless communication protocols followed by the wireless communication module include: for example, Near Field Communication (NFC) technology, Infrared (IR) technology, Global System for Mobile communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Bluetooth (BT), Global Navigation Satellite System (GNSS), etc.

[0137] In an embodiment of the present application, a computer-readable storage medium may also be provided, storing program instructions that, when run, execute the steps in the previous embodiments.

[0138] That is, the method steps in the above embodiments are implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and will be stored in a local recording medium, so that the method represented herein can be stored on such a recording medium and processed by software using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA).

[0139] In summary, the embodiments of the present application provide a log-based event acquisition method, device, equipment, and medium. By obtaining a rule tree file and constructing a rule tree based on the rule tree file, each node of the rule tree corresponds to a rule construction, and the node relationship between nodes corresponds to the dependency relationship between rules. The log data is input into the root node of the rule tree and the corresponding rules are matched at each node along the matching path. Events are obtained according to the matching results of the nodes. By establishing a rule-related rule tree, the log data can pass through each node along the matching path. Since the node relationship corresponds to the dependency relationship between rules, the efficiency is effectively improved compared with the previous method of traversing the rule set. In addition, when constructing the rule tree, the rule tree structure can be optimized by methods such as node balancing and regular text splitting to form multiple nodes, which is also more conducive to improving the matching efficiency. The node type can also be set according to the rule requirements to call the corresponding method during matching to meet the rule requirements.

[0140] Let's further analyze the technical effects that can be achieved by the solution in the embodiments of the present application:

[0141] 1) Regular matching is performed through the upper and lower layer relationships of the rule tree nodes corresponding to the upstream and downstream relationships of the rules, reducing unnecessary node matching and avoiding traversing all rules;

[0142] 2) The rule tree is compiled once and can be saved as a file for multiple uses, and multiple related files can be read and compiled together;

[0143] 3) Support complex functions such as regular matching requirements for key-type grouped statistics;

[0144] 4) In a rule, complex regular text is split into different rules, which are transformed into multiple simple regular texts, which is beneficial to improving the execution efficiency of regular matching;

[0145] 5) The node type can be extended, and the rule information stored in the node can be flexibly increased or decreased, facilitating the support of different requirement scenarios;

[0146] 6) Node balancing processing can be implemented based on natural language processing to further reduce the number of nodes in the matching path.

[0147] The above embodiments only illustrate the principle and its effects of the present application by way of example, and are not used to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present application should still be covered by the claims of the present application.

Claims

1. A log-based event acquisition method, characterized in that, Including: Obtain a pre-generated rule tree file, and construct a rule tree based on the rule tree file; wherein, each node of the rule tree corresponds to a rule construction; the node relationship between each node corresponds to the dependency relationship between rules, and the rule tree file is generated by the following method: traverse each rule, parse the content of the rule to obtain rule information, and adjust the dependency relationship between rules through preprocessing, and the preprocessing includes at least one or more of the following methods: create a rule copy with the same content as the downstream rule to match the number of upstream rules, or split the regular text of the rule into multiple sub-regular texts to form a new rule, or split the regular text of the downstream rule into multiple sub-regular texts to form a rule set, and create multiple rule sets with the same content as the rule set to match the number of upstream rules; generate the rule tree file based on the rule information and the adjusted dependency relationship; Input the log data into the root node of the rule tree and perform corresponding rule matching at each node along the matching path; Obtain an event according to the matching result of the node.

2. The method for obtaining events based on logs according to claim 1, wherein The method for generating the rule tree file includes: Traverse each rule, parse the content of the rule to obtain rule information, and determine the node relationship based on the dependency relationship between rules; the rule information includes: rule attributes and methods, and the rule attributes include regularized text for matching; Construct a first rule tree according to the rule information and dependency relationship of each rule; wherein, each node corresponds to storing the rule information of a rule; Balance the dependency relationship of each node in the first rule tree to obtain a second rule tree; Form a rule tree file according to the second rule tree.

3. The method for obtaining events based on logs according to claim 2, wherein The balancing of the dependency relationship of each node in the first rule tree includes: Generate a corresponding text feature matrix based on the regularized text of each node; Determine a preset number of target feature dimensions with the largest amount of information based on the text feature matrix; Determine each text segment corresponding in the regularized text according to the preset number of target feature dimensions; Construct an intermediate layer between the current layer and the upper layer according to each text segment; wherein, the intermediate layer includes: a preset number of first nodes, each first node is respectively constructed corresponding to a text segment; and a second node; Based on the matching relationship between the regularized text of each current node in the current layer and each text segment, determine the dependency relationship between each current node and the first node and the second node respectively; wherein, the current node that obtains a match forms a dependency relationship with the first node, and the current node that does not obtain a match forms a dependency relationship with the second node.

4. The method for obtaining events based on logs according to claim 3, wherein The text segment is a vocabulary; the generating a corresponding text feature matrix based on the regularized text of each node includes: Map the regularized text corresponding to each node into a text feature vector through a bag-of-words model and stack them to form a text feature matrix.

5. The method for obtaining events based on logs according to claim 1, characterized in that Each node has a node type, and the node type is related to the requirements of the corresponding rule.

6. The method for obtaining events based on logs according to claim 5, wherein The node types include: specific types and general types; the specific types include at least one of: frequency-related types, keyword grouping-related types, and frequency and keyword grouping-related types.

7. The log-based event acquisition method according to claim 5 or 6, wherein The node also has a corresponding rule method, which is used to be called to obtain an event according to the matching result.

8. The method for obtaining events based on logs according to claim 1, wherein Each of the nodes caches the matching results generated by itself.

9. A log-based event acquisition device, characterized in that Including: A rule tree construction module, configured to obtain a pre-generated rule tree file and construct a rule tree based on the rule tree file; wherein, each node of the rule tree corresponds to the construction of a rule; the node relationships among the nodes correspond to the dependency relationships among the rules, and the rule tree file is generated by the following method: traversing each rule, parsing the content of the rule to obtain rule information, and preprocessing to adjust the dependency relationships among the rules, the preprocessing includes at least one or more of the following methods: creating a rule copy with the same content as the downstream rule to match the number of upstream rules, or splitting the regular text of the rule into multiple sub-regular texts to form new rules, or splitting the regular text of the downstream rule into multiple sub-regular texts to form a rule set, and creating multiple rule sets with the same content as the rule set to match the number of upstream rules; generating the rule tree file based on the rule information and the adjusted dependency relationships; A rule tree matching module, configured to input log data into the root node of the rule tree and perform corresponding rule matching at each node along the matching path; An event obtaining module, configured to obtain an event according to the matching result of the node.

10. A computer device, characterized in that, Including: A communicator, a memory, and a processor; The communicator is used for external communication; The memory is used for storing program instructions; the processor is used for running the program instructions to execute the log-based event obtaining method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, Stored with program instructions, which are executed when the program instructions are run to execute the log-based event obtaining method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for recognizing state of communication device, and communication system and storage medium

    WO2021129024A1