Log analysis method and device, electronic equipment, medium and computer program product

Through automated log grouping and parsing logic matching, the problems of long development cycle and low flexibility caused by manual participation in the existing technology are solved, and efficient and flexible log analysis is achieved.

CN119940338APending Publication Date: 2025-05-06SANGFOR TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998430.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology requires manual participation in the log parsing process, resulting in a long development cycle and low flexibility, making it difficult to quickly adapt to log structure changes and new log access.

Method used

By grouping multiple logs based on log features and obtaining the parsing logic corresponding to each log group, we realize automatic parsing of logs and reduce manual intervention.

Benefits of technology

Improves the efficiency and flexibility of log parsing, and can quickly adapt to log structure changes and new log access without re-formulating the parsing logic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940338A_ABST
    Figure CN119940338A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a log analysis method and device, electronic equipment, a medium and a computer program product, and the log analysis method comprises the steps: grouping a plurality of logs based on the characteristics of the logs, and obtaining a plurality of log groups; wherein the logs in each log group in the plurality of log groups have the same characteristics; obtaining an analysis logic corresponding to each log group; and analyzing the logs in each log group based on the analysis logic to obtain an analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data analysis, and in particular relates to a log parsing method, device, electronic device, medium and computer program product. Background Art

[0002] Currently, in the log parsing process, developers are required to analyze the log structure and set the log parsing logic based on the log structure. If new logs need to be parsed, or if existing logs need to be parsed and the parsing logic needs to be adjusted, the manual parsing process has a long development cycle and low flexibility in log parsing. Summary of the invention

[0003] Embodiments of the present application provide a log parsing method, device, electronic device, medium, and computer program product.

[0004] The present invention provides a log parsing method, which includes:

[0005] Based on the features of the logs, multiple logs are grouped to obtain multiple log groups; wherein the logs in each log group in the multiple log groups have the same features;

[0006] Obtain the parsing logic corresponding to each log group;

[0007] The logs in each log group are parsed based on the parsing logic to obtain parsing results.

[0008] In some embodiments, the logs in each log group are parsed based on the parsing logic to obtain a parsing result, including: extracting the first key field of any first log in the first log group based on the parsing logic corresponding to any first log group in each log group; determining the similarity between the first key field and preset extraction information, and when the similarity is greater than a first similarity threshold, using the first key field as the parsing result of the first log.

[0009] In some embodiments, the method further includes: obtaining a standard result based on a standard parsing engine of the first log; comparing the parsing result with the standard result to obtain a comparison difference; and modifying the preset extraction information when the comparison difference is greater than a difference threshold.

[0010] It can be seen that by comparing with the standard results, when the comparison difference is greater than the difference threshold, modifying the preset extraction information helps to obtain accurate log parsing results.

[0011] In some embodiments, the method also includes: obtaining a standard result based on a standard parsing engine of the first log; comparing the parsing result with the standard result to obtain a comparison difference; when the comparison difference is greater than a difference threshold, searching for multiple second key fields in the first log; selecting a target key field from the multiple second key fields so that the comparison difference between the target key field and the standard result is less than a difference threshold; using the target key field as the parsing result of the first log; wherein the similarity between each second key field and the preset extraction information is greater than a second similarity threshold, and the second similarity threshold is less than the first similarity threshold.

[0012] It can be seen that by comparing with the standard results, when the comparison difference is greater than the difference threshold, redetermining the parsing results in multiple key fields helps to improve the accuracy of log parsing.

[0013] In some embodiments, the method further includes: obtaining a preset scenario logic; the scenario logic is used to parse a log containing a preset field; and when the log contains the preset field, parsing the log based on the scenario logic.

[0014] It can be seen that log parsing through preset scenario logic helps to extract information in specific formats and preset fields, which is conducive to obtaining accurate log parsing results.

[0015] In some embodiments, the characteristics of the log include one or more of the format of the log, the structure of the log, and the delimiter of the log.

[0016] In some embodiments, the log-based features group multiple logs to obtain multiple log groups, including: grouping multiple logs based on the format of each log and / or the structure of the log to obtain multiple first groups; grouping the logs in the first group based on the separator of each log in the first group to obtain multiple log groups.

[0017] It can be seen that the method provided in this embodiment can automatically divide multiple logs, which helps to automatically parse the logs by matching the corresponding parsing logic through the divided log groups, thereby improving the log parsing efficiency.

[0018] The present application also provides a log parsing device, the device comprising:

[0019] A processing module, configured to group a plurality of logs based on the features of the logs to obtain a plurality of log groups; wherein the logs in each log group of the plurality of log groups have the same features;

[0020] An acquisition module, used to acquire the parsing logic corresponding to each log group;

[0021] The parsing module is used to parse the logs in each log group based on the parsing logic to obtain parsing results.

[0022] An embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory for storing a computer program that can be run on the processor; wherein:

[0023] The processor is used to run the computer program to execute any one of the above-mentioned log parsing methods.

[0024] An embodiment of the present application provides a computer storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned log parsing methods is implemented.

[0025] An embodiment of the present application provides a computer program product, including a computer program, which implements any of the above-mentioned log parsing methods when executed by a processor.

[0026] The embodiments of the present application provide a log parsing method, device, electronic device, medium and computer program product, which can improve the log parsing efficiency by grouping logs and parsing logs with the same characteristics through corresponding parsing logic, without the need for developers to analyze the log structure or modify the parsing logic. When the log to be parsed changes, the corresponding log grouping is determined by the characteristics of the updated log, and the corresponding parsing logic is determined based on the log grouping, without the need to reformulate the parsing logic, thereby improving the flexibility of log parsing. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 A flow chart of a log parsing method provided in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of the log structure provided for an embodiment of the present application;

[0029] Figure 3 A schematic diagram of a method for parsing a preset field provided in an embodiment of the present application;

[0030] Figure 4 A schematic diagram of key logic of log parsing provided in an embodiment of the present application;

[0031] Figure 5 A flow chart of a log grouping method provided in an embodiment of the present application;

[0032] Figure 6 A comparison diagram of the analysis structure stacking process and analysis algorithm logic provided in the embodiment of the present application;

[0033] Figure 7 A log parsing flow chart provided for an embodiment of the present application;

[0034] Figure 8 A schematic diagram of the structure of a log parsing device provided in an embodiment of the present application;

[0035] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The Security Information and Event Management (SIEM) security log governance method is a comprehensive security strategy that aims to detect and respond to network threats by collecting, analyzing and managing security log data. The primary task of the SIEM system is to collect various security log data in the enterprise network. These logs may come from different devices and systems. At present, there are two main methods for SIEM security log governance: one is to use SIEM equipment manufacturers to build in rules and release patch packages in advance: in the SIEM equipment development stage, the manufacturer surveys various security equipment output log samples, and the developer analyzes the log structure and builds the parsing logic into the SIEM device; after the SIEM device is released, if the customer has new security logs that need to be accessed, or the existing access logs change and need to adjust the parsing logic, it is necessary to initiate a customized demand to the SIEM manufacturer, and the manufacturer develops a patch package to upgrade the SIEM device to meet the user's needs. The other is to achieve customization through the secondary development platform provided by the SIEM equipment manufacturer: SIEM manufacturers provide standard access methods and a set of low-threshold custom development environments; after the customer accesses the log, he writes the parsing rules according to the log structure. After the log changes, the customer modifies the parsing rules to complete the adaptation.

[0037] The above methods all have the following problems during implementation: when encountering deep customized analysis needs, the customization cost is high and the customization cycle is long; there is a lack of flexibility, and each time the log structure changes, human participation is required to modify the parsing rules synchronously; the participation of personnel with development experience is required to complete the log parsing task.

[0038] In view of the above problems, the embodiment of the present application provides a log parsing method, which can solve the disadvantages of requiring manual participation in development when parsing security log text and the problem of long development cycle.

[0039] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings and examples. It should be understood that the embodiments provided herein are only used to explain the embodiments of the present application and are not intended to limit the embodiments of the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than providing all embodiments for implementing the present application. In the absence of conflict, the technical solutions recorded in the embodiments of the present application can be implemented in any combination.

[0040] It should be noted that, in the embodiments of the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a method or device including a series of elements includes not only the elements explicitly recorded, but also includes other elements not explicitly listed, or also includes elements inherent to the implementation of the method or device. In the absence of further restrictions, an element defined by the sentence "includes..." does not exclude the presence of other related elements (such as steps in the method or units in the device, such as a unit in the device may be a part of a circuit, a part of a processor, a part of a program or software, etc.) in the method or device including the element.

[0041] The log parsing method provided in the embodiment of the present application includes a series of steps, but the log parsing method provided in the embodiment of the present application is not limited to the recorded steps. Similarly, the log parsing device provided in the embodiment of the present application includes a series of modules, but the device provided in the embodiment of the present application is not limited to including the modules explicitly recorded, and can also include modules that need to be set up to obtain relevant information or perform processing based on the information.

[0042] The present application embodiment provides a log parsing method, such as Figure 1 As shown, Figure 1 A flow chart of a log parsing method is shown. Figure 1 The log parsing methods shown include:

[0043] Step 101: Based on the features of the logs, a plurality of logs are grouped to obtain a plurality of log groups; wherein the logs in each of the plurality of log groups have the same features.

[0044] After receiving the input batch of original log texts to be parsed, based on the method given in this step, the batch logs are first grouped based on the characteristics of the logs. The characteristics of the logs can be determined according to different parsing logics, and the logs that can use the same parsing logic are grouped into one group to obtain multiple log groups.

[0045] Each of the multiple log groups contains one or more logs. The logs in each log group can be generated by the same code or system component, and the logs in each log group contain similar information formats and constant parts. Specifically, after receiving a batch of logs to be parsed, the logs can be grouped according to the log generation code, system architecture, log specifications and other information to obtain multiple log groups.

[0046] Step 102: Obtain the parsing logic corresponding to each log group.

[0047] Based on the log groups obtained in step 101, the logs in each log group have the same characteristics, that is, the logs in each log group can be parsed based on the same parsing logic. Based on the characteristics of the logs in each log group, the parsing logic corresponding to each log group is obtained, and the parsing logic can be used to parse each log in a log group.

[0048] Here, the parsing logic corresponding to each log group may be pre-set. For multiple logs in a log group, the same features of the multiple logs may include the same log format, the same content structure, the same keywords, etc. Therefore, based on the same features of each log in a log group, a parsing logic can be set for parsing.

[0049] For logs with the same fixed format, you can uniformly design regular expressions or preset keywords to extract key information from each log in a group; you can use the same constant parts of each log in a log group to analyze the structure of the log and distinguish between constants and variables in the log; or you can preset parsing logic to extract timestamp information from each log in a log group, etc.

[0050] By using the clear log features in the log grouping obtained in the above steps, we can unify the design of parsing logic and implement log parsing, which helps to improve the log parsing efficiency.

[0051] Step 103: Parse the logs in each log group based on the parsing logic to obtain parsing results.

[0052] After obtaining the preset parsing logic corresponding to each log group, each log in the log group can be parsed based on the corresponding parsing logic. In actual applications, a parser can be applied to parse the logs in each log group based on the corresponding parsing logic.

[0053] In actual applications, when multiple logs need to be parsed, based on the method provided in this embodiment, the features of the logs contained in the batch of logs to be parsed can be first obtained, and the parsing logic corresponding to each feature can be obtained based on the features of the logs, and the parsing logic corresponding to each feature can be combined to obtain a parsing logic set, and the multiple logs to be parsed can be parsed through the parsing logic set. Specifically, the parsing of the logs to be parsed can be achieved by the logs hitting the corresponding parsing logic in the parsing logic set. When there is a parsing logic that cannot be hit in the log to be parsed, that is, the preset parsing logic cannot parse the log, it is considered that the log parsing has failed, and relevant alarm information is output.

[0054] The embodiment of the present application provides a log parsing method, which can realize automatic grouping of logs according to the characteristics of the logs. By presetting the parsing logic based on the grouping of the logs, and directly obtaining the corresponding parsing logic based on the characteristics of the grouped logs, the logs can be automatically parsed without manual intervention, thereby improving the efficiency of log parsing. When it is necessary to update the log to be parsed, it is only necessary to determine the characteristics of the updated log to be parsed, and the parsing logic corresponding to the updated log to be parsed can be directly obtained, so as to realize the parsing of the updated log to be parsed, without the need to manually redetermine the characteristics of the log, and without the need to reconfigure the parsing logic, thereby improving the flexibility of log parsing.

[0055] In practical applications, steps 101 to 103 can be implemented based on a processor, and the processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.

[0056] In some embodiments, the above-mentioned parsing of the logs in each log group is performed based on the parsing logic to obtain the parsing result, including: extracting the first key field of any first log in the first log group based on the parsing logic corresponding to any first log group in each log group; determining the similarity between the first key field and the preset extraction information, and when the similarity is greater than the first similarity threshold, using the first key field as the parsing result of the first log.

[0057] After obtaining the log grouping, since the valid information in the log is usually concentrated in the main part of the log, before parsing the log, the main part with the payload can also be extracted based on the log header. Specifically, if the log is stored in plain text and the log body has obvious identifiers in the text (such as specific keywords, delimiters or formats), text parsing technology can be used to extract the main body of the log, for example, regular expressions can be used to match and extract the log body. If the log is stored in a structured format such as JSON (JavaScript Object Notation), XML (eXtensible Markup Language), etc., then the corresponding parsing library can be used to parse the log and directly access the fields of the log body. For example, for logs in JSON format, a JSON parsing library (such as Python's JSON module) can be used to parse the log string into a dictionary or object, and then directly access the fields containing the log body.

[0058] For example, when processing a JSON log, such as Figure 2 As shown, the location, format, and characteristics of the payload (i.e., the information to be extracted) in JSON can also be determined by analyzing the structure of the JSON log. Extraction rules are formulated based on the characteristics of the payload. These rules should clearly specify how to traverse the JSON structure, how to identify the payload, and how to extract the value of the payload. The extraction rules can be determined based on the path, field name, data type, etc. of the JSON. The extraction rules are abstracted into extraction patterns so that they can be reused in different situations. The extraction pattern can be a regular expression, an XPath expression, a custom parsing function, or a parsing template, etc.

[0059] After obtaining the payload of the log, for some logs in specific formats, such as some logs that may be in binary format or specific data structures, such logs are first deserialized before parsing, and a data structure object that can be operated by the program is obtained based on the payload. For example, for JSON logs, an open source library can be used to process and obtain a key-value pair mapping object.

[0060] In addition, for some key-value (KV) storage formats, you can first split the string containing the key-value data into a list by line, then traverse each line, use the equal sign as a delimiter to extract the key and value, and store them in a dictionary. For the numeric value flat format, you can use the recognized delimiter to get an array with a payload.

[0061] When performing log parsing on any first log, the first log to be parsed is processed based on the above method, and after obtaining the payload and key value information, the first key field of the first log is extracted. Taking the first key field as an address-related key field as an example, since the log may contain multiple similar first key fields, for example, a log may simultaneously include the client Internet Protocol (IP) address, server IP address, source IP address (Source IP), target IP address (Destination IP), etc. When the first log contains multiple different IP addresses, the first key field related to the IP address can be screened out in the first log based on the preset parsing logic. When the target key field to be parsed in the first log is the source IP address, the source IP address can be used as the preset extraction information, and the first key field and the preset extraction information are calculated for similarity to determine that the source IP address in the first log is the parsing result of the first log.

[0062] For example, assuming that the preset extraction information is set to the source IP, since the settings of the key fields for the source IP may be different in different logs, in actual applications, for the analysis result corresponding to the extraction of the source IP, the preset extraction information may be "Source IP", "srcIp", "S_IP", etc. When the first key field obtained by the analysis logic is "Destination IP", the similarity calculation is performed with "Source IP", "srcIp", and "S_IP" respectively through "Destination IP", and when any similarity is greater than the first similarity threshold, "Destination IP" is used as the analysis result of the first log, otherwise it is not used as the analysis result; when another first key field obtained by the analysis logic is "Source IP", the similarity calculation is performed with "Source IP", "srcIp", and "S_IP" respectively through "Source IP", and when any similarity is greater than the first similarity threshold, "Source IP" is used as the analysis result of the first log.

[0063] In actual applications, the key fields of the first log obtained by the parsing logic may also include the log name, timestamp, data sent by the client to the server, etc., which are not specifically limited in this embodiment. After obtaining multiple key fields, the similarity is calculated with each key field respectively through the preset extraction information, and when the similarity is greater than the first similarity threshold, the corresponding parsing result is output. For example, for JSON, KV mode, compare and check whether the semantics of the key name conforms to the preset extraction information, and extract the key name that is considered to be the source IP. This key name needs to be a key field similar to "srcIp".

[0064] It can be seen that the method provided in this embodiment can automatically obtain the log parsing result through the parsing logic and preset extraction information, thereby improving the log parsing efficiency.

[0065] In some embodiments, the above method further includes: obtaining a standard result based on a standard parsing engine of the first log; comparing the parsing result with the standard result to obtain a comparison difference; and modifying the preset extraction information when the comparison difference is greater than a difference threshold.

[0066] In some embodiments, the above method also includes: obtaining a standard result based on a standard parsing engine of the first log; comparing the parsing result with the standard result to obtain a comparison difference; when the comparison difference is greater than a difference threshold, searching for multiple second key fields in the first log; selecting a target key field from the multiple second key fields so that the comparison difference between the target key field and the standard result is less than the difference threshold; using the target key field as the parsing result of the first log; wherein the similarity between each second key field and the preset extracted information is greater than a second similarity threshold, and the second similarity threshold is less than the first similarity threshold.

[0067] After obtaining the parsing results of the logs based on the above steps, you can use the playback verification service to pass the parsing rules and the original logs to be parsed into the verification engine. The verification engine here is the corresponding standard parsing engine. Compared with the parsing engine used to parse logs, the verification engine is an independent and lightweight deployment. It is used to simulate the real parsing environment in small batches multiple times to give the actual parsing results.

[0068] When the comparison difference is greater than the difference threshold, it is considered that the parsing result obtained based on the parsing method given in the above embodiment is significantly different from the expected parsing result. In this case, the reasons for the inaccurate parsing result generally include the following two factors: the first factor is that the preset extraction information is incorrect, that is, the expected parsing result cannot be effectively identified based on the preset extraction information; the second factor is that the correct first key field is not extracted.

[0069] For the first factor, the preset extraction information needs to be modified, and the similarity is recalculated with the first key field through the modified preset extraction information until the difference between the determined parsing result and the standard result is less than the difference threshold.

[0070] For the second factor, it is necessary to re-extract the second key field in the first log, that is, to lower the first similarity threshold while ensuring that the preset extraction information remains unchanged, and obtain the second similarity threshold. By lowering the similarity threshold, it is determined whether there is a target key field that is the same or similar to the standard result among the multiple second key fields with lower similarity to the preset extraction information, and the target key field is used as the parsing result.

[0071] In order to improve the accuracy of the subsequent log analysis, in the process of determining the target key field, different weights can be assigned to the first key field and each second key field. The similarity between each key field and the preset extracted information is obtained, and the similarity between each key field and the preset extracted information is multiplied by the weight of the key field to obtain the first weighted similarity. The key field with the largest first weighted similarity is selected as the target key field to determine the analysis result. By continuously adjusting the weight of the first key field and the weight of each second key field, the difference between the target key field and the standard result is made less than the difference threshold. Based on this method, in the subsequent log analysis, when analyzing this type of key field, the analysis result can be determined directly by multiplying the determined weight with the similarity between the key field and the preset extracted information.

[0072] It can be seen that the method provided in the above embodiment helps to obtain a more accurate analysis result by comparing the analysis result with the standard result and correcting the analysis logic.

[0073] In some embodiments, the above method also includes: obtaining preset scenario logic; the scenario logic is used to parse the log containing preset fields; when the log contains the preset fields, the log is parsed based on the scenario logic.

[0074] Since the formats of logs or the ways in which logs are generated vary greatly, for special logs, such as logs containing preset fields, scenario logic for parsing preset fields is set, so that in the special case where the log contains preset fields, the log is parsed based on the scenario logic. In practical applications, based on the method given in the above implementation, after parsing the log, it can be determined whether the log contains preset fields based on the scenario logic given in this embodiment. If the log contains preset fields, the log is parsed based on the scenario logic; if the log does not contain preset fields, the log parsing is completed.

[0075] Figure 3 A method for parsing a preset field is shown. It can be seen that when the source IP of the log needs to be parsed, Figure 3 The preset field "host" shown also contains the source IP information. Therefore, after the preset field "host" is detected based on the scenario logic, the scenario logic can be used to parse the source IP information contained in the field "host" in the log as "3.3.3.3".

[0076] In actual applications, if multiple logs are parsed at the same time, multiple parsing results may be obtained. When the parsing result of a certain log is needed, a field set can be extracted according to the obtained standard document, and the content of the extracted field set can be trimmed to exclude the original fields that do not meet the extraction target.

[0077] Based on the log parsing method given in the above embodiment, Figure 4 A schematic diagram of key logic of log parsing is shown, including:

[0078] Step 401: Obtain third-party logs.

[0079] Obtain multiple logs provided by a third party.

[0080] Step 402: Identify the characteristics of the log.

[0081] Based on the characteristics of the logs, the logs are grouped according to their characteristics. When the characteristics of the logs cannot be identified, a prompt is issued.

[0082] Step 403: Locate the log payload location.

[0083] Locate the payload position of the log, that is, extract the body of the log, and extract the payload through the body of the log.

[0084] Step 404: Extract the segmentation method within the payload.

[0085] The segmentation method in the log is extracted from the payload, that is, the delimiter of the log is obtained, and the log is further grouped according to the log segmentation method.

[0086] Step 405: Extract the key-value pairs contained in the log.

[0087] Extract the key-value pairs of each log in each log group.

[0088] Step 406: Map key-value pairs to standard fields.

[0089] The key-value pair is processed based on the parsing logic, and the key name in the key-value pair that satisfies the parsing logic is used as the first key field extracted. The parsing result of the log is determined based on the preset extraction information and the first key field. In a specific implementation, the first key field in the log can be determined by performing a similarity calculation based on the log field standard document and the information in the log.

[0090] Step 407: Count the field contents and extract special scenario processing logic.

[0091] The scenario logic provided in the above embodiment is used to parse the preset fields in the log.

[0092] Step 408: Identify irrelevant fields and clean them up.

[0093] Based on the processing flow given in the above steps, step 403 can be used as the pre-extraction logic of the parsing logic, steps 404 to 405 can be used as the deserialization logic in the parsing logic, step 406 can be used as the field mapping logic in the parsing logic, step 407 can be used as the post-processing logic in the parsing logic, and step 408 can be used as the cache clearing logic. In actual applications, different logics can be combined and used in combination according to the specific characteristics of the log to obtain the parsing results of the log.

[0094] In some embodiments, the characteristics of the log include one or more of the format of the log, the structure of the log, and the delimiter of the log.

[0095] Specifically, the log format includes but is not limited to JSON format, XML format, log database format, text format, etc., and the log structure includes but is not limited to KV format, hierarchical structure, text line structure, etc. Different logs can be defined to use different separators, such as blank lines, specific characters such as "=", "\", "|", etc.

[0096] In some embodiments, the above-mentioned log-based features group multiple logs to obtain multiple log groups, including: grouping multiple logs based on the format of each log and / or the structure of the log to obtain multiple first groups; grouping the logs in the first group based on the separator of each log in the first group to obtain multiple log groups.

[0097] Based on the characteristics of the logs, after obtaining a batch of logs to be parsed, firstly, the multiple logs are grouped for the first time according to the format of the logs and / or the results of the logs to obtain multiple first groups. After obtaining multiple first groups, for any first group, the logs in the first group are grouped for the second time according to the delimiter of the logs, and finally multiple log groups are obtained. Based on the method provided in this embodiment, each log in each log group finally obtained has the same log format, the same log structure and the same log delimiter.

[0098] Based on the log grouping method provided in this embodiment, Figure 5 A flow chart of a log grouping method is shown, including:

[0099] Step 501: Obtain three-party logs.

[0100] Obtain multiple logs provided by a third party.

[0101] Step 502: Process a single log.

[0102] Select a single log from multiple logs for processing.

[0103] Step 503: Determine whether the JSON feature is matched.

[0104] First, it is determined whether the JSON feature is hit, that is, whether the log is in JSON format. If the log hits the JSON feature, step 504 is executed, otherwise step 507 is executed.

[0105] Step 504: Get the JSON schema grouping.

[0106] Step 505: Extract log features.

[0107] When a log matches the JSON feature, extract the delimiter feature of the log.

[0108] Step 506: Perform statistical analysis to distinguish log features.

[0109] The characteristics of the delimiters used by each log in the JSON pattern grouping are statistically analyzed, and based on the delimiter characteristics, the logs in the JSON pattern grouping are further grouped, and step 516 is executed.

[0110] Step 507: Determine whether the key-value feature is matched.

[0111] When the log does not match the JSON feature, it is further determined whether the log matches the key-value feature, that is, whether the content in the log is in the key-value format. If the log matches the key-value feature, step 508 is executed; otherwise, step 511 is executed.

[0112] Step 508: Get KV mode grouping.

[0113] Step 509: Extract log features.

[0114] When a log matches the key-value feature, the delimiter feature of the log is extracted.

[0115] Step 510: Statistical analysis to distinguish log features.

[0116] The characteristics of the delimiter used by each log in the KV mode grouping are statistically analyzed, and based on the delimiter characteristics, the logs in the KV mode grouping are further grouped, and step 516 is executed.

[0117] Step 511: Determine whether the value tile feature is hit.

[0118] When the log does not match the JSON feature and the key-value feature, it is further determined whether the log matches the value flat feature, that is, whether the content in the log is in the form of value flat. If the log matches the value flat feature, step 512 is executed; otherwise, step 515 is executed.

[0119] Step 512: Get the value split pattern grouping.

[0120] Step 513: Extract log features.

[0121] When a log hits the value tile feature, the delimiter feature of the log is extracted.

[0122] Step 514: Statistical analysis to distinguish log features.

[0123] The characteristics of the delimiter used by each log in the value split mode grouping are statistically analyzed, and based on the delimiter characteristics, the logs in the value split mode grouping are further grouped, and step 516 is executed.

[0124] Step 515: Determine whether the log has been processed.

[0125] After the processing of the current multiple logs is completed, step 517 is executed; otherwise, the process returns to step 502 to continue grouping the remaining logs.

[0126] Step 516: Obtain grouping results.

[0127] It can be seen that through the division method in the above steps, logs can be divided in detail based on different log modes.

[0128] Step 517: Obtain identification grouping results.

[0129] Step 518: Determine whether there is a grouping result.

[0130] When there is a log grouping, that is, when the grouping result is obtained in step 516 , step 519 is executed; otherwise, step 521 is executed.

[0131] Step 519: Locate the log payload location.

[0132] Based on the method provided in the above embodiment, the main part of the log is located.

[0133] Step 520: Execute the pre-extraction logic.

[0134] After locating the log body, extract the log payload based on the log body.

[0135] Step 521: Identification failed.

[0136] Based on the log grouping method provided in this embodiment, it is possible to identify the features of a single log for multiple logs in turn, determine whether the payload in a single log text is JSON, a key-value combination, or a value flat format, and at the same time, for logs of the same format, different log headers, key-value delimiters, value flat delimiters and other features may be included. At this time, the log features are further extracted and grouped, and corresponding patterns are defined to give log groupings under different patterns. If all logs are traversed but there is no successful pattern recognition result, it means that the log cannot be classified by the preset log features. In this case, the features of the log can be increased, and log grouping can be further implemented based on the features of the newly added log. In actual applications, with Figure 5 According to the group identification method shown, when a log does not match any of the JSON features, key-value features, or value features, it is considered that the log does not belong to the standard security log and the log parsing task fails.

[0137] In combination with the log parsing method given in the above embodiment, Figure 6 The corresponding analysis structure stacking process and analysis algorithm logic comparison diagram are shown, as shown in Figure 6 The analysis structure stacking process is consistent with the method given in the above embodiment. Based on the analysis result stacking, the analysis algorithm logic can be obtained. In the corresponding algorithm logic, according to the different log modes extracted, that is, the different log groups in the above embodiment, it is determined whether the mode 1 is hit. If it is hit, the log is parsed based on the analysis algorithm of the current hit mode 1. Otherwise, it continues to determine whether the mode 2 is hit. If it is hit, the log is parsed based on the analysis algorithm of the current hit mode 2. Otherwise, it continues to determine whether the mode 3 is hit.

[0138] based on Figure 6 In the example shown, the groups of the identified logs are arranged and combined, and the parsing logic syntax is stacked according to the combination results. Here, the parsing logic syntax can be flexibly matched according to different parsing engines. Figure 6 The process shown on the left arranges logical objects and generates corresponding parsing rules according to the specified engine syntax. When the parsing task starts, the parsing engine is used to load the parsing rules and connect to the third-party security log source. Figure 6 The algorithm logic shown on the right processes the log to complete the log parsing.

[0139] In practical applications, based on Figure 6The parsing algorithm logic shown in the figure can skip the process of rendering the parsing logic after obtaining the specific parsing logic, and directly load the parsing logic object into the parsing engine to take effect. When the parsing task is started, the traditional parsing engine compiles the parsing object of the runtime state according to the loaded rule syntax, and the parsing engine needs to be redeveloped at this time; in specific implementation, the parsing logic combination process given in this embodiment can be embedded in the parsing engine, and the access task can be divided into a learning stage and a production stage. The parsing logic object is obtained in the learning stage, and after the object is successfully obtained, it automatically enters the production stage and starts to execute the formal log parsing task.

[0140] After determining the parsing logic, use the configuration delivery pipeline that comes with the parsing environment to deliver the parsing logic that has been verified through playback to the production environment engine, so that it can take effect in the specified engine as expected by the user.

[0141] In combination with the log parsing method given in the above embodiment, Figure 7 A log parsing flow chart is shown, including:

[0142] Step 701: Batch logs are input.

[0143] Step 702: Perform pattern recognition and extract parsing logic.

[0144] Based on the log field standard document, the log pattern is analyzed to extract the log features. If the extraction is successful, the log is parsed. If the extraction fails, the log parsing is considered to have failed.

[0145] Step 703: group the logs.

[0146] Based on the characteristics of the logs, the logs are grouped to obtain log groups. For specific implementation methods, please refer to Figure 5 The processing method given.

[0147] Step 704: Parse the rule code.

[0148] That is, the corresponding parsing logic is obtained according to the characteristics of different logs, such as Figure 6 The parsing algorithm logic is shown.

[0149] Step 705: The parsing engine loads rules, replays logs, and confirms that the parsing is as expected.

[0150] After obtaining the analysis result based on the analysis algorithm, the analysis result is compared with the standard result. When the comparison difference is greater than the difference threshold, the preset extraction information is modified, or the analysis result is re-determined based on the method given in the above embodiment to obtain a new analysis logic.

[0151] Step 706: Output the parsing result.

[0152] The log is parsed through the data processing engine based on the new parsing logic to obtain the log parsing result, indicating that the log parsing is successful.

[0153] This embodiment provides a log parsing method, which can automatically generate parsing rules. The parsing rules are flexibly adapted to various log structures, and the log parsing efficiency can be improved.

[0154] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.

[0155] Based on the log parsing method proposed in the above embodiment, the present application embodiment also provides a log parsing device, such as Figure 8 As shown, the log parsing device includes:

[0156] The processing module 801 is used to group multiple logs based on the features of the logs to obtain multiple log groups; wherein the logs in each log group in the multiple log groups have the same features.

[0157] The acquisition module 802 is used to acquire the parsing logic corresponding to each log group.

[0158] The parsing module 803 is used to parse the logs in each log group based on the parsing logic to obtain parsing results.

[0159] In practical applications, the processing module 801, the acquisition module 802, and the parsing module 803 can be implemented based on a processor and a communication device.

[0160] In some embodiments, the parsing module 803 is specifically used to extract the first key field of any first log in the first log group based on the parsing logic corresponding to any first log group in each log group; determine the similarity between the first key field and the preset extraction information, and when the similarity is greater than the first similarity threshold, use the first key field as the parsing result of the first log.

[0161] In some embodiments, the parsing module 803 is further used to obtain a standard result based on a standard parsing engine of the first log; compare the parsing result with the standard result to obtain a comparison difference; and modify the preset extraction information when the comparison difference is greater than a difference threshold.

[0162] In some embodiments, the parsing module 803 is also used to obtain a standard result based on a standard parsing engine of the first log; compare the parsing result with the standard result to obtain a comparison difference; when the comparison difference is greater than a difference threshold, search for multiple second key fields in the first log; select a target key field from the multiple second key fields so that the comparison difference between the target key field and the standard result is less than the difference threshold; use the target key field as the parsing result of the first log; wherein the similarity between each second key field and the preset extracted information is greater than the second similarity threshold, and the second similarity threshold is less than the first similarity threshold.

[0163] In some embodiments, the acquisition module 802 is also used to acquire preset scenario logic; the scenario logic is used to parse logs containing preset fields; the parsing module 803 is also used to parse logs based on the scenario logic when the logs contain preset fields.

[0164] In some embodiments, the processing module 801 is specifically used to group multiple logs based on the format of each log and / or the structure of the log to obtain multiple first groups; based on the separator of each log in the first group, group the logs in the first group to obtain multiple log groups.

[0165] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the same method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0166] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a computer device (which can be a terminal, a server, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0167] An embodiment of the present application also provides an electronic device. Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Fig. 9 As shown, the electronic device 90 may include:

[0168] The memory 901 is used to store executable instructions.

[0169] The processor 902 is used to implement any one of the above-mentioned log parsing methods when executing the executable instructions stored in the memory 901.

[0170] The processor 902 may be at least one of an ASIC, a DSP, a DSPD, a PLD, a FPGA, a CPU, a controller, a microcontroller, and a microprocessor.

[0171] The above-mentioned computer-readable storage medium or memory 901 can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and other memories; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0172] An embodiment of the present application further provides a computer storage medium, on which computer executable instructions are stored, and the computer executable instructions are used to implement any one of the log parsing methods provided in the above embodiments.

[0173] Correspondingly, an embodiment of the present application further provides a computer program product, which includes computer executable instructions, and the computer executable instructions are used to implement any one of the log parsing methods provided in the above embodiments.

[0174] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0175] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0176] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0177] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0178] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0179] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0180] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A log parsing method, characterized in that: The method comprises: Based on the features of the logs, multiple logs are grouped to obtain multiple log groups; wherein the logs in each log group in the multiple log groups have the same features; Obtain the parsing logic corresponding to each log group; The logs in each log group are parsed based on the parsing logic to obtain parsing results.

2. The method according to claim 1, characterized in that The parsing of the logs in each log group based on the parsing logic to obtain parsing results includes: Extracting a first key field of any first log in each log group based on the parsing logic corresponding to any first log group in each log group; The similarity between the first key field and preset extraction information is determined, and when the similarity is greater than a first similarity threshold, the first key field is used as a parsing result of the first log.

3. The method according to claim 2, characterized in that The method further comprises: Obtaining a standard result based on a standard parsing engine of the first log; Comparing the analysis result with the standard result to obtain a comparison difference; When the comparison difference is greater than a difference threshold, the preset extraction information is modified.

4. The method according to claim 2, characterized in that: The method further comprises: Obtaining a standard result based on a standard parsing engine of the first log; Comparing the analysis result with the standard result to obtain a comparison difference; When the comparison difference is greater than the difference threshold, multiple second key fields are searched in the first log; a target key field is selected from the multiple second key fields so that the comparison difference between the target key field and the standard result is less than the difference threshold; the target key field is used as the parsing result of the first log; wherein the similarity between each second key field and the preset extraction information is greater than a second similarity threshold, and the second similarity threshold is less than the first similarity threshold.

5. The method according to claim 1, characterized in that The method further comprises: Obtaining preset scenario logic; the scenario logic is used to parse logs containing preset fields; When the log contains the preset field, the log is parsed based on the scenario logic.

6. The method according to claim 1, characterized in that The log features include one or more of the log format, the log structure, and the log separator.

7. The method according to claim 1, characterized in that The multiple logs are grouped based on the features of the logs to obtain multiple log groups, including: Based on the format of each log and / or the structure of the log, the multiple logs are grouped to obtain multiple first groups; The logs in the first group are grouped based on the separator of each log in the first group to obtain multiple log groups.

8. A log analysis device, characterized in that: The device comprises: A processing module, configured to group a plurality of logs based on the features of the logs to obtain a plurality of log groups; wherein the logs in each log group of the plurality of log groups have the same features; An acquisition module, used to acquire the parsing logic corresponding to each log group; The parsing module is used to parse the logs in each log group based on the parsing logic to obtain parsing results.

9. An electronic device, characterized in that: The electronic device comprises a processor and a memory for storing a computer program that can be run on the processor; wherein, The processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.

10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 7 when executed by a processor.