Log analysis rule generation method, device and equipment and computer readable storage medium

Through a large language model, the log analysis rules are automatically created, which solves the inefficiency problem of relying on experts to manually create rules in the existing technology, and realizes the automation and generalization of log analysis rules.

CN119940339APending Publication Date: 2025-05-06SANGFOR TECH INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411999033.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, experts need to have deep network security knowledge and programming skills to manually create log parsing rules, which are inefficient and difficult to adapt to logs from different manufacturers and formats.

Method used

A large language model is used to process third-party logs, determine the mapping relationship between the third-party log and the target log field name and value, and generate automated log parsing rules.

Benefits of technology

Without the need for a priori knowledge in specific fields, large language model technology is used to generate log parsing rules directly from raw log data and log specification standard documents, which improves efficiency and realizes the standardization and generalization of the log parsing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940339A_ABST
    Figure CN119940339A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a log analysis rule generation method. The method comprises the steps that a third-party log corresponding to third-party equipment is acquired; processing the third-party log based on a large language model to obtain processed log information; based on the large language model, the log information and the target log field name, determining a first mapping relation between the third-party log and the target log field name; wherein the target log field name has a first format; based on the large language model and the target log field value, determining a second mapping relation between the third-party log and the target log field value; wherein the target log field value has a first format; and based on the large language model, the log information, the first mapping relationship and the second mapping relationship, generating a log analysis rule for the third-party log. The embodiment of the invention further discloses a log analysis rule generation device and equipment and a computer readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a log parsing rule generation method, device, equipment and computer-readable storage medium. Background Art

[0002] With the rapid development of information technology and the increasing complexity of network security threats, how to effectively process and parse network and terminal logs from different manufacturers and in different formats has become a thorny issue. At present, the relevant technologies mainly rely on rule writing experts in the field of network security to manually create log parsing rules, and then parse network and terminal logs from different manufacturers and in different formats based on the manually created log parsing rules. However, the method of parsing logs through log parsing rules manually created by experts in the relevant technologies not only requires the experts to have deep network security knowledge and programming skills, but also has low efficiency in generating log parsing rules. Summary of the invention

[0003] In order to solve the above technical problems, the embodiments of the present application hope to provide a log parsing rule generation method, device, equipment and computer-readable storage medium, which can solve the problem that the related technology not only requires experts to have deep network security knowledge and programming skills, but also the efficiency of generating log parsing rules is low.

[0004] The technical solution of this application is implemented as follows:

[0005] A log parsing rule generation method, the method comprising:

[0006] Obtain third-party logs corresponding to third-party devices;

[0007] Processing the third-party log based on the large language model to obtain processed log information;

[0008] Based on the large language model, the log information and the target log field name, determining a first mapping relationship between the third-party log and the target log field name; wherein the target log field name has a first format;

[0009] Based on the large language model and the target log field value, determine a second mapping relationship between the third-party log and the target log field value; wherein the target log field value has the first format;

[0010] A log parsing rule for the third-party log is generated based on the large language model, the log information, the first mapping relationship, and the second mapping relationship.

[0011] In the above solution, the third-party log includes the first log set and / or the second log set, and the third-party log is processed based on the large language model to obtain processed log information, including:

[0012] For the first log set, based on the large language model and the prompt word corresponding to the first log in the first log set, format each of the first logs to obtain a log in a second format; wherein the log in the second format includes third-party log field information; the first log set includes multiple first logs;

[0013] For the second log set, the second log in the second log set is parsed based on the large language model to obtain third-party log field information of the second log; wherein the second log set includes multiple second logs and description information of each second log.

[0014] In the above solution, the step of formatting each of the first logs based on the large language model and the prompt words corresponding to the first logs in the first log set to obtain a log in the second format includes:

[0015] Determine a target log from the plurality of first logs;

[0016] Based on the large language model and the first target prompt words corresponding to the target logs, each of the first logs is formatted to obtain a log having the second format.

[0017] In the above solution, the formatting of each first log to obtain a log having the second format based on the large language model and the first target prompt word corresponding to the target log includes:

[0018] Determine a first regular expression of the target log based on the large language model and the first target prompt word;

[0019] For logs matching the first regular expression among the plurality of first logs, formatting the matching logs based on the first regular expression to obtain logs having the second format corresponding to the matching logs;

[0020] For logs that do not match the first regular expression among the plurality of first logs, formatting is performed on each of the unmatched logs based on the large language model and the second target prompt words corresponding to the unmatched logs to obtain logs having the second format.

[0021] In the above scheme, the formatting of each of the unmatched logs based on the large language model and the second target prompt word corresponding to the unmatched logs to obtain a log having the second format includes:

[0022] Determining a first sub-target log from the unmatched logs;

[0023] Determine a second regular expression for the first sub-goal log based on the large language model and a first sub-prompt word corresponding to the first sub-goal log;

[0024] For a first sub-log that matches the second regular expression in the unmatched log, formatting the matched first sub-log based on the second regular expression to obtain a log having the second format corresponding to the matched first sub-log;

[0025] For a second sub-log in the unmatched log that does not match the second regular expression, determine a second sub-target log from the unmatched second sub-log;

[0026] Based on the large language model and the second sub-prompt word corresponding to the second sub-target log, a third regular expression of the second sub-target log is determined until a log having the second format corresponding to each of the unmatched second sub-logs is determined.

[0027] In the above solution, the second log in the second log set is parsed based on the large language model to obtain the third-party log field information in the second log, including:

[0028] Determine a log document title corresponding to each of the second logs; wherein the log document title is used to characterize the type of the log field of the second log;

[0029] Merging the plurality of second logs based on the log document titles to obtain a target log;

[0030] For each of the target logs, when the target log is determined to be a security log based on the large language model, a third-party log field name is determined from the target log; wherein the third-party log field information includes the third-party log field name.

[0031] In the above solution, determining the first mapping relationship between the third-party log and the target log field name based on the large language model, the log information and the target log field name includes:

[0032] constructing a first prompt word based on the third-party log field name in the third-party log field information and the target field name;

[0033] Based on the first prompt word and the large language model, a first mapping relationship between the third-party log field name and the target field name is determined.

[0034] In the above solution, determining the second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value includes:

[0035] For the first log set, based on the first mapping relationship and the concerned field name in the target field name, determine the target third-party log field name from the third-party log field names;

[0036] Determine a third-party log field enumeration value corresponding to the target third-party log field name;

[0037] Constructing a second prompt word based on the third-party log field enumeration value and the target log field value;

[0038] Based on the second prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

[0039] In the above solution, determining the second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value includes:

[0040] For the second log set, obtaining target field information of the second log from the second log set; wherein the target field information includes a third-party log field enumeration value corresponding to each third-party log field name;

[0041] Constructing a third prompt word based on the third-party log field enumeration value and the target log field value;

[0042] Based on the third prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

[0043] In the above scheme, the log parsing rule generation method also includes:

[0044] Get the parsing rule verification file;

[0045] The log parsing rule is verified using the parsing rule verification file to obtain a verification result.

[0046] A log parsing rule generating device, the device comprising:

[0047] An acquisition unit, used to acquire third-party logs corresponding to third-party devices;

[0048] A parsing unit, used for processing the third-party log based on a large language model to obtain processed log information;

[0049] A mapping unit, configured to determine a first mapping relationship between the third-party log and the target log field name based on the large language model, the log information and the target log field name; wherein the target log field name has a first format;

[0050] The mapping unit is further used to determine a second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value; wherein the target log field value has the first format;

[0051] A generating unit is used to generate a log parsing rule for the third-party log based on the large language model, the log information, the first mapping relationship and the second mapping relationship.

[0052] A log parsing rule generating device, the device comprising: a processor, a memory and a communication bus;

[0053] The communication bus is used to realize the communication connection between the processor and the memory;

[0054] The processor is used to execute the log parsing rule generation program stored in the memory to implement the steps of the above-mentioned log parsing rule generation method.

[0055] A computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the above-mentioned log parsing rule generation method.

[0056] The log parsing rule generation method, apparatus, device and computer-readable storage medium provided in the embodiments of the present application first obtain a third-party log corresponding to a third-party device, and then process the third-party log based on a large language model to obtain processed log information; then determine a first mapping relationship between the third-party log and the target log field name based on the large language model, the log information and the target log field name; wherein the target log field name has a first format; and determine a second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value, and the target log field value has a first format; then generate a log parsing rule for the third-party log based on the large language model, the log information, the first mapping relationship and the second mapping relationship. In this way, in the absence of prior knowledge in a specific field, the large language model technology is used to realize an automated process of directly generating log parsing rules from raw log data and log specification standard documents, rather than relying on experts with deep network security knowledge and programming skills to manually create log parsing rules as in the related art, which not only realizes the standardization and generalization of the log parsing process, but also improves the efficiency of generating log parsing rules. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A flow chart of a log parsing rule generation method provided in an embodiment of the present application;

[0058] Figure 2 A flowchart of another log parsing rule generation method provided in an embodiment of the present application;

[0059] Figure 3 A flowchart of another log parsing rule generation method provided in an embodiment of the present application;

[0060] Figure 4 A schematic diagram of processing a second log in a log parsing rule generation method provided in an embodiment of the present application;

[0061] Figure 5 A schematic diagram of formatting processing for a first log in a log parsing rule generation method provided in an embodiment of the present application;

[0062] Figure 6 A schematic diagram of verifying a log parsing rule in a log parsing rule generation method provided in an embodiment of the present application;

[0063] Figure 7 A schematic diagram of the structure of a log parsing rule generation device provided in an embodiment of the present application;

[0064] Figure 8 A schematic diagram of the structure of a log parsing rule generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0066] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0067] It should be noted that with the rapid development of information technology and the increasing complexity of network security threats, enterprises and organizations are facing unprecedented security challenges. Extended Detection and Response (XDR) is an open and flexible security management platform that can integrate log data from various security devices and applications to provide a comprehensive perspective for security analysis and threat detection. XDR is a platform that integrates security information and event management (SIEM) log access, security detection engine, security operation module, and linkage disposal. It generates security alarms based on secondary analysis of access logs, further explores potential threats based on the security capabilities of reported security devices, and links with disposal devices to complete the closed loop of detection and disposal actions, bringing semi-automatic operation capabilities to network security in the covered environment, saving manpower while improving the quality of network security construction. It is generally believed that XDR is a mainstream product implementation form of security operation platform and security detection platform.

[0068] However, there is a key challenge here: how to effectively process and parse network and terminal logs from different vendors and in different formats. Due to the diversity and complexity of these network and terminal logs from different vendors and in different formats, data unification and standardization has become a thorny issue. These heterogeneous log data need to be converted into a unified, system-understandable standard format for subsequent security analysis and incident response. At present, the industry mainly relies on rule writing experts in the field of network security to manually create log parsing rules. These rules are usually based on the following frameworks or engines: parsing frameworks based on Java / Flink, parsing frameworks based on data collection engines (such as Logstash), and custom parsing frameworks based on other programming languages. The following disadvantages exist: 1) Inefficiency: Manually writing parsing rules is a time-consuming and tedious process, especially when faced with a large number of logs in different formats; 2) High professional threshold: Writing high-quality parsing rules requires deep network security knowledge and programming skills, which limits the number of people who can do this job; 3) Difficulty in scaling: As the scale of enterprises expands and the amount of data increases, the method of manually writing and maintaining parsing rules can hardly meet the rapidly growing needs; 4) Difficulty in adapting to changes: As new security devices and log formats continue to emerge, it becomes increasingly difficult to manually update and maintain parsing rules.

[0069] Based on this, this application solves the problems of low efficiency and difficulty in scaling when manually writing network security device log parsing rules (codes), and proposes an automated log parsing rule generation solution based on a large language model. It supports log data and log specification documents as input, and generates log parsing rules without prior knowledge.

[0070] Specifically, the embodiment of the present application provides a log parsing rule generation method, which can be applied to a log parsing rule generation device, referring to Figure 1 As shown, the method comprises the following steps:

[0071] Step 101: Obtain a third-party log corresponding to a third-party device.

[0072] In an embodiment of the present application, the third-party log includes a first log set and / or a second log set; the third-party log corresponding to the third-party device may refer to network logs and terminal logs from different manufacturers and in different formats, and the format of the third-party log is diverse and complex; the third-party log data may include a first log set, and the first log set only includes the first log (i.e., the original log data); the third-party log data may include a second log set, and the second log set includes a third-party log specification document or a third-party log standard document, and the third-party log specification document or the third-party log standard document includes multiple second logs and description information of each second log; the third-party log data may also include a first log set and a second log set.

[0073] Step 102: Process the third-party log based on the large language model to obtain processed log information.

[0074] In an embodiment of the present application, the processed log information may refer to log information obtained by processing a third-party log through a large language model, and the large language model may be used to process the first log set and / or the second log set included in the third-party log respectively, and the processing methods for the first log and the second log are different; processing the first log set based on the large language model may refer to format conversion of the first log in the first log set to obtain a log with a format different from that of the original first log; processing the second log set based on the large language model may refer to parsing the content of the second log set and extracting log field information from the second log set.

[0075] It should be noted that the Large Language Model (LLM) is a natural language processing model based on deep learning technology. It can understand and generate human language through training with massive text data, and perform various complex language tasks, such as text generation, question answering, translation, etc., and LLM usually adopts the Transformer architecture, with billions or even hundreds of billions of parameters, which can capture the deep semantics and contextual relationships of language. In a feasible implementation, LLM can be a generative pre-trained transformer (GPT) series, etc., and GPT is a natural language processing model architecture based on deep learning, which is pre-trained with large-scale text data and can understand and generate human language.

[0076] Step 103: Determine a first mapping relationship between the third-party log and the target log field name based on the large language model, the log information, and the target log field name.

[0077] The target log field name has a first format.

[0078] In an embodiment of the present application, the target log field name may refer to a preset local standard specification field name, and the first format may refer to a local standard specification format; the first mapping relationship refers to the mapping relationship between the third-party log field name and the target log field name, and the first mapping relationship may include a first mapping relationship for the first log set and / or the second log set; the log information and the target log field name may be processed first, and then processed with the large language model to obtain the mapping relationship between the third-party log field name and the target log field name; it should be noted that network and terminal logs from different manufacturers and in different formats (i.e., heterogeneous log data) are converted into a local standard specification format for subsequent security analysis and incident response.

[0079] Step 104: Determine a second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value.

[0080] The target log field value has a first format.

[0081] In an embodiment of the present application, the target log field value may refer to a preset local standard specification field value; the second mapping relationship may refer to a mapping relationship between a third-party log field value and a target log field value, and the second mapping relationship may include a second mapping relationship for the first log set and / or the second log set; the target log field value may be processed first and then processed with the large language model to obtain a mapping relationship between the third-party log field value and the target log field value.

[0082] Step 105: Generate log parsing rules for third-party logs based on the large language model, the log information, the first mapping relationship, and the second mapping relationship.

[0083] In an embodiment of the present application, the log information, the first mapping relationship and the second mapping relationship can be processed by a code generation module in a large language model, that is, the code generation module can be used to perform code conversion on the log information to obtain a first code; and the code generation module can be used to perform code conversion on the first mapping relationship to obtain a second code; and the code generation module can be used to perform code conversion on the second mapping relationship to obtain a third code; in the case where the third-party log includes a first log set, the first code, the second code and the third code for the first log set can be merged to obtain a log parsing rule (log parsing code) for the first log set; in the case where the third-party log includes a second log set, the first code, the second code and the third code for the second log set can be merged to obtain a log parsing rule for the second log set; in the case where the third-party log includes a first log set and a second log set, the log parsing rule for the first log set and the log parsing rule for the second log set can be merged to obtain a final log parsing rule.

[0084] The log parsing rule generation method provided in the embodiment of the present application uses large language model technology to realize an automated process of directly generating log parsing rules from raw log data and log specification standard documents without prior knowledge in a specific field, rather than relying on experts with deep network security knowledge and programming skills to manually create log parsing rules as in the related art. This not only realizes the standardization and generalization of the log parsing process, but also improves the efficiency of generating log parsing rules.

[0085] Based on the above embodiments, the present application embodiment provides another log parsing rule generation method, referring to Figure 2 and Figure 3 As shown, the method may include the following steps:

[0086] Step 201: The log parsing rule generating device obtains a third-party log corresponding to a third-party device.

[0087] The third-party log includes a first log set and / or a second log set.

[0088] In the embodiment of the present application, the third-party log may include only the first log set, the third-party log may include only the second log set, or the third-party log may include the first log set and the second log set. In the case where the third-party log includes the first log set and the second log set, since there are generally a large number of second logs with rich types in the second log set, and they can be extracted as the log data set, therefore, Figure 4As shown, the log document content in the second log set may be segmented and then input into the large language model to obtain a second log in the log document, and the second log may be used as a data supplement to the first log set.

[0089] It should be noted that step 202 may be performed after step 201, and step 203 may also be performed after step 201;

[0090] Step 202: For the first log set, the log parsing rule generating device formats each first log based on the large language model and the prompt words corresponding to the first logs in the first log set to obtain a log in a second format.

[0091] The log in the second format includes third-party log field information; and the first log set includes multiple first logs.

[0092] In an embodiment of the present application, the first log set includes multiple first logs, and the multiple first logs can have multiple different formats, such as key-value format, json format, natural language text format, value flat format, etc. and different log header formats; the log in the second format can specifically refer to a log in the json format of a key-value pair, and the third-party log field information can include a third-party log field name and a third-party log field value, and the third-party log field name corresponds to the key, and the third-party log field value corresponds to the value; a prompt word can be first constructed for the first log, and the prompt word can be input into the large language model to parse the first log to obtain a log in the second format.

[0093] It should be noted that, based on the large language model and the prompt word corresponding to the first log in the first log set, before parsing the first log to obtain a log in the second format, you can first select several logs from the multiple first logs in the first log set and input them into the large language model to obtain the log type of the first log, including security logs (network security logs, terminal security logs, mixed security logs, etc.) and non-security logs (protocol audit logs, operation logs, login logs, etc.).

[0094] It should be noted that step 202 can be implemented in the following ways:

[0095] Step 202A: The log parsing rule generating device determines a target log from a plurality of first logs.

[0096] Step 202B: The log parsing rule generating device formats each first log based on the large language model and the first target prompt word corresponding to the target log to obtain a log in a second format.

[0097] In the embodiment of the present application, the target log may refer to a log randomly selected from multiple first logs; the first target prompt word may refer to a prompt word constructed for the target log; a prompt word may be constructed for the target log, and based on the constructed prompt word and the large language model, each first log is formatted to obtain a log in the second format; that is, each first log can be formatted to obtain a log in the second format.

[0098] It should be noted that if Figure 5 As shown, step 202B can be implemented in the following manner:

[0099] Step 202B1: The log parsing rule generating device determines a first regular expression of the target log based on the large language model and the first target prompt word corresponding to the target log.

[0100] Step 202B2: For the logs matching the first regular expression among the multiple first logs, the log parsing rule generating device formats the matching logs based on the first regular expression to obtain logs having a second format corresponding to the matching logs.

[0101] Step 202B3: For logs in the multiple first logs that do not match the first regular expression, the log parsing rule generation device formats each unmatched log based on the large language model and the second target prompt word corresponding to the unmatched log to obtain a log with a second format.

[0102] In an embodiment of the present application, the first target prompt word may refer to the requirement information of the regular expression of the format corresponding to the target log; the second target prompt word may refer to the requirement information of the regular expression of the format corresponding to the unmatched logs in multiple first logs; the matched logs may refer to the logs that match the first regular expression in multiple first logs, and the unmatched logs may refer to the logs that do not match the first regular expression in multiple first logs; it should be noted that logs of different formats have different regular expressions.

[0103] In an embodiment of the present application, a target log can be randomly selected from multiple first logs, and then the first target prompt word corresponding to the target log is input into the large language model to obtain a first regular expression of the format of the target log, and then the first regular expression is used to match all the first logs to obtain logs that match the first regular expression and logs that do not match the first regular expression, and the matched logs are formatted to obtain logs in the second format corresponding to the matched logs (i.e., key-value json format); for the unmatched logs, each unmatched log needs to be formatted based on the large language model and the second target prompt word corresponding to the unmatched log to obtain a log in the second format corresponding to each unmatched log.

[0104] It should be noted that step 202B3 can be implemented in the following ways:

[0105] Step 202b1: The log parsing rule generating device determines a first sub-target log from unmatched logs.

[0106] Step 202b2: Determine a second regular expression for the first sub-goal log based on the large language model and the first sub-prompt word corresponding to the first sub-goal log.

[0107] Step 202b3: For the first sub-log that matches the second regular expression in the unmatched log, the log parsing rule generating device formats the matched first sub-log based on the second regular expression to obtain a log having a second format corresponding to the matched first sub-log.

[0108] Step 202b4: For the second sub-log in the unmatched log that does not match the second regular expression, the log parsing rule generating device determines a second sub-target log from the unmatched second sub-log.

[0109] Step 202b5: Determine a third regular expression for the second sub-target log based on the large language model and the second sub-prompt word corresponding to the second sub-target log, until a log having the second format corresponding to each unmatched second sub-log is determined.

[0110] In an embodiment of the present application, the first sub-target log may refer to any log in the unmatched logs; the first sub-prompt word may refer to the prompt word corresponding to the first sub-target log; for the unmatched logs, the first sub-target log may be randomly selected from the unmatched logs, and then the first sub-prompt word corresponding to the first sub-target log may be input into the large language model to obtain a second regular expression in the format of the first sub-target log, and then the second regular expression is used to match the unmatched logs to obtain a first sub-log that matches the second regular expression and a second sub-log that does not match the second regular expression, and for the first sub-log that matches the second regular expression, the first sub-log is matched based on the second regular expression. The log is formatted to obtain a log in the second format (i.e., key-value json format) corresponding to the first sub-log. For the second sub-log that matches the second regular expression, a log is randomly selected from the second sub-log, i.e., the second sub-target log. Then, a third regular expression of the format of the second sub-target log is constructed. Then, based on the third regular expression, matching is continued with the unmatched second sub-logs until a log in the second format corresponding to each unmatched second sub-log is determined, that is, until the log formats of all the first logs in the first log set are covered (i.e., until the format conversion of each first log in the first log set is completed).

[0111] Step 203: For the second log set, the log parsing rule generating device parses the second log in the second log set based on the large language model to obtain third-party log field information of the second log.

[0112] The second log set includes multiple second logs and description information of each second log.

[0113] In the embodiment of the present application, the second log in the second log set exists in the form of a log document, and the content of each log document in the second log set can be parsed, and the third-party log field information can be extracted from the log document based on the parsing result. Figure 4 As shown, before parsing the second log set, the format of the log document corresponding to the second log in the second log set may be converted, and the formats of all log documents may be converted into the same document format (such as docx format).

[0114] It should be noted that step 203 can be implemented in the following ways:

[0115] Step 203C: The log parsing rule generating device determines the log document title corresponding to each second log.

[0116] The log document title is used to characterize the type of the log field of the second log.

[0117] Step 203D: The log parsing rule generating device merges the multiple second logs based on the log document titles to obtain a target log.

[0118] Step 203E: For each target log, the log parsing rule generating device obtains the third-party log field name from the target log when determining that the target log is a security log based on the large language model.

[0119] The third-party log field information includes the third-party log field name.

[0120] In the embodiment of the present application, since there may be a large number of non-security related log fields in the log document, the log document title can be used as the basis for dividing the log document, and the contents under the same log document title are considered to belong to the same type of log field. Figure 4 As shown, the content of the log document can be parsed to determine the log document title of the log document, and then the log documents corresponding to the same log document title are spliced ​​to obtain the text information under the same log document title (i.e., the target log). Then, as shown in Figure 4 As shown, the large language model is used to determine whether the text information under the document title belongs to the security log. If it belongs to the security log, all field information under the document title (ie, the third-party log field name) is extracted from the text information under the document title.

[0121] It should be noted that security log types may include network security logs, terminal security logs, mixed security logs, etc.; non-security log types may include protocol audit logs, operation logs, login logs, etc.

[0122] It should be noted that steps 204 to 205 may be performed after steps 202 and 203;

[0123] Step 204: The log parsing rule generating device constructs a first prompt word based on the third-party log field name and the target field name in the third-party log field information.

[0124] Step 205: The log parsing rule generating device determines a first mapping relationship between the third-party log field name and the target field name based on the first prompt word and the large language model.

[0125] In an embodiment of the present application, for the first log set, after formatting each first log in the first log set, all the first logs can be parsed into a key-value json format; the target field name can refer to a local standard specification field name; the key (i.e., the third-party log field name) can be first determined from the key-value json format, and then a first prompt word can be constructed based on the third-party log field name and the target field name, and then the first prompt word can be input into the large language model to request the large language model to obtain a mapping relationship between the third-party log field name and the target field name; it should be noted that, in the case where the first mapping relationship between the third-party log field name and the target field name cannot be directly determined, the third-party log field name can also be extracted and spliced ​​before determining the mapping relationship.

[0126] In an embodiment of the present application, for the second log set, since the third-party log field name can be directly extracted after parsing the second log set, the first prompt word can be directly constructed based on the obtained third-party log field name and the local standard specification field name, and then the first prompt word is input into the large language model, and the large language model is requested to obtain the first mapping relationship between the third-party log field name and the target field name.

[0127] It should be noted that after step 205, the second mapping relationship for the first log set can be implemented through steps 206 to 209, and the second mapping relationship for the second log set can be implemented through steps 210 to 212:

[0128] Step 206: For the first log set, the log parsing rule generating device determines a target third-party log field name from the third-party log field names based on the first mapping relationship and the concerned field name in the target field name.

[0129] Step 207: The log parsing rule generating device determines a third-party log field enumeration value corresponding to the target third-party log field name.

[0130] Step 208: The log parsing rule generating device constructs a second prompt word based on the third-party log field enumeration value and the target log field value.

[0131] Step 209: The log parsing rule generating device determines a second mapping relationship between the third-party log field enumeration value and the target log field value based on the second prompt word and the large language model.

[0132] In an embodiment of the present application, for the first log, the target third-party log field name is the field name that needs to be paid attention to that is preset in the third-party log field name; some enumeration type log fields, such as attack success status, confidence, threat level, action and other log fields, need to map their log field values ​​to local standard specification field enumeration values ​​through mapping. Therefore, the target third-party log field name corresponding to the field name of interest in the target field name can be first determined based on the field name of interest in the target field name and the first mapping relationship, and then for each target third-party log field name, the third-party log field value corresponding to the target third-party log field name is merged to obtain the third-party log field enumeration value corresponding to the target third-party log field name, and then a second prompt word is constructed based on the third-party log field enumeration value and the target log field value, and then the second prompt word is input into the large language model, requesting the large language model to determine the mapping relationship between the third-party log field enumeration value and the target log field value.

[0133] Step 210: For the second log set, the log parsing rule generating device obtains target field information of the second log from the second log set.

[0134] The target field information includes a third-party log field enumeration value corresponding to each third-party log field name.

[0135] Step 211: The log parsing rule generating device constructs a third prompt word based on the third-party log field enumeration value and the target log field value.

[0136] Step 212: The log parsing rule generating device determines a second mapping relationship between the third-party log field enumeration value and the target log field value based on the third prompt word and the large language model.

[0137] In the embodiment of the present application, for the second log, the target field information of the log document may refer to a standard field table in the log document, and the standard field table includes an enumeration type of each third-party log field name and its corresponding third-party log field enumeration value, such as Figure 4 As shown, specifically, the content of each log document can be analyzed to obtain the target field information of each log document; then a third prompt word can be constructed based on the third-party log field enumeration value and the target log field value, and the third prompt word can be input into the large language model to request the large language model to obtain the mapping relationship between the third-party log field enumeration value and the target log field value.

[0138] Step 213: The log parsing rule generating device generates a log parsing rule for the third-party log based on the large language model, the log information, the first mapping relationship, and the second mapping relationship.

[0139] It should be noted that the above embodiment may further include the following steps:

[0140] Step 214: The log parsing rule generating device obtains a parsing rule verification file.

[0141] Step 215: The log parsing rule generating device uses the parsing rule verification file to verify the log parsing rule to obtain a verification result.

[0142] In the embodiment of the present application, after obtaining the log parsing rules, the log parsing rules can also be processed as follows: Figure 5 The rule automatic verification shown, that is, in the process of gradually updating the log parsing rules, it is necessary to ensure that the generated log parsing rules are syntactically correct, executable, and the output results meet the expected rules (i.e., code). Specifically, the parsing rule verification service module can automatically load the modifiable parsing rule configuration file at any time point where verification is required, provide a way to input logs and monitor output results, and achieve the purpose of verifying rule codes and batch filtering log data sets.

[0143] It should be noted that in the log formatting and mapping relationship generation process, in addition to using the prompt words mentioned above, supervised fine-tuning (SFT) and retrieval-augmented generation (RAG) can also be used. Among them, the SFT method can be implemented by supervised learning of the pre-trained model on a specific log data set, specifically including: collecting labeled data related to log formatting, building a training set, inputting logs, outputting corresponding formatting codes, and fine-tuning the model based on these data; in the log mapping relationship generation process, the RAG method first uses a retrieval module to obtain context information related to the input log from a pre-built knowledge base, and then inputs this information together with the input log into the mapping generation module, thereby generating a more accurate and context-related mapping relationship. In the specific implementation, by building a knowledge base containing common log formats and mapping relationships, the retrieval algorithm is used to quickly find relevant information to enhance the context understanding ability of the generated model. In this way, the accuracy and relevance of the generated content can be improved by combining an external knowledge base and a pre-trained model.

[0144] It should be noted that, for the description of the same steps and the same contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.

[0145] The log parsing rule generation method provided in the embodiment of the present application uses large language model technology to realize an automated process of directly generating log parsing rules from raw log data and log specification standard documents without prior knowledge in a specific field, rather than relying on experts with deep network security knowledge and programming skills to manually create log parsing rules as in the related art. This not only realizes the standardization and generalization of the log parsing process, but also improves the efficiency of generating log parsing rules.

[0146] Based on the above embodiments, the present application provides a log parsing rule generation device, which can be applied to Figure 1 and Figure 2 In the log parsing rule generation method provided in the corresponding embodiment, refer to Figure 7 As shown, the log parsing rule generating device 3 may include: an acquiring unit 31, a parsing unit 32, a mapping unit 33 and a generating unit 34, wherein:

[0147] An acquisition unit 31 is used to acquire a third-party log corresponding to a third-party device;

[0148] The parsing unit 32 is used to process the third-party log based on the large language model to obtain processed log information;

[0149] A mapping unit 33, configured to determine a first mapping relationship between a third-party log and a target log field name based on the large language model, the log information, and the target log field name; wherein the target log field name has a first format;

[0150] The mapping unit 33 is further used to determine a second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value; wherein the target log field value has a first format;

[0151] The generating unit 34 is used to generate a log parsing rule for a third-party log based on the large language model, the log information, the first mapping relationship and the second mapping relationship.

[0152] In other embodiments of the present application, the parsing unit 32 is further configured to perform the following steps:

[0153] For the first log set, based on the large language model and the prompt word corresponding to the first log in the first log set, each first log is formatted to obtain a log in a second format; wherein the log in the second format includes third-party log field information; the first log set includes multiple first logs;

[0154] For the second log set, the second log in the second log set is parsed based on the large language model to obtain third-party log field information in the second log; wherein the second log set includes multiple second logs and description information of each second log.

[0155] In other embodiments of the present application, the parsing unit 32 is further configured to perform the following steps:

[0156] determining a target log from a plurality of first logs;

[0157] Based on the large language model and the first target prompt word corresponding to the target log, each first log is formatted to obtain a log in a second format.

[0158] In other embodiments of the present application, the parsing unit 32 is further configured to perform the following steps:

[0159] Determine a first regular expression of the target log based on the large language model and the first prompt word;

[0160] For logs matching the first regular expression among the plurality of first logs, formatting the matching logs based on the first regular expression to obtain logs having a second format corresponding to the matching logs;

[0161] For logs that do not match the first regular expression among the plurality of first logs, formatting is performed on each of the unmatched logs based on the large language model and the second target prompt words corresponding to the unmatched logs to obtain logs in a second format.

[0162] In other embodiments of the present application, the parsing unit 32 is further configured to perform the following steps:

[0163] determining a first sub-target log from the unmatched logs;

[0164] Determine a second regular expression for the first sub-log based on the large language model and the first sub-prompt word corresponding to the first sub-target log;

[0165] For the first sub-log that matches the second regular expression in the unmatched log, formatting the matched first sub-log based on the second regular expression to obtain a log having a second format corresponding to the matched first sub-log;

[0166] For a second sub-log that does not match the second regular expression in the unmatched log, determine a second sub-target log from the unmatched second sub-log;

[0167] Based on the large language model and the second sub-prompt word corresponding to the second sub-target log, a third regular expression of the second sub-target log is determined until a log having a second format corresponding to each unmatched second sub-log is determined.

[0168] In other embodiments of the present application, the parsing unit 32 is further configured to perform the following steps:

[0169] Determine a log document title corresponding to each second log; wherein the log document title is used to characterize the type of the log field of the second log;

[0170] Merging multiple second log documents based on log document titles to obtain a target log;

[0171] For each target log, when it is determined based on the large language model that the target log document is a security log, a third-party log field name is obtained from the target log document; wherein the third-party log field information includes the third-party log field name.

[0172] In other embodiments of the present application, the mapping unit 33 is further configured to perform the following steps:

[0173] constructing a first prompt word based on the third-party log field name and the target field name in the third-party log field information;

[0174] Based on the first prompt word and the large language model, a first mapping relationship between the third-party log field name and the target field name is determined.

[0175] In other embodiments of the present application, the mapping unit 33 is further configured to perform the following steps:

[0176] For the first log set, based on the first mapping relationship and the concerned field name in the target field name, determine the target third-party log field name from the third-party log field names;

[0177] Determine the third-party log field enumeration value corresponding to the target third-party log field name;

[0178] Constructing a second prompt word based on the third-party log field enumeration value and the target log field value;

[0179] Based on the second prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

[0180] In other embodiments of the present application, the mapping unit 33 is further configured to perform the following steps:

[0181] For the second log set, obtaining target field information of each second log from the second log set; wherein the target field information includes a third-party log field enumeration value corresponding to each third-party log field name;

[0182] Constructing a third prompt word based on the third-party log field enumeration value and the target log field value;

[0183] Based on the third prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

[0184] In other embodiments of the present application, the mapping unit 33 is further configured to perform the following steps:

[0185] Get the parsing rule verification file;

[0186] The log parsing rules are verified using the parsing rule verification file to obtain the verification results.

[0187] It should be noted that the specific implementation process of the steps performed by each module in the embodiment of the present application can be referred to Figure 1 and Figure 2 The implementation process of the log parsing rule generation method provided in the corresponding embodiment will not be repeated here.

[0188] The log parsing rule generating device provided in the embodiment of the present application utilizes large language model technology to realize an automated process of directly generating log parsing rules from raw log data and log specification standard documents without prior knowledge in a specific field, rather than relying on experts with deep network security knowledge and programming skills to manually create log parsing rules as in the related art. This not only realizes the standardization and generalization of the log parsing process, but also improves the efficiency of generating log parsing rules.

[0189] Based on the above embodiments, the embodiments of the present application provide a log parsing rule generation device, which can be applied to Figure 1 and Figure 2 In the log parsing rule generation method provided in the corresponding embodiment, refer to Figure 8 As shown, the log parsing rule generating device 4 may include: a processor 41, a memory 42 and a communication bus 43, wherein:

[0190] The communication bus 43 is used to realize the communication connection between the processor 41 and the memory 42;

[0191] The processor 41 is used to execute the log parsing rule generation program in the memory 42 to implement the following steps:

[0192] Obtain third-party logs corresponding to third-party devices;

[0193] Process third-party logs based on the large language model to obtain processed log information;

[0194] Determine a first mapping relationship between the third-party log and the target log field name based on the large language model, the log information, and the target log field name; wherein the target log field name has a first format;

[0195] Determine a second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value; wherein the target log field value has a first format;

[0196] Based on the large language model, the log information, the first mapping relationship and the second mapping relationship, a log parsing rule for the third-party log is generated.

[0197] In other embodiments of the present application, the processor 41 is used to execute the third-party log in the log parsing rule generation program in the memory 42, including the first log set and / or the second log set, and process the third-party log based on the large language model to obtain processed log information to implement the following steps:

[0198] For the first log set, based on the large language model and the prompt word corresponding to the first log in the first log set, each first log is formatted to obtain a log in a second format; wherein the log in the second format includes third-party log field information; the first log set includes multiple first logs;

[0199] For the second log set, the second log in the second log set is parsed based on the large language model to obtain third-party log field information in the second log; wherein the second log set includes multiple second logs and description information of each second log.

[0200] In other embodiments of the present application, the processor 41 is used to execute the prompt words corresponding to the first log in the first log set based on the large language model in the log parsing rule generation program in the memory 42, and format each first log to obtain a log in the second format, so as to implement the following steps:

[0201] determining a target log from a plurality of first logs;

[0202] Based on the large language model and the first target prompt word corresponding to the target log, each first log is formatted to obtain a log in a second format.

[0203] In other embodiments of the present application, the processor 41 is used to execute the first target prompt word corresponding to the large language model and the target log in the log parsing rule generation program in the memory 42, and format each first log to obtain a log with a second format, so as to implement the following steps:

[0204] Determine a first regular expression of the target log based on the large language model and the first target prompt word;

[0205] For logs matching the first regular expression among the plurality of first logs, formatting the matching logs based on the first regular expression to obtain logs having a second format corresponding to the matching logs;

[0206] For logs in the plurality of first logs that do not match the first regular expression, each unmatched log is formatted based on the large language model and the second target prompt word corresponding to the unmatched log to obtain a log in a second format.

[0207] In other embodiments of the present application, the processor 41 is used to execute the second target prompt word corresponding to the large language model and the unmatched logs in the log parsing rule generation program in the memory 42, and format each unmatched log to obtain a log with a second format, so as to implement the following steps:

[0208] determining a first sub-target log from the unmatched logs;

[0209] Determine a second regular expression for the first sub-target log based on the large language model and the first sub-prompt word corresponding to the first sub-target log;

[0210] For the first sub-log that matches the second regular expression in the unmatched log, formatting the matched first sub-log based on the second regular expression to obtain a log having a second format corresponding to the matched first sub-log;

[0211] For a second sub-log that does not match the second regular expression in the unmatched log, determine a second sub-target log from the unmatched second sub-log;

[0212] Based on the large language model and the second sub-prompt word corresponding to the second sub-target log, a third regular expression of the second sub-target log is determined until a log having a second format corresponding to each unmatched second sub-log is determined.

[0213] In other embodiments of the present application, the processor 41 is used to execute the log parsing rule generation program in the memory 42 to parse the second log in the second log set based on the large language model to obtain the third-party log field information in the second log, so as to implement the following steps:

[0214] Determine a log document title corresponding to each second log; wherein the log document title is used to characterize the type of the log field of the second log;

[0215] Merging multiple second logs based on log document titles to obtain a target log;

[0216] For each target log, when it is determined based on the large language model that the target log is a security log, a third-party log field name is determined from the target log; wherein the third-party log field information includes the third-party log field name.

[0217] In other embodiments of the present application, the processor 41 is used to execute the log parsing rule generation program in the memory 42 to determine the first mapping relationship between the third-party log and the target log field name based on the large language model, log information and target log field name, so as to implement the following steps:

[0218] constructing a first prompt word based on the third-party log field name and the target field name in the third-party log field information;

[0219] Based on the first prompt word and the large language model, a first mapping relationship between the third-party log field name and the target field name is determined.

[0220] In other embodiments of the present application, the processor 41 is used to execute the log parsing rule generation program in the memory 42 to determine the second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value, so as to implement the following steps:

[0221] For the first log set, based on the first mapping relationship and the concerned field name in the target field name, determine the target third-party log field name from the third-party log field names;

[0222] Determine the third-party log field enumeration value corresponding to the target third-party log field name;

[0223] Constructing a second prompt word based on the third-party log field enumeration value and the target log field value;

[0224] Based on the second prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

[0225] In other embodiments of the present application, the processor 41 is used to execute the log parsing rule generation program in the memory 42 to determine the second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value, so as to implement the following steps:

[0226] For the second log set, obtaining target field information of the second log from the second log set; wherein the target field information includes a third-party log field enumeration value corresponding to each third-party log field name;

[0227] Constructing a third prompt word based on the third-party log field enumeration value and the target log field value;

[0228] Based on the third prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

[0229] In other embodiments of the present application, the processor 41 is used to execute the log parsing rule generation method in the log parsing rule generation program in the memory 42 to implement the following steps:

[0230] Get the parsing rule verification file;

[0231] The log parsing rules are verified using the parsing rule verification file to obtain the verification results.

[0232] It should be noted that the specific description of the steps performed by the processor can be referred to Figure 1 and Figure 2 The implementation process of the information processing method provided in the corresponding embodiment will not be repeated here.

[0233] The log parsing rule generating device provided in the embodiment of the present application utilizes large language model technology to realize an automated process of directly generating log parsing rules from raw log data and log specification standard documents without prior knowledge in a specific field, rather than relying on experts with deep network security knowledge and programming skills to manually create log parsing rules as in the related art. This not only realizes the standardization and generalization of the log parsing process, but also improves the efficiency of generating log parsing rules.

[0234] Based on the foregoing embodiments, the embodiments of the present application provide a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement Figure 1 and Figure 2The steps in the log parsing rule generation method provided by the corresponding embodiment. It should be noted that the above-mentioned computer-readable storage medium can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and other memories; it can also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0235] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0236] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0237] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0238] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0239] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0240] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0241] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A log parsing rule generation method, characterized in that: The method comprises: Obtain third-party logs corresponding to third-party devices; Processing the third-party log based on the large language model to obtain processed log information; Based on the large language model, the log information and the target log field name, determining a first mapping relationship between the third-party log and the target log field name; wherein the target log field name has a first format; Based on the large language model and the target log field value, determining a second mapping relationship between the third-party log and the target log field value; wherein the target log field value has the first format; A log parsing rule for the third-party log is generated based on the large language model, the log information, the first mapping relationship, and the second mapping relationship.

2. The method according to claim 1, characterized in that The third-party log includes a first log set and / or a second log set, and the third-party log is processed based on the large language model to obtain processed log information, including: For the first log set, based on the large language model and the prompt word corresponding to the first log in the first log set, format each of the first logs to obtain a log in a second format; wherein the log in the second format includes third-party log field information; the first log set includes multiple first logs; For the second log set, the second log in the second log set is parsed based on the large language model to obtain third-party log field information of the second log; wherein the second log set includes multiple second logs and description information of each second log.

3. The method according to claim 2, characterized in that The step of formatting each of the first logs based on the large language model and the prompt words corresponding to the first logs in the first log set to obtain a log in a second format includes: Determine a target log from the plurality of first logs; Based on the large language model and the first target prompt words corresponding to the target logs, each of the first logs is formatted to obtain a log having the second format.

4. The method according to claim 3, characterized in that The step of formatting each of the first logs based on the large language model and the first target prompt word corresponding to the target log to obtain a log having the second format includes: Determine a first regular expression of the target log based on the large language model and the first target prompt word; For logs matching the first regular expression among the plurality of first logs, formatting the matching logs based on the first regular expression to obtain logs having the second format corresponding to the matching logs; For logs that do not match the first regular expression among the plurality of first logs, formatting is performed on each of the unmatched logs based on the large language model and the second target prompt words corresponding to the unmatched logs to obtain logs having the second format.

5. The method according to claim 4, characterized in that The step of formatting each of the unmatched logs based on the large language model and the second target prompt word corresponding to the unmatched logs to obtain a log having the second format includes: Determining a first sub-target log from the unmatched logs; Determine a second regular expression for the first sub-goal log based on the large language model and a first sub-prompt word corresponding to the first sub-goal log; For a first sub-log that matches the second regular expression in the unmatched log, formatting the matched first sub-log based on the second regular expression to obtain a log having the second format corresponding to the matched first sub-log; For a second sub-log in the unmatched log that does not match the second regular expression, determine a second sub-target log from the unmatched second sub-log; Based on the large language model and the second sub-prompt word corresponding to the second sub-target log, a third regular expression of the second sub-target log is determined until a log having the second format corresponding to each of the unmatched second sub-logs is determined.

6. The method according to claim 2, characterized in that The parsing the second log in the second log set based on the large language model to obtain the third-party log field information of the second log includes: Determine a log document title corresponding to each of the second logs; wherein the log document title is used to characterize the type of the log field of the second log; Merging the plurality of second logs based on the log document titles to obtain a target log; For each of the target logs, when the target log is determined to be a security log based on the large language model, a third-party log field name is determined from the target log; wherein the third-party log field information includes the third-party log field name.

7. The method according to any one of claims 5 or 6, characterized in that: The determining, based on the large language model, the log information, and the target log field name, a first mapping relationship between the third-party log and the target log field name includes: constructing a first prompt word based on the third-party log field name in the third-party log field information and the target field name; Based on the first prompt word and the large language model, a first mapping relationship between the third-party log field name and the target field name is determined.

8. The method according to claim 7, characterized in that The determining, based on the large language model and the target log field value, a second mapping relationship between the third-party log and the target log field value includes: For the first log set, based on the first mapping relationship and the concerned field name in the target field name, determine the target third-party log field name from the third-party log field names; Determine a third-party log field enumeration value corresponding to the target third-party log field name; Constructing a second prompt word based on the third-party log field enumeration value and the target log field value; Based on the second prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

9. The method according to claim 6, characterized in that The determining, based on the large language model and the target log field value, a second mapping relationship between the third-party log and the target log field value includes: For the second log set, obtaining target field information of the second log from the second log set; wherein the target field information includes a third-party log field enumeration value corresponding to each third-party log field name; Constructing a third prompt word based on the third-party log field enumeration value and the target log field value; Based on the third prompt word and the large language model, a second mapping relationship between the third-party log field enumeration value and the target log field value is determined.

10. The method according to claim 1, characterized in that The method further comprises: Get the parsing rule verification file; The log parsing rule is verified using the parsing rule verification file to obtain a verification result.

11. A log parsing rule generation device, characterized in that: The device comprises: An acquisition unit, used to acquire third-party logs corresponding to third-party devices; A parsing unit, used for processing the third-party log based on a large language model to obtain processed log information; A mapping unit, configured to determine a first mapping relationship between the third-party log and the target log field name based on the large language model, the log information and the target log field name; wherein the target log field name has a first format; The mapping unit is further used to determine a second mapping relationship between the third-party log and the target log field value based on the large language model and the target log field value; wherein the target log field value has the first format; A generating unit is used to generate a log parsing rule for the third-party log based on the large language model, the log information, the first mapping relationship and the second mapping relationship.

12. A log parsing rule generation device, characterized in that: The device comprises: a processor, a memory and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the log parsing rule generation program stored in the memory to implement the steps of the log parsing rule generation method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the log parsing rule generation method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Data analysis rule matching method and device, storage medium and electronic equipment

    CN120658810A