Log processing method and device, electronic equipment and storage medium

CN116820881BActive Publication Date: 2026-09-04LINKAGE TECHNOLOGY (NANJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310780007.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-28
Publication Date
2026-09-04
Estimated Expiration
2043-06-28

AI Technical Summary

Technical Problem

[0004](1)当前的日志模板的更新时间较长、时效性较差

Benefits of technology

[0034]一方面,日志处理方法应用于分布式设备集群中的至少一台设备。通过设置分布式设备集群,可按照至少一条日志消息的数量启动适配数量的设备,然后在各设备上同时进行匹配和训练。该手段实现了对大量日志的分流处理,该分流处理为提升日志消息的匹配和训练效率奠定基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116820881B_ABST
    Figure CN116820881B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a log processing method and device, electronic equipment and storage medium, and relate to the field of system operation and maintenance. The method comprises: obtaining at least one log message, and determining, for each log message, whether there is a log template matching the log message from a preset template library; taking the log message without the matching log template as a first log message, performing regular processing on real-time variable information recorded in a content field of the first log message, and classifying each first log message to obtain at least one first log message set. For each first log message set, at least one new log template of the first log message set is obtained, and each new log template is updated to the preset template library. Embodiments of the present application can improve matching efficiency by sharing the preset template library, reduce the update time by first classifying and then determining the log template according to the category, and improve the training effect of the log template by regular processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of system operation and maintenance technology, and more specifically, to a log processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the widespread adoption of microservices, cloud platforms, and containerization, system operation and maintenance logs are becoming increasingly diverse and complex. Furthermore, the volume of logs is also growing. On the one hand, detecting anomalies in a large number of different types of operation and maintenance logs is quite challenging. On the other hand, users expect rapid anomaly alerts for complex log formats.

[0003] In the above scenario, log anomaly detection using the Drain algorithm has the following problems:

[0004] (1) The current log template has a long update time and poor timeliness.

[0005] (2) The interaction between the inference and training functions provided by the Drain algorithm is in the form of files, which has poor performance.

[0006] (3) The log content contains special information that interferes with the template training of the Drain algorithm, resulting in poor template training effect. Summary of the Invention

[0007] The purpose of this application is to solve one of the above-mentioned technical problems.

[0008] On one hand, embodiments of this application provide a log processing method applied to at least one device in a distributed device cluster, the cluster including a preset template library shared by all devices, the preset template library including multiple log templates; the method includes:

[0009] Obtain at least one log message, which includes at least one required field and real-time variable information recorded in each required field; the required field includes a content field. For each log message, determine whether a log template matching the log message exists in a preset template library. Log messages for which no matching log template exists are designated as first log messages. Regularize the real-time variable information recorded in the content field of the first log messages and classify each first log message to obtain at least one set of first log messages. Update the preset template library using log templates trained based on each set of first log messages.

[0010] Optionally, each log template includes at least one required field, as well as historical variable information or historical constant information recorded in each required field; the historical variable information and historical constant information are obtained from the real-time variable information recorded in the required fields of historical log messages.

[0011] This includes determining whether a log template matching the log message exists in the preset template library, including:

[0012] For each log template in the preset template library, the historical variable information or historical constant information recorded in each necessary field of the log template is compared with the real-time variable information recorded in each necessary field of the log message to obtain the similarity data of each necessary field of the log template. If there is a log template whose similarity data for each necessary field is greater than the preset similarity threshold, it is determined that there is a log template that matches the log message.

[0013] Optionally, at least one required field includes a feature field, wherein the real-time variable information recorded in the feature field has a distinguishability of each first log message reaching a preset threshold.

[0014] The first log messages are classified to obtain at least one set of first log messages, including:

[0015] Based on whether the real-time variable information recorded in the feature fields is the same, each first log message is classified to obtain at least one set of first log messages. Among them, the real-time variable information recorded in the preset necessary fields of each first log message in the first log message set is the same.

[0016] Optionally, each first log message can be classified to obtain at least one set of first log messages, including:

[0017] Feature information is extracted from the real-time variable information recorded in the content fields of each first log message. The first log messages are then classified according to whether their corresponding feature information is the same, resulting in at least one set of first log messages. Within each set of first log messages, the feature information of all first log messages is identical.

[0018] Optionally, a preset template library for updating log templates trained on various first log message sets is provided, including:

[0019] The training period is determined based on the number of at least one log message and the current training level; the training level characterizes the timeliness of the training process for the first log message set. For each type of the first log message set, the first log message set is trained according to the training period to obtain at least one new log template. The at least one new log template is then updated to the preset template library.

[0020] Optionally, the method further includes:

[0021] Within a preset statistical period, perform the following statistical operations: count the first number of new log templates and the second number of first log messages corresponding to each new log template. Based on the first number and each second number, obtain log monitoring information for the first log messages.

[0022] Optionally, the method further includes:

[0023] Log messages with matching log templates are used as second log messages, and log templates in the preset template library that match the log messages are marked.

[0024] The method also includes:

[0025] Within a preset statistical period, perform the following statistical operations: count the third number of tagged log templates in the preset template library, and the fourth number of corresponding second log messages for each tagged log template. Based on the third number and each fourth number, obtain log monitoring information for the second log messages.

[0026] On the other hand, embodiments of this application provide a log processing apparatus, which includes:

[0027] The acquisition module is used to acquire at least one log message. The log message includes at least one required field and real-time variable information recorded in each required field. The required field includes a content field.

[0028] The matching module is used to determine whether a log template that matches the log message exists in the preset template library for each log message.

[0029] The preprocessing and classification module is used to treat log messages that do not have a matching log template as first log messages, perform regular expression processing on the real-time variable information recorded in the content field of the first log messages, and classify each first log message to obtain at least one set of first log messages.

[0030] The update module is used to update the preset template library based on log templates trained on various first log message sets.

[0031] This application also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the above-described log processing method.

[0032] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the log processing method described above.

[0033] The beneficial effects of the technical solutions provided in this application are:

[0034] On one hand, the log processing method is applied to at least one device in a distributed device cluster. By setting up a distributed device cluster, an appropriate number of devices can be started according to the number of at least one log message, and then matching and training can be performed simultaneously on each device. This approach achieves the diversion of a large number of logs, which lays the foundation for improving the efficiency of log message matching and training.

[0035] On the other hand, a distributed device cluster includes a pre-defined template library shared by all devices. Sharing a pre-defined template library among devices, rather than each maintaining its own, allows all devices to match log messages within the same "log template set," improving the success rate of matching log messages with log templates.

[0036] On the other hand, regular expression processing is performed on the real-time variable information recorded in the content field of the first log message. This allows for the effective extraction of special information that may interfere with training in the first log message. Training based on this effective information can reduce or avoid interference from special information.

[0037] On the other hand, the first log messages are classified to obtain at least one set of first log messages of each category, and then training is performed based on each category of first log message set. Since the first log messages in each category have a high degree of similarity, training on the same category of log messages can reduce the difficulty and time of training, thereby quickly obtaining log templates. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0039] Figure 1 A flowchart illustrating a log processing method provided in an embodiment of this application;

[0040] Figure 2 A log message provided in an embodiment of this application;

[0041] Figure 3 This is a schematic diagram of the structure of a log processing system provided in an embodiment of this application;

[0042] Figure 4 This is a schematic diagram of the structure of a log processing device provided in an embodiment of this application;

[0043] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0044] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0045] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” indicates implementation as “A,” or implementation as “A,” or implementation as “A and B.”

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0047] First, let's introduce and explain several terms used in this application:

[0048] Drain is an online log parsing method based on a fixed-depth tree. When a new raw log message arrives, Drain preprocesses it using simple regular expressions based on domain knowledge. Then, Drain searches for log groups (i.e., leaf nodes of the tree) according to special design rules encoded in the internal nodes of the tree. If a suitable log group is found, the log message will match the log event stored in that log group; otherwise, a new log group will be created based on the log information.

[0049] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0050] This application provides a log processing method, such as... Figure 1As shown. This method can be applied to at least one device in a distributed device cluster, where the cluster includes a shared preset template library for each device, and the preset template library includes multiple log templates. The method includes the following steps S110–S140.

[0051] S110, obtain at least one log message, the log message includes at least one necessary field and real-time variable information recorded in each necessary field; the necessary field includes the content field.

[0052] Specifically, at least one log message is obtained from at least one message source. Each message source provides log messages according to a preset message protocol. The preset message protocol defines the necessary and unnecessary fields for a log message. The variable information recorded in the necessary fields of a log message must be non-empty; if it is empty, the log message is not processed. The variable information recorded in the unnecessary fields of a log message may or may not be empty.

[0053] In one example, a log message is provided, such as Figure 2 As shown. The required fields in the log message include: timestamp (log event stamp); organization (organization); logType (log type); method (method name); name (log object); severityText (log level); and body (log content). The non-required fields in the log message include: logTypeName (log type name); objType (log object type); and logPath (log file path). The body field is the content field.

[0054] S120, For each log message, determine from the preset template library whether there is a log template that matches the log message.

[0055] Each log template includes at least one required field, and historical variable information or historical constant information recorded in each required field; the historical variable information and historical constant information are obtained from the real-time variable information recorded in the required fields of historical log messages. Historical log messages are log messages obtained and processed before the current time.

[0056] In one example, a historical log message is provided, the content of which includes:

[0057] method: csclacDispatcherControlle;

[0058] body={Ack={MessageID=0025519582014131,

[0059] CommandCode=ACK,

[0060] ResponseCode = 100

[0061] ErrorMessage = 501;

[0062] severityText:INFO;

[0063] Other fields and their contents.

[0064] The log template obtained from this historical log message includes the following content:

[0065] method: csclacDispatcherControlle;

[0066] body = {Ack = {MessageID = <num>,

[0067] CommandCode=ACK,

[0068] ResponseCode= <num>,

[0069] ErrorMessage= <num>}};

[0070] severityText: INFO.

[0071] In the log template, the fields "method" and "severityText" record historical variable information, while the field "body" records historical constant information.

[0072] Optionally, for each log template in the preset template library, the historical variable information or historical constant information recorded in each necessary field of the log template is compared with the real-time variable information recorded in each necessary field of the log message to obtain the similarity data of each necessary field of the log template. If there is a log template whose similarity data for each necessary field is greater than the preset similarity threshold, it is determined that there is a log template that matches the log message.

[0073] Specifically, the similarity between log messages and log templates is determined by comparing the similarity of information recorded in the same fields of log templates and log messages.

[0074] Specifically, for the historical variable information recorded in the necessary fields of the log template, the historical variable information is compared with the real-time variable information recorded in the corresponding fields of the log message. If they are consistent, the similarity is 100%. Otherwise, the similarity is 0.

[0075] Specifically, for the historical constant information recorded in the necessary fields of the log template, the similarity between the historical constant information and the real-time variable information recorded in the corresponding fields of the log message is compared. More specifically, constants can be extracted from the real-time variable information recorded in the corresponding fields of the log message to obtain all constants in the real-time variable information; further, the constants in the historical constant information are compared with the constants obtained from the real-time variable information to determine the similarity; if the constants in the historical constant information correspond one-to-one with the constants obtained from the real-time variable information, the similarity is determined to be 100%; if there are non-corresponding fields, the similarity is determined by combining the non-corresponding fields with the corresponding fields.

[0076] Each device also shares a preset similarity threshold. If a log template exists with similarity data for all necessary fields exceeding the preset similarity threshold, it is determined that a log template matching the log message exists; if no log template exists with similarity data for all necessary fields exceeding the preset similarity threshold, it is determined that no log template matching the log message exists. The preset similarity threshold has a maximum value of 100% and a minimum value of 0%.

[0077] S130, log messages for which no matching log template exists are taken as first log messages, the real-time variable information recorded in the content field of the first log message is processed by regular expression, and each first log message is classified to obtain at least one set of first log messages.

[0078] In one example, the content field of a log message is provided to record the log content:

[0079] IDMM{

[0080] "220";["state":true"astRunTime":"20220630000016""lastDealTime":"20220629212250"},

[0081] "221":{"state";true,"lastRunTime":"20220630000015","lastDealTime":"20220629203437"}

[0082] }

[0083] After performing regular expression processing on this log message, the resulting log content is as follows:

[0084] IDMM{

[0085] "XX";["state":XX"LastRunTime":"XX""lastDealTime":"XX"]

[0086] }

[0087] In this process, log messages with matching log templates are used as second log messages, and log templates in the preset template library that match the log messages are marked.

[0088] S140 is a preset template library for updating log templates trained on various first log message sets.

[0089] In one implementation of this embodiment, step S140 includes the following steps Sa1 to Sa3.

[0090] Sa1 determines the training period based on the number of at least one log message and the current training level.

[0091] The training level represents the timeliness of the training process for the first log message set. For example, during the initial launch of a new application, a large number of new log templates are generated. At this time, a high timeliness is required, meaning the training time needs to be set shorter to obtain new log templates and update the preset template library. In this case, the training level is set to a high level. After a period of time, when the new application enters a stable operating phase, fewer log types are generated. Therefore, a high timeliness is not necessary, and the training period can be set longer.

[0092] There is a pre-defined correspondence, which includes multiple training periods, the number of log messages in each training period, and the training level. A larger number of log messages corresponds to a longer training period, while a higher training level corresponds to a shorter training period.

[0093] In one example, the training period is determined to be 2 minutes based on the number of at least one log message and the current training level.

[0094] Sa2: For each type of first log message set, train the first log message set according to the training cycle to obtain at least one new log template.

[0095] In one example, the Drain algorithm can be trained based on various types of first log message sets to obtain the log messages corresponding to each type of first log message set. For information on the training process of the Drain algorithm, please refer to relevant technical documentation.

[0096] Sa3 updates at least one new log template to the preset template library.

[0097] This application also provides an embodiment to illustrate how to classify the first log message.

[0098] For log messages, the log content is a crucial component, reflecting a wealth of information and being a key focus of the training process. Therefore, this embodiment provides an implementation method for classification based on log content.

[0099] Feature information is extracted from the real-time variable information recorded in the content fields of each first log message; the first log messages are classified according to whether their corresponding feature information is the same, resulting in at least one set of first log messages. Within each set of first log messages, the feature information of all first log messages is identical.

[0100] Optionally, feature information of a preset length can be extracted from the real-time variable information recorded in the content field, starting from the beginning position.

[0101] For log messages, the similarity is particularly high when other fields record the same information. Therefore, this embodiment provides a method for classification based on feature fields.

[0102] Among them, at least one necessary field of the log message includes a feature field, and the real-time variable information recorded in the feature field has a distinguishability of each first log message reaching a preset threshold.

[0103] In one example, the "method" field records the name of the method that generated the log message. Generally, different methods provide different log message templates. Therefore, the "method" field can be used to distinguish the first log message from log messages generated by various methods.

[0104] Based on whether the real-time variable information recorded in the feature fields is the same, each first log message is classified to obtain at least one set of first log messages; wherein, the real-time variable information recorded in the preset necessary fields of each first log message in the first log message set is the same.

[0105] After processing the log messages, the most important step is to perform anomaly detection based on the processing results to determine if there are any abnormal logs such as "sudden increases or decreases". To this end, this application also provides an embodiment to illustrate the anomaly detection process.

[0106] In one implementation of this embodiment, the method further includes:

[0107] Within the preset statistical period, perform the following statistical operations:

[0108] Count the first number of new log templates and the second number of first log messages corresponding to each new log template; based on the first number and each second number, obtain log monitoring information for the first log messages.

[0109] Optionally, each device can share at least one monitoring threshold. The monitoring threshold can be either a maximum monitoring threshold or a minimum monitoring threshold.

[0110] Specifically, the first quantity and each of the second quantities are compared with the corresponding highest monitoring threshold to obtain log monitoring information mainly showing sudden increases. The first quantity and each of the second quantities are compared with the corresponding lowest monitoring threshold to obtain log monitoring information mainly showing sudden decreases.

[0111] For example, if the first quantity exceeds the corresponding monitoring threshold, a sudden increase in new log templates is identified. Further, each second quantity is compared with its corresponding monitoring threshold; if a second quantity exceeds its corresponding monitoring threshold, a sudden increase in log messages for the log template corresponding to that second quantity is identified.

[0112] The log monitoring information includes identification information for various new log templates, as well as the second number of log messages corresponding to each new log template.

[0113] In another implementation of this embodiment, the method further includes:

[0114] Within the preset statistical period, perform the following statistical operations:

[0115] The third number of tagged log templates in the preset template library and the fourth number of corresponding second log messages for each tagged log template are counted; based on the third number and each fourth number, log monitoring information for the second log messages is obtained.

[0116] Specifically, the third and fourth quantities are compared with their corresponding highest monitoring thresholds to obtain log monitoring information primarily showing sudden increases. Conversely, the third and fourth quantities are compared with their corresponding lowest monitoring thresholds to obtain log monitoring information primarily showing sudden decreases.

[0117] In another implementation of this embodiment, the method further includes:

[0118] Output the log monitoring information for the first and second log messages.

[0119] To more clearly illustrate the technical effects of the log processing method provided in the embodiments of this application, an example of a log processing system is also provided in the embodiments of this application. For example... Figure 3 As shown, the log processing system comprises three parts: an external log system, a log anomaly detection device, and an alarm ticket system. The log anomaly detection device includes a message input interface module, a Drain algorithm module, a data caching module (including a preset template library), a time-series anomaly detection module, and a message output interface. Each log system corresponds to a device in a distributed device cluster. An external log system constitutes a message source.

[0120] This example is also based on Figure 3 The log anomaly detection device shown provides an execution flow for a log processing method. This execution flow includes the following steps S1001 to S1009.

[0121] S1001, access log messages conforming to the preset message protocol.

[0122] Specifically, log messages are obtained from message sources such as "number portability" through the message input interface.

[0123] S1002, use the Drain algorithm to reason about the log template to determine whether a matching log template exists in the preset template library.

[0124] If it does not exist, proceed to S1003. If it exists, identify the log message as the second log message, mark the matching log template in the preset template library, and proceed to S1007.

[0125] S1003, perform regular expression processing on the log content of each first log message.

[0126] S1004, classify the first log messages to obtain at least one set of first log messages.

[0127] S1005: Use the Drain algorithm to classify and train various types of first log message sets to obtain new log templates, and then execute S1006 and S1007 respectively.

[0128] S1006, Update the new log template to the preset template library.

[0129] S1007, count the first number of new log templates and the second number of first log messages corresponding to each new log template; count the third number of marked log templates and the fourth number of log messages corresponding to each marked log template.

[0130] S1008, Statistics Data, obtain log monitoring information.

[0131] Within a preset statistical period, obtain various data related to new log templates, as well as various data related to tagged log templates.

[0132] S1009, Error message output.

[0133] From the first quantity, each of the second quantities, the third quantity, and each of the fourth quantities, determine whether there is any abnormal information of sudden increase or decrease.

[0134] If such anomalies exist, they can be output to the alarm work order system to generate alarm work orders and notify relevant personnel to handle the anomalies.

[0135] Figure 4 A log processing device 400 is shown, applied to at least one device in a distributed device cluster. The cluster includes a preset template library shared by all devices, and the preset template library includes multiple log templates. The device 400 includes the following modules:

[0136] The module 410 is used to obtain at least one log message, which includes at least one necessary field and real-time variable information recorded in each necessary field; the necessary field includes a content field.

[0137] The matching module 420 is used to determine, for each log message, whether there is a log template that matches the log message from the preset template library;

[0138] The preprocessing and classification module 430 is used to take log messages that do not have a matching log template as first log messages, perform regular expression processing on the real-time variable information recorded in the content field of the first log messages, and classify each first log message to obtain at least one set of first log messages.

[0139] Update module 440 is used to update the preset template library based on log templates trained on various first log message sets.

[0140] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.

[0141] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of a log processing method. Compared with the prior art, the following improvements are made: matching efficiency can be improved by sharing a preset template library; update time can be reduced by classifying logs first and then determining log templates by category; and the training effect of log templates can be improved by using regular expression processing.

[0142] See Figure 5 This application also provides a specific example of an electronic device. Figure 5 The illustrated electronic device 5000 includes a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, for example, via a bus 5002. Optionally, the electronic device 5000 may further include a transceiver 5004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 5004 is not limited to one type, and the structure of the electronic device 5000 does not constitute a limitation on the embodiments of this application.

[0143] Processor 5001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 5001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0144] Bus 5002 may include a path for transmitting information between the aforementioned components. Bus 5002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 5002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0145] The memory 5003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.

[0146] The memory 5003 is used to store computer programs that execute the embodiments of this application, and its execution is controlled by the processor 5001. The processor 5001 is used to execute the computer programs stored in the memory 5003 to implement the steps shown in the foregoing method embodiments.

[0147] Electronic devices include, but are not limited to: servers

[0148] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.

[0149] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0150] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.

[0151] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.< / num> < / num> < / num>

Claims

1. A log processing method, characterized in that, It is applied to at least one device in a distributed device cluster, wherein the cluster includes a preset template library shared by each device, and the preset template library includes multiple log templates; The method includes: Obtain at least one log message, the log message including at least one necessary field and real-time variable information recorded in each necessary field; the necessary field includes a content field; For each log message, determine from the preset template library whether there is a log template that matches the log message; Log messages for which no matching log template exists are taken as first log messages. The real-time variable information recorded in the content field of the first log message is processed by regular expression, and each first log message is classified to obtain at least one set of first log messages. The preset template library is updated based on log templates trained from various first log message sets; The process of updating the preset template library using log templates trained based on various first log message sets specifically includes: The training period is determined based on the number of at least one log message and the current training level; the training level characterizes the timeliness of the training process for the first log message set. For each type of first log message set, train the first log message set according to the training period to obtain at least one new log template; Update the at least one new log template to the preset template library.

2. The method according to claim 1, characterized in that, Each log template includes at least one required field, as well as historical variable information or historical constant information recorded in each required field; the historical variable information and historical constant information are obtained from the real-time variable information recorded in the required fields of historical log messages; Determining whether a log template matching the log message exists in the preset template library includes: For each log template in the preset template library, the historical variable information or historical constant information recorded in each necessary field of the log template is compared with the real-time variable information recorded in each necessary field of the log message to obtain the similarity data of each necessary field of the log template. If there is a log template whose similarity data for each necessary field is greater than the preset similarity threshold, it is determined that there is a log template that matches the log message.

3. The method according to claim 1, characterized in that, The at least one necessary field includes a feature field, and the real-time variable information recorded by the feature field has a distinguishability of each first log message reaching a preset threshold. The process of classifying each first log message to obtain at least one set of first log messages includes: Based on whether the real-time variable information recorded in the feature fields is the same, each first log message is classified to obtain the at least one type of first log message set; Among them, the real-time variable information recorded in the preset necessary fields of each first log message in the first log message set is the same.

4. The method according to claim 1, characterized in that, The process of classifying each first log message to obtain at least one set of first log messages includes: Feature information is extracted from the real-time variable information recorded in the content fields of each first log message; Based on whether the corresponding feature information of each first log message is the same, the first log messages are classified to obtain the at least one type of first log message set; In each category of the first log message set, the characteristic information of each first log message is the same.

5. The method according to claim 1, characterized in that, The method further includes: Within the preset statistical period, perform the following statistical operations: Count the first number of new log templates, and the second number of the first log messages corresponding to each new log template; Based on the first quantity and each of the second quantities, log monitoring information for the first log message is obtained.

6. The method according to claim 2, characterized in that, The method further includes: Log messages with matching log templates are used as second log messages, and log templates in the preset template library that match the log messages are marked. The method further includes: Within the preset statistical period, perform the following statistical operations: The third number of marked log templates in the preset template library and the fourth number of second log messages corresponding to each marked log template are counted. Based on the third quantity and each of the fourth quantities, log monitoring information for the second log message is obtained.

7. A log processing device, characterized in that, It is applied to at least one device in a distributed device cluster, wherein the cluster includes a preset template library shared by each device, and the preset template library includes multiple log templates; The device includes: The acquisition module is used to acquire at least one log message, the log message including at least one necessary field and real-time variable information recorded in each necessary field; the necessary field includes a content field; The matching module is used to determine, for each log message, whether there is a log template that matches the log message from the preset template library; The preprocessing and classification module is used to take log messages that do not have a matching log template as first log messages, perform regular expression processing on the real-time variable information recorded in the content field of the first log messages, and classify each first log message to obtain at least one set of first log messages. The update module is used to update the preset template library based on log templates trained on various first log message sets; The update module is specifically used for: The training period is determined based on the number of at least one log message and the current training level; the training level characterizes the timeliness of the training process for the first log message set. For each type of first log message set, train the first log message set according to the training period to obtain at least one new log template; Update the at least one new log template to the preset template library.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the log processing method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the log processing method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Log abnormity early warning method and device, electronic equipment and storage medium

    CN114064434A

  • 5G weblog compression method and device, terminal equipment and storage medium

    CN115757308A