Method, electronic device and computer program product for analyzing log files
By converting log records into log identifiers and using machine learning models for similarity comparison, the problems of incomplete rules and the unavailability of NLP in log file analysis are solved, achieving efficient and accurate log monitoring and fault detection.
Patent Information
- Application Number
- CN202010730016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-27
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2040-07-27
AI Technical Summary
Existing rule-based log file analysis mechanisms cannot exhaustively enumerate all log records, resulting in the omission of log records that do not conform to the predetermined rules. Furthermore, the analysis results of log files by natural language processing technology may be unusable, increasing human resources and rule maintenance costs.
By determining the template of log records and converting it into log identifiers, a similarity comparison model is used to select the log identifier to be analyzed and its context log identifiers corresponding to the predetermined event, and then analysis is performed in combination with machine learning or deep learning models.
It enables accurate and efficient monitoring of log records, timely detection of system problems, saving manpower costs, and avoiding misidentification by traditional methods.
Smart Images

Figure CN113986643B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to the field of computers, and more specifically, to a method, an electronic device and a computer program product for analyzing a log file. BACKGROUND
[0002] Currently, a log recording mechanism in a large IT system is used to record events that occur when operating the system or other software, and to generate a log file consisting of a number of log records. In turn, the process of analyzing log records in the log file is the process of analyzing the log file. The analysis process of the log file is an important means in the system monitoring and troubleshooting process. For example, in the system monitoring process, when a critical log (such as thread crash and critical service failure) occurs, an alarm and a notification can be sent. Usually, the detection of the critical log is a rule-based matching process, that is, the monitoring system scans the log records and issues a notification when a log record that meets a specific rule occurs. However, the rule-based matching process cannot exhaust all possibilities and still requires a large amount of human cost. SUMMARY
[0003] Embodiments of the present disclosure provide a method, an electronic device and a corresponding computer program product for analyzing a log file.
[0004] In a first aspect of the present disclosure, a method for analyzing a log file is provided. The method can include determining, based on a plurality of reference templates, a respective template of a plurality of log records in the log file. The method can further include determining the plurality of log records as a plurality of log identities corresponding to the respective templates, respectively. The method further includes determining, from the plurality of log identities, a to-be-analyzed log identity corresponding to a predetermined event. In addition, the method can further include selecting, from a plurality of reference log identities corresponding to the plurality of reference templates, a target reference log identity, the first similarity of the target reference log identity to the to-be-analyzed log identity being higher than a first threshold similarity.
[0005] In some embodiments, the method can further include processing a log record corresponding to the to-be-analyzed log identity based on a predetermined diagnosis strategy of the target reference log identity for the predetermined event.
[0006] In some embodiments, selecting the target reference log identifier can include: obtaining a context log identifier associated with the log identifier to be analyzed and a reference context log identifier associated with the target reference log identifier; determining a second similarity between the context log identifier and the reference context log identifier; determining a comprehensive similarity based on the first similarity and the second similarity; and in accordance with a determination that the comprehensive similarity is higher than a second threshold similarity, processing log records corresponding to the log identifier to be analyzed based on a predetermined diagnosis strategy of the target reference log identifier for the predetermined event.
[0007] In some embodiments, the context log identifier is separated from the log identifier to be analyzed by a time interval that is less than a threshold time interval, and the reference context log identifier is separated from the target reference log identifier by a time interval that is less than the threshold time interval.
[0008] In some embodiments, determining the respective templates of the plurality of log records can include: obtaining the plurality of reference templates from a reference template database; and in response to a first log record in the plurality of log records matching a first reference template in the plurality of reference templates, determining the first reference template as a template of the first log record.
[0009] In some embodiments, determining the plurality of log records as the plurality of log identifiers respectively includes: obtaining a mapping relationship between the plurality of reference templates and the plurality of reference log identifiers from the reference template database; obtaining a reference log identifier corresponding to the first reference template based on the mapping relationship; and determining the reference log identifier corresponding to the first reference template as a log identifier of the first log record.
[0010] In some embodiments, determining the log identifier to be analyzed corresponding to the predetermined event includes: determining the log identifier to be analyzed at a time when the predetermined event occurs based on timestamp information of the log file.
[0011] In some embodiments, the predetermined event includes at least one of: a thread crash event; a user report event; and a system metric abnormal event.
[0012] In a second aspect of the disclosure, an electronic device is provided. The device can include at least one processing unit, and at least one memory coupled to the at least one processing unit and storing machine executable instructions that, when executed by the at least one processing unit, cause the device to perform actions that can include determining, based on a plurality of reference templates, respective templates of a plurality of log records in a log file. The method can further include determining the plurality of log records as a plurality of log identities corresponding to the respective templates, respectively. The method further includes determining, from the plurality of log identities, a log identity to be analyzed corresponding to a predetermined event. In addition, the method can further include selecting, from a plurality of reference log identities corresponding to the plurality of reference templates, a target reference log identity, the target reference log identity having a first similarity to the log identity to be analyzed higher than a first threshold similarity.
[0013] In a third aspect of the disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer readable medium and comprises machine executable instructions that, when executed, cause a machine to perform the steps of the method according to the first aspect.
[0014] The summary is provided to introduce a selection of concepts, in a simplified form, that are further described below in the detailed description. This summary is not intended to identify key or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0015] The above and other objects, features and advantages of the disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which like reference characters designate the same components in several views.
[0016] Figure 1 A schematic diagram showing an example environment in which embodiments of the disclosure can be implemented is shown;
[0017] Figure 2 A schematic diagram showing a detailed example environment in which embodiments of the disclosure can be implemented is shown;
[0018] Figure 3 A flowchart showing a process for analyzing a log file according to embodiments of the disclosure is shown;
[0019] Figure 4 An example graph showing a correspondence between log records, templates, log identities according to embodiments of the disclosure is shown;
[0020] Figure 5 A flowchart showing a process for analyzing a log file according to embodiments of the disclosure is shown; and
[0021] Figure 6 A block diagram of a computing device capable of implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0022] Preferred embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0023] The term "comprising" and variations thereof as used herein are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an overly literal sense unless expressly so defined herein.
[0024] In order to monitor a system, a rule-based log file analysis mechanism is usually established. Although a conventional rule-based log file analysis mechanism can monitor log records that meet rules pre-set by an administrator perfectly, the rule-based log file analysis mechanism cannot exhaust all log records and is not flexible enough for different types of log records for which the administrator has not pre-set rules. Therefore, newly emerging log records that do not meet the pre-set rules can be missed by the system monitoring. In addition, in order to set more perfect rules, the administrator also studies the newly emerging log records and updates the rules, which also increases the construction and maintenance cost of the rules.
[0025] In addition, although natural language processing (NLP) technology has been applied to identify the semantics of log records in order to perform analysis of log files, the analysis result of the log files based on the natural language processing technology can be unavailable due to the fact that most log records are generated line by line by a specific code library, i.e., the expression of the log records is different from the expression of human language.
[0026] To at least partially address the aforementioned and other potential problems and shortcomings, embodiments of this disclosure propose a scheme for analyzing individual log records in a log file. In this scheme, firstly, a template (pattern, also called a "schema") corresponding to the log record is determined. Then, based on a pre-set mapping relationship between the template and log identifiers, the log record is converted into a corresponding log identifier. Next, at least one set of log records is extracted from the log file by selecting the log identifier to be analyzed corresponding to a predetermined event and its surrounding context log identifiers. The log identifier to be analyzed and the context log identifiers are then compared with target reference log identifiers and reference context log identifiers stored in a database, thereby selecting the representation set with the highest or highest similarity to this set, providing a reference for subsequent operations on the predetermined event. Therefore, this scheme can accurately and efficiently monitor log records, thereby enabling timely identification of system needs and saving manpower costs. The following first combines... Figure 1 Discuss the basic concept of this disclosure.
[0027] Figure 1 A schematic diagram of an example environment 100 in which various embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 includes computing device 110, log file 120, and comparison result 130. Furthermore, computing device 110 also includes comparison model 140. Log file 120 may contain a collection of log records used to display the operational processes of an IT system, such as log record 150-1, log record 150-2, ..., log record 150-N (hereinafter collectively referred to as log record 150). Computing device 110 may receive log file 120 and convert the log records 150 therein into log identifiers, and determine comparison result 130 using comparison model 140 in computing device 110. Comparison result 130 may identify a target reference log identifier in the database that has the highest or higher similarity to one or more log identifiers converted from log records 150. A processing strategy for the corresponding log record can be determined based on the diagnostic strategy for a predetermined event using this target reference log identifier.
[0028] exist Figure 1In particular, the key to generating the comparison result 130 based on the log file 120 lies in two points. Firstly, the comparison model 140 can determine the to-be-analyzed log identifier and the context log identifier corresponding to the predetermined event, and compare the to-be-analyzed log identifier, the context log identifier with the target reference log identifier and the reference context log identifier in the database respectively, so as to determine the comprehensive comparison result 130. Secondly, the comparison model 140 in the computing device 110 is a model based on known algorithms, for example, the corresponding template of the log record can be determined based on the Hamming distance algorithm or the Hashing algorithm, and the similarity comparison result between the to-be-analyzed log identifier sequence and the corresponding reference log identifier sequence in the database can be determined based on the Levenshtein distance algorithm. Preferably, the comparison model 140 can also be a pre-trained machine learning or deep learning model, such as BERT (Bidirectional Encoder Representations from Transformers) and the like. The construction and use of the comparison model 140 will be described below. Figure 2 The construction and use of the comparison model 140 will be described below.
[0029] Figure 2 A schematic diagram of a detailed example environment 200 in which embodiments of the present disclosure can be implemented is shown. As with the example environment 100, the example environment 200 can include a computing device 110, a log file 120, and a comparison result 130. Figure 1 Similarly, the example environment 200 can include a computing device 110, a log file 120, and a comparison result 130. The difference is that the example environment 200 can generally include a model training system 260 and a model application system 270. As an example, the model training system 260 and / or the model application system 270 can be implemented by a computing device 110 as shown in FIG. 1, for example. It should be understood that the structure and function of the example environment 200 described for illustrative purposes only are not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions. Figure 1 or Figure 2 It should be understood that the structure and function of the example environment 200 described for illustrative purposes only are not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions.
[0030] As mentioned previously, the scheme for comparing log identifiers according to the present disclosure can be divided into two stages: a model training stage and a model application stage. In the model training stage, the model training system 260 can train the comparison model 140 for comparing log identifiers using a log data set 250. In the model application stage, the model application system 270 can receive the trained comparison model 140 and the log file 120, thereby generating the comparison result 130. In certain embodiments, the log data set 250 can be a large number of annotated log records or corresponding log identifiers.
[0031] It should be understood that the comparison model 140 can also be constructed as a deep learning network for comparing log identifiers. Such a deep learning network can also be referred to as a deep learning model. In some embodiments, the deep learning network for comparing log identifiers may include multiple networks, where each network may be a multi-layer neural network, which may consist of a large number of neurons. Through the training process, the corresponding parameters of the neurons in each network can be determined.
[0032] The technical solutions described above are for illustrative purposes only and are not intended to limit the invention. To more clearly explain the principles of the above solutions, the following will refer to... Figure 3 Let's describe the process of analyzing log files in more detail.
[0033] Figure 3 A flowchart of a process or method 300 for analyzing log files according to embodiments of the present disclosure is shown. In some embodiments, method 300 may be performed in... Figure 6 Implemented in the device shown. As an example, method 300 can be implemented in the device shown. Figure 1 , Figure 2 This is implemented in the computing device 110 shown. Now refer to... Figure 1 describe Figure 3 The illustrated process or method 300 for analyzing log files according to an embodiment of this disclosure is shown. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of this disclosure.
[0034] At 310, computing device 110 can determine the corresponding template for multiple log records 150 in log file 120 based on multiple reference templates stored in a database. In some embodiments, computing device 110 can obtain multiple reference templates from a reference template database, and when a first log record (e.g., log record 150-2) among the multiple log records 150 matches a first reference template among the multiple reference templates, the first reference template is determined as the template for the first log record.
[0035] Figure 4 This is an example diagram 400 illustrating the correspondence between log records, templates, and log identifiers according to embodiments of this disclosure. Figure 4 As shown, the "Log Records" column 430 displays multiple log records, some of which follow specific rules. Wildcards can be used to identify templates for log records with the same specific rules; these templates are the wildcard segments in the "Template" column 420. Therefore, several log records can be matched with a single template. The "Log Identifier" column 410 displays multiple strings consisting of numbers and letters, which can be used to uniquely indicate the identity information (ID) of the corresponding template.
[0036] At 320, the computing device 110 can determine the plurality of log records as a plurality of log identifiers corresponding to the respective templates, respectively. As shown in Figure 4 FIG. 3, one template can correspond to one log identifier. It should be understood that although Figure 4 the correspondence among log records, templates, and log identifiers is shown in FIG. 3, in order to save storage resources, only the correspondence between templates and log identifiers can be stored in the database.
[0037] In some embodiments, the computing device 110 can obtain a mapping relationship between a plurality of reference templates and a plurality of reference log identifiers from the reference template database, as shown in Figure 4 FIG. 4. Based on the mapping relationship, the computing device 110 can obtain a reference log identifier corresponding to the first reference template in the above embodiment. Further, the computing device 110 can convert the first log record (e.g., log record 150-2) in the above embodiment into the reference log identifier corresponding to the first reference template. As an example, when the computing device receives the log file 120, the received log file can be scanned, and each log record therein can be compared with the logs or templates thereof stored in the database, and finally the log identifier corresponding to the matched template is returned. If a certain log record does not match any template, a new template can be generated, and the mapping relationship between the reference templates and the reference log identifiers is updated. In addition, in order to simplify the processing process, certain reference log identifiers in the database can be pre-labeled as “key log identifiers” corresponding to predetermined events, so that the to-be-analyzed log identifier corresponding to the predetermined event can be determined more quickly in subsequent operations.
[0038] At 330, the computing device 110 can determine the to-be-analyzed log identifier corresponding to the predetermined event from the plurality of log identifiers. In some embodiments, the computing device 110 can determine the to-be-analyzed log identifier when the predetermined event occurs based on the timestamp information of the log file. As an example, when a predetermined event such as a user report (a report initiated by a user who finds that there may be a system problem) occurs, the log record generated at that moment can be found, and then the log identifier thereof is determined as the to-be-analyzed log identifier. It should be understood that the user report is only an example of the predetermined event. The predetermined event can also include a thread crash event, a system indicator abnormal event, etc. In this way, log records that are not easily accurately identified by NLP models can be converted into log identifiers or log identifier sequences that can be accurately identified.
[0039] In addition, in other embodiments, as described in the above embodiments, since certain reference log identifiers in the database are pre-labeled as “key log identifiers”, as long as it is found that there is a log identifier matching the “key log identifier” in the plurality of log identifiers, it can be determined as the to-be-analyzed log identifier.
[0040] At 340, the computing device 110 can select a target reference log identifier from the plurality of reference log identifiers corresponding to the plurality of reference templates, the target reference log identifier having a first similarity to the log identifier to be analyzed that is higher than a first threshold similarity. Preferably, the computing device 110 can determine the reference log identifier that is most similar or identical to the log identifier to be analyzed as the target reference log identifier.
[0041] When the target reference log identifier is determined, the computing device 110 can process the log record corresponding to the log identifier to be analyzed based on a predetermined diagnostic strategy for the predetermined event of the target reference log identifier. The computing device 110 can automatically process the corresponding log record with reference to the predetermined diagnostic strategy, thereby improving system performance and saving labor costs.
[0042] It should be understood that in most cases, the log identifier to be analyzed is only one log identifier and is used to represent one log record. In this case, only one pair of log identifiers needs to be compared to determine the first similarity. However, there can also be a case where the log identifier to be analyzed contains multiple log identifiers and is used to represent multiple log records, for example, a fault report generated by the Backtrace function, which contains multiple log records, i.e., multiple converted log identifiers. At this time, the multiple log identifiers contained in the log identifier to be analyzed and the corresponding reference log identifiers in the database can be compared using algorithms such as the Levenshtein distance algorithm. Of course, machine learning or deep learning models can also be used for comparison.
[0043] It should also be understood that the above description with reference to Figure 3 is only for the operation of one or a group of log identifiers to be analyzed corresponding to the predetermined event. The one or the group of log identifiers to be analyzed can be regarded as the "kernel" of the partial log file. However, in addition to the "kernel", the context log identifiers thereof, i.e., other log identifiers that are a predetermined time period or a predetermined number of entries away from the "kernel", also have reference value. The specific embodiments of integrating context log identifiers will be described below with reference to Figure 5 .
[0044] Figure 5 A flowchart of a process or method 500 for analyzing a log file according to an embodiment of the present disclosure is shown. In certain embodiments, the method 500 can be implemented in the devices shown. As an example, the method 500 can be implemented in the computing device 110 shown. Figure 6 . Figure 1 , Figure 2 The computing device 110 shown. Referring now to the description of Figure 1 Figure 5 A process or method 500 for analyzing log files according to embodiments of the present disclosure is shown. For ease of understanding, the specific data mentioned in the following description are all exemplary and do not serve to limit the protection scope of the present disclosure.
[0045] At 510, the computing device 110 can obtain context log identifiers associated with the above-mentioned log identifier to be analyzed and reference context log identifiers associated with the above-mentioned target reference log identifier. In some embodiments, the time interval between the context log identifiers and the log identifier to be analyzed is less than a threshold time interval. For example, the context log identifiers include all log identifiers half an hour before the log identifier to be analyzed and all log identifiers half an hour after the log identifier to be analyzed. Correspondingly, the time interval between the reference context log identifiers and the target reference log identifier is less than the above-mentioned threshold time interval. For example, the reference context log identifiers include all reference log identifiers half an hour before the target reference log identifier and all reference log identifiers half an hour after the target reference log identifier in the database.
[0046] It should also be understood that there can be idle time intervals in the log identifiers near the log identifier to be analyzed, i.e., there can be default entries in the determined context log identifiers. In this case, the reference value of the context log identifiers can be reduced. Therefore, the context log identifiers within a predetermined number of entries from the log identifier to be analyzed can also be determined according to the number of entries of the log identifiers.
[0047] At 520, the computing device 110 can determine a second similarity between the context log identifiers and the reference context log identifiers. As mentioned above, the context log identifiers are log identifiers near the log identifier to be analyzed. Since the context log records have been converted into context log identifier sequences by finding the mapping relationship between the log identifiers in the database and the templates, such sequences are similar to word sequences in articles. Therefore, the context log identifier sequences can be processed using NLP techniques. The NLP techniques can include “bag of words”, TF-IDF, BERT, etc. As an example, the weighted frequency of each log identifier can be calculated using, for example, the TF-IDF algorithm to obtain a TF-IDF vector after extracting the context log identifiers. Finally, the similarity between the context log identifiers and the reference context log identifiers can be calculated by comparing the TF-IDF vectors. It should be understood that machine learning or deep learning models can also be used for comparison.
[0048] At 530, the computing device 110 can determine a comprehensive similarity based on the first similarity and the second similarity. For example, the comprehensive similarity can be determined by performing a weighted average of the first similarity and the second similarity. In addition, to make the comparison result more biased towards the first similarity (i.e., biased towards the log identification to be analyzed), the comprehensive similarity can be determined as a weighted average of the first similarity and the second similarity only when the first similarity is greater than a predetermined threshold, and otherwise only the first similarity is considered.
[0049] Further, at 540, the computing device 110 can determine whether the comprehensive similarity is higher than a second threshold similarity. If the comprehensive similarity is higher than the second threshold similarity, proceed to 550. At 550, the computing device 110 can process the log record corresponding to the log identification to be analyzed based on a predetermined diagnostic policy for the predetermined event of the target reference log identification. In some embodiments, the comprehensive similarity represents a similarity between a log segment in the log file to be analyzed and a reference log segment in the plurality of reference log records. The set of log identifications of the log segment includes the log identification to be analyzed and the contextual log identification, and the set of log identifications of the reference log segment corresponds to the target reference log identification information and the reference contextual log identification information.
[0050] Note that the above operations provide a technical solution of comparing the log identification to be analyzed and its contextual log identification with a set of log identifications in a database. The comparison can be used as further log analysis. The log analysis can be used to determine similar system problems, determine associated knowledge bases, and thus resolve similar service requests.
[0051] By implementing the above process, the comparison with the reference log records in the database can be simplified by converting the log records in the log file into log identifications corresponding to templates. In addition, the above process can find a target reference log identification most relevant to the one or the set of log identifications to be analyzed by comparing the one or the set of log identifications to be analyzed most relevant to the predetermined event with the log identifications in the database, and thus determine a processing manner for the log identification to be analyzed based on a solution corresponding to the target reference log identification. In addition, to make the comparison result more certain and more accurate, the comparison result of the context can also be integrated. Since the contextual log records have been converted into contextual log identifications, the contextual log identifications can be processed using NLP techniques, thus avoiding the misidentification of the log records caused by the traditional use of NLP techniques.
[0052] Figure 6A schematic block diagram of an example device 600 that can be used to implement embodiments of the present disclosure is shown. As shown, the device 600 includes a central processing unit (CPU) 601 that can perform various suitable actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or computer program instructions loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data used by the device 600, in addition to the computer program instructions, can also be stored in the RAM 603. The CPU 601, the ROM 602, and the RAM 603 are connected to each other by a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0053] Various components in the device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; the storage unit 608, such as a magnetic disk, a magneto-optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices over a computer network, such as the Internet, and / or various telecommunication networks.
[0054] The various processes and procedures described above, such as the methods 300 and / or 500, can be performed by the processing unit 601. For example, in some embodiments, the methods 300 and / or 500 can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the CPU 601, one or more actions of the methods 300 and / or 500 described above can be performed.
[0055] The present disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for performing various aspects of the present disclosure.
[0056] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0057] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0058] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0059] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0060] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0061] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0062] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0063] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are within the scope of the disclosure. The selection of terms is intended to best describe the principles of the embodiments, practical application, or technical improvements over the technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for analyzing log files, comprising: determining respective templates of a plurality of log records in a log file based at least in part on a plurality of reference templates; determining the plurality of log records as a plurality of log signatures respectively associated with the respective templates; determining, from the plurality of log signatures, a first log signature to be analyzed corresponding to a first log record of the plurality of log records associated with a predetermined event; selecting, from a plurality of reference log signatures corresponding to the plurality of reference templates, a target reference log signature having a first similarity to the first log signature to be analyzed higher than a first threshold similarity; obtaining a context associated with the first log signature, the context including one or more additional log signatures associated with one or more additional log records of the plurality of log records collected before and / or after the first log record; and diagnosing one or more system issues associated with the first log record based at least in part on: (i) analyzing a first log signature sequence including the first log signature and the one or more additional log signatures; and (ii) analyzing a predetermined diagnostic policy of the target reference log signature; wherein analyzing the first log signature sequence comprises: generating the first log signature sequence including the first log signature and the one or more additional log signatures; generating a second log signature sequence including a target reference log signature and one or more additional reference context log signatures associated with the target reference log signature; and determining a second similarity between the first log signature sequence and the second log signature sequence using one or more machine learning models.
2. The method of claim 1, further comprising: processing the first log record corresponding to the first log signature to be analyzed based at least in part on the predetermined diagnostic policy of the target reference log signature for the predetermined event.
3. The method of claim 1, wherein selecting the target reference log signature comprises: obtaining the one or more additional reference context log signatures associated with the target reference log signature; determining a comprehensive similarity based at least in part on the first similarity and the second similarity; in accordance with a determination that the comprehensive similarity is higher than a second threshold similarity, processing the first log record corresponding to the first log signature to be analyzed based at least in part on the predetermined diagnostic policy of the target reference log signature for the predetermined event.
4. The method of claim 3, wherein the one or more additional log signatures are separated from the first log signature to be analyzed by a time interval less than a threshold time interval, and the one or more additional reference context log signatures are separated from the target reference log signature by a time interval less than the threshold time interval.
5. The method of claim 1, wherein determining the respective templates of the plurality of log records comprises: obtaining the plurality of reference templates from a reference template database; and determining, in response to the first log record in the plurality of log records matching a first reference template in the plurality of reference templates, the first reference template as a template of the first log record. 6.The method of claim 5, wherein determining the plurality of log records as the plurality of log identities respectively comprises: obtaining, from the reference template database, a mapping relationship between the plurality of reference templates and the plurality of reference log identities; obtaining, based at least in part on the mapping relationship, a reference log identity corresponding to the first reference template; and determining the reference log identity corresponding to the first reference template as the first log identity of the first log record. 7.The method of claim 1, wherein determining the first log identity to be analyzed corresponding to the predetermined event comprises: determining, based at least in part on timestamp information of the log file, the first log identity to be analyzed at a time when the predetermined event occurs. 8.The method of claim 1, wherein the predetermined event comprises at least one of: a thread crash event; a user report event; and a system metric abnormal event. 9.An electronic device comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing machine executable instructions that, when executed by the at least one processing unit, cause the device to perform acts comprising: determining, based at least in part on a plurality of reference templates, respective templates of a plurality of log records in a log file; determining the plurality of log records as a plurality of log identities associated with the respective templates respectively; determining, from the plurality of log identities, a first log identity to be analyzed corresponding to a predetermined event associated with a first log record in the plurality of log records; selecting, from a plurality of reference log identities corresponding to the plurality of reference templates, a target reference log identity having a first similarity to the first log identity to be analyzed higher than a first threshold similarity; obtaining a context associated with the first log identity, the context containing one or more additional log identities associated with one or more additional log records collected before and / or after the first log record; and diagnosing one or more system issues associated with the first log record based at least in part on: (i) analyzing a first log identity sequence containing the first log identity and the one or more additional log identities; and (ii) analyzing a predetermined diagnosis policy of the target reference log identity; wherein analyzing the first log identity sequence comprises: generating the first log identity sequence containing the first log identity and the one or more additional log identities; generating a second log identity sequence containing a target reference log identity and one or more additional reference context log identities associated with the target reference log identity; and determine a second similarity between the first log identifier sequence and the second log identifier sequence using one or more machine learning models.
10. The device of claim 9, the actions further comprising: processing the first log records corresponding to the first log identifier to be analyzed based at least in part on the predetermined diagnostic policy for the predetermined event identified by the target reference log identifier.
11. The device of claim 9, wherein selecting the target reference log identifier comprises: obtaining the one or more additional reference contextual log identifiers associated with the target reference log identifier; determining a comprehensive similarity based at least in part on the first similarity and the second similarity; in accordance with a determination that the comprehensive similarity is higher than a second threshold similarity, processing the first log records corresponding to the first log identifier to be analyzed based at least in part on the predetermined diagnostic policy for the predetermined event identified by the target reference log identifier.
12. The device of claim 11, wherein the one or more additional log identifiers are separated from the first log identifier to be analyzed by a time interval that is less than a threshold time interval, and the one or more additional reference contextual log identifiers are separated from the target reference log identifier by a time interval that is less than the threshold time interval.
13. The device of claim 9, wherein determining the respective templates for the plurality of log records comprises: obtaining the plurality of reference templates from a reference template database; and in response to the first log record of the plurality of log records matching a first reference template of the plurality of reference templates, determining the first reference template as the template for the first log record.
14. The device of claim 13, wherein determining the plurality of log identifiers for the plurality of log records, respectively, comprises: obtaining a mapping relationship between the plurality of reference templates and the plurality of reference log identifiers from the reference template database; obtaining a reference log identifier corresponding to the first reference template based at least in part on the mapping relationship; and determining the reference log identifier corresponding to the first reference template as the first log identifier for the first log record.
15. The device of claim 9, wherein determining the first log identifier to be analyzed corresponding to the predetermined event comprises: determining the first log identifier to be analyzed at which the predetermined event occurs based at least in part on timestamp information of the log file.
16. The device of claim 9, wherein the predetermined event comprises at least one of: a thread crash event; a user report event; and a system metric abnormality event.
17. A computer program product tangibly stored on a non-transitory computer- readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the steps of the method of any one of claims 1-8.
Citation Information
Patent Citations
Clustering of log messages
US20190370347A1