Log statement recommendation method and device, electronic equipment and distributed system
By generating recommended log statements based on code operation scenarios in a distributed system and using neural network models for fault detection, the problem of poor log data quality is solved, and efficient fault detection and intelligent operation and maintenance are achieved.
Patent Information
- Application Number
- CN202311861564.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the log data quality of the distributed system is not ideal, and there are problems such as inconsistent statement formats, high redundant information, and a lot of random noise, which leads to high operation and maintenance costs and difficult to form an effective closed loop, affecting the effective development of fault diagnosis and intelligent operation and maintenance.
By generating recommended log statements based on the preset correspondence between code operation scenarios and log specification information in a distributed system, a neural network model is used for fault detection, and structured and unstructured information conversion of log events is realized, forming a closed loop from log statement generation to analysis.
It improves the quality of log data and fault detection efficiency, reduces dependence on professionals, adapts to different environments and code operation scenarios, reduces labor costs, and supports efficient fault location and detection.
Smart Images

Figure CN120295892A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a method and apparatus for recommending log statements, an electronic device, and a distributed system. Background Art
[0002] Log data is one of the most important basic operation and maintenance data of software systems, which records detailed runtime information during the operation of software systems in real time. Log data reflects fine-grained application states and program execution logics across components, thereby reflecting whether there are abnormalities in software systems. Therefore, tasks such as fault prediction, sub-healthy detection, anomaly detection, and performance optimization of software systems can be performed by analyzing log data.
[0003] For distributed systems, log data has the advantages of diagnosing faults at the code segment level, detecting program execution logics, and capturing fault details. At the same time, in many cases, log data is the only available fault diagnosis data source. However, log data has problems with unsatisfactory data quality, such as inconsistent statement formats, excessive redundant information, and a large amount of random noise, resulting in too high costs for traditional rule-matching operation and maintenance methods. Summary of the Invention
[0004] Embodiments of this application provide a method and apparatus for recommending log statements, an electronic device, and a distributed system, which solve the problem of too low quality of current log statements.
[0005] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, a method for recommending log statements is provided, which is applied to a node device in a distributed system. The distributed system includes multiple node devices. Each node device implements the functions corresponding to the code by running the code. The method includes: running a first code in a first running scenario, where the first running scenario is any one of multiple preset code running scenarios of the distributed system. Based on a preset correspondence between the code running scenario and log specification information, determining first log specification information corresponding to the first running scenario. The log specification information corresponding to a code running scenario includes the characteristics of historical log statements generated when any node device in the distributed system runs historical code in this code running scenario. The log specification information is used to represent the generation rules of log statements. Generating recommended log statements corresponding to the first code according to the generation rules represented by the first log specification information.
[0007] The log statement recommendation method provided by the embodiments of this application is universal, scalable, and interpretable, and can be applied to the fault location scenarios of operating environments with different environments and different product forms, such as cloud service scenarios, embedded scenarios, etc. The log statement recommendation method provided by the embodiments of this application is universal and can be applied to devices capable of printing logs. The log statement recommendation method provided by the embodiments of this application is scalable. For example, it can be applied to different distributed systems and different code running scenarios. At the same time, after the log statement recommendation method provided by this application obtains the first code running in the first running scenario of the code file, based on the preset correspondence between the running scenario of the code statement and the log specification information, it determines the first log specification information corresponding to the first running scenario, and generates the recommended log statement corresponding to the first code statement according to the generation rule of the log statement characterized by the first specification information. In this process, the log specification information can be refined based on the code running scenario, and log statements can be recommended for each code running scenario, avoiding differences in log specification information due to different code running scenarios. In this process, no professional intervention is required, and it does not rely on manual and expert experience, which can save labor costs.
[0008] In a possible implementation manner, the log statement recommendation method provided by this application further includes: obtaining the historical code run by each node device in the distributed system and the historical log statements corresponding to the historical code. Classify the code running scenarios of the historical code to obtain multiple preset code running scenarios. Respectively extract the log statement features for the historical log statements corresponding to the historical code in each code running scenario. Based on the log statement features corresponding to each code running scenario, generate the log specification information corresponding to each code running scenario to obtain the preset correspondence.
[0009] In this possible implementation manner, the master node device uses the historical code run by the slave node device and the historical log statements corresponding to the historical code to extract the log statement features for the historical log statements corresponding to the historical code in each code running scenario under each code running scenario classification directory. A correspondence can be generated based on the log statement features and the log specification information corresponding to each code running scenario. By refining the log specification information based on the code running scenario, the difference in log specification information caused by different code running scenarios can be avoided.
[0010] In a possible implementation, the historical log statements corresponding to the historical code have recommendation tags, where the recommendation tag is a positive tag or a negative tag opposite to the positive tag. The historical log statements are used to reflect the running situation of the corresponding historical code. The recommendation tag is used to characterize the accuracy of the running situation reflected by the historical log statement. For the historical log statements corresponding to the historical code in each code running scenario, log statement features are extracted, including: for the target historical log statements corresponding to the historical code in each code running scenario, log statement features are extracted. The target historical log statement is a historical log statement with a positive recommendation tag.
[0011] In this possible implementation, by annotating the historical log statements corresponding to the historical code with recommendation tags, a preset correspondence between the code running scenario of the historical log statement with a positive recommendation tag and the log specification information of the historical log statement can be generated. Based on this preset correspondence, the quality of the recommended log statements can be continuously improved. Even if the service content of the distributed system changes, recommended code statements can be generated based on the newly generated preset correspondence.
[0012] In a possible implementation, the above-mentioned log statement features include: one or more of the error category, statement format, and printing position of the log statement.
[0013] In a possible implementation, the log statement recommendation method provided in this application further includes: based on the recommended log statements corresponding to the first code, determining the first log statements corresponding to the first code and printing the first log statements.
[0014] In a possible implementation, the log statement recommendation method provided in this application further includes: extracting the semantic information of the first log statements from the first log statements corresponding to the first code. Based on the semantic information of at least one first log statement and the printing order between at least one first log statement, a first log event model is generated. The first log event model includes the structured information and unstructured information of the log events corresponding to at least one first log statement. The first log event model is input into a preset fault detection model based on a neural network. The preset fault detection model is used to perform fault detection on the running situation of the first code, and the information output by the preset fault detection model includes the fault detection result of the running situation of the first code.
[0015] In this possible implementation, by generating a first log event model based on the semantic information of at least one first log statement and the printing order among at least one first log statement, the structured or unstructured information of the log event can be automatically transformed into a unified log event model, which is convenient for subsequent use of the log event model to detect faults in the running status of the first code. For the manufacturers of node devices in different distributed systems, the log specifications and professional terms can be unified, and the key information in the log can be extracted using the unified log event model, improving the detection efficiency of fault detection for the running status of the first code and saving the time of analysts. At the same time, if the service content in the distributed system changes rapidly, for example, the keywords in the log statement change, but the semantic information of the log statement does not change, a first log event model can still be generated based on the semantic information of at least one first log statement and the printing order among at least one first log statement, and then the structured or unstructured information of the log event can be automatically transformed into a unified log event model, and the log event model can be used to detect faults in the running status of the first code.
[0016] In a possible implementation, the log statement recommendation method provided in this application further includes: determining that the recommended label of at least one first log statement is a first recommended label according to the fault detection result of the running status of the first code output by a preset fault detection model. The first recommended label is used to characterize the accuracy of the running status of the first code reflected by at least one first log statement.
[0017] In this possible implementation, the node device can determine that the recommended label of at least one first log statement is a first recommended label according to the fault detection result of the running status of the first code output by a preset fault detection model, and extract log statement features from the first log statements with a positive recommended label among the first recommended labels. Based on the log statement features corresponding to the first code running scenario with a positive recommended label, generate the log specification information corresponding to the first code running scenario with a positive recommended label, and obtain the preset correspondence between the first code running scenario with a positive recommended label and the log specification information of the first code with a positive recommended label. It can be used to generate the recommended log statements corresponding to other codes of the node device according to the generation rules characterized by the log specification information corresponding to the first code running scenario, and can form an effective closed loop from log statement generation to log data analysis, continuously improving the quality of the recommended log statements.
[0018] In a possible implementation manner, the log statement recommendation method provided by this application further includes: for each code running scenario, extracting the semantic information of historical log statements from the historical log statements corresponding to the historical code in the code running scenario. Based on the semantic information of at least one historical log statement and the printing order among at least one historical log statement in the code running scenario, a log event model is generated. The log event model includes the structured information and unstructured information of the log events corresponding to at least one historical log statement. According to the log event models in various code running scenarios, a preset fault detection model is obtained by training a neural network-based fault detection model.
[0019] In the embodiments of this application, a node device can extract the semantic information of a first log statement from the first log statement corresponding to the first code, generate a first log event model based on the semantic information of at least one first log statement and the printing order among at least one first log statement, and input the first log event model into a preset fault detection model based on a neural network to obtain a fault detection result of the running situation of the first code, so as to more efficiently detect faults during the code running process and locate the position and generation time of the faults. At the same time, the log statement recommendation method provided by the embodiments of this application is interpretable. By using the preset fault detection model to match the historical log event model with the first event model to perform fault detection on the running situation of the first code, developers can better understand the fault detection process of the preset fault detection model for the first event model.
[0020] In a second aspect, the embodiments of this application provide a log statement recommendation device, which is applied to a node device in a distributed system. The distributed system includes multiple node devices. Each node device realizes the function corresponding to the code by running the code. The device includes: a running module, a determining module, and a generating module.
[0021] Among them, the running module is used to run the first code in the first running scenario, and the first running scenario is any one of multiple preset code running scenarios of the distributed system.
[0022] The determining module is used to determine the first log specification information corresponding to the first running scenario based on a preset correspondence between the code running scenario and the log specification information. The log specification information corresponding to a code running scenario includes the characteristics of historical log statements generated when any node device in the distributed system runs historical code in this code running scenario. The log specification information is used to characterize the generation rule of log statements.
[0023] The generating module is used to generate recommended log statements corresponding to the first code according to the generation rule characterized by the first log specification information.
[0024] In a possible implementation, the log statement recommendation device provided by the embodiments of the present application further includes: an acquisition module, a classification module, and an extraction module.
[0025] Among them, the acquisition module is used to acquire the historical code run by each node device in the distributed system and the historical log statements corresponding to the historical code.
[0026] The classification module is used to classify the code running scenarios of the historical code to obtain a variety of preset code running scenarios.
[0027] The extraction module is used to extract the log statement features for the historical log statements corresponding to the historical code under each code running scenario respectively.
[0028] The generation module is further used to generate the log specification information corresponding to each code running scenario based on the log statement features corresponding to each code running scenario, and obtain a preset corresponding relationship.
[0029] In a possible implementation, the historical log statements corresponding to the historical code have recommendation tags, and the recommendation tags are positive tags or negative tags opposite to the positive tags. The historical log statements are used to reflect the running situation of the corresponding historical code. The recommendation tags are used to characterize the accuracy of the running situation reflected by the historical log statements.
[0030] Specifically, the extraction module is used to extract the log statement features for the target historical log statements corresponding to the historical code under each code running scenario respectively. The target historical log statements are the historical log statements with positive recommendation tags.
[0031] In a possible implementation, the log statement features include one or more of the error category, statement format, and printing position of the log statement.
[0032] In a possible implementation, the determination module is further used to determine the first log statement corresponding to the first code based on the recommended log statement corresponding to the first code, and print the first log statement.
[0033] In a possible implementation, the extraction module is further used to extract the semantic information of the first log statement from the first log statement corresponding to the first code.
[0034] The generation module is further used to generate a first log event model based on the semantic information of at least one first log statement and the printing order between at least one first log statement. The first log event model includes the structured information and unstructured information of the log events corresponding to at least one first log statement.
[0035] The log statement recommendation device provided by the embodiments of the present application further includes:
[0036] An input module for inputting a first log event model into a preset fault detection model based on a neural network, where the preset fault detection model is used to detect faults in the running condition of a first piece of code, and the information output by the preset fault detection model includes the fault detection result of the running condition of the first piece of code.
[0037] In a possible implementation, the determination module is further configured to determine that the recommended label of at least one first log statement is a first recommended label according to the fault detection result of the running condition of the first piece of code output by the preset fault detection model. The first recommended label is used to characterize the accuracy of the running condition of the first piece of code reflected by at least one first log statement.
[0038] In a possible implementation, the extraction module is further configured to, for each code running scenario, extract the semantic information of the historical log statements from the historical log statements corresponding to the historical code in the code running scenario.
[0039] The generation module is further configured to generate a log event model based on the semantic information of at least one historical log statement and the printing order between at least one historical log statement in the code running scenario. The log event model includes the structured information and unstructured information of the log events corresponding to at least one historical log statement.
[0040] The log statement recommendation device provided in the embodiments of the present application further includes:
[0041] A training module for training a fault detection model based on a neural network according to the log event models in various code running scenarios to obtain a preset fault detection model.
[0042] In a third aspect, the embodiments of the present application provide an electronic device, which has the function of implementing the log statement recommendation method in the first aspect or any one of the first aspects. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0043] In a fourth aspect, the embodiments of the present application provide a distributed system, which includes at least one node device. The node device may include an electronic device as in the third aspect. Each node device includes a processor and a memory. The processor of at least one node device is configured to execute instructions stored in the memory of at least one node device, so that the distributed system executes the log statement recommendation method in the first aspect and any one of the first aspects.
[0044] In a fifth aspect, there is provided a computer program product containing instructions, which, when run by a distributed system, enables the distributed system to execute the log statement recommendation method in the first aspect and any one of the first aspects.
[0045] In a sixth aspect, a computer-readable storage medium is provided, in which computer program instructions are stored. When executed in a distributed system, the distributed system can execute the log statement recommendation method in the first aspect and any one of the items in the first aspect as described above.
[0046] Among them, for the technical effects brought by any one of the design manners in the second aspect to the sixth aspect, reference may be made to the technical effects brought by different design manners in the first aspect, which will not be elaborated herein. Description of the Drawings
[0047] Figure 1 A schematic diagram of a log exception type provided by an embodiment of the present application;
[0048] Figure 2 A schematic diagram of a source of log data provided by an embodiment of the present application;
[0049] Figure 3 A schematic diagram of a log exception type provided by an embodiment of the present application;
[0050] Figure 4 A schematic flowchart of a log intelligent analysis method provided by the related art;
[0051] Figure 5 A schematic diagram of a structure of a distributed system provided by an embodiment of the present application;
[0052] Figure 6 A schematic diagram of a structure of a node device provided by an embodiment of the present application;
[0053] Figure 7 A schematic flowchart of a log statement recommendation method provided by an embodiment of the present application;
[0054] Figure 8 Another schematic flowchart of a log statement recommendation method provided by an embodiment of the present application;
[0055] Figure 9 Another schematic flowchart of a log statement recommendation method provided by an embodiment of the present application;
[0056] Figure 10 Another schematic flowchart of a log statement recommendation method provided by an embodiment of the present application;
[0057] Figure 11 A schematic diagram of a structure of a log statement recommendation device provided by an embodiment of the present application;
[0058] Figure 12 Another schematic diagram of a structure of a log statement recommendation device provided by an embodiment of the present application;
[0059] Figure 13 Another structural schematic diagram of a node device provided by an embodiment of the present application;
[0060] Figure 14 A structural schematic diagram of a distributed system provided by an embodiment of the present application;
[0061] Figure 15 Another structural schematic diagram of a distributed system provided by an embodiment of the present application. Detailed implementation manners
[0062] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; "and / or" in the present application is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A and B may be singular or plural. And, in the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression below refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, c may be single or multiple. In addition, in order to clearly describe the technical solutions in the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit to be different. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner for easy understanding.
[0063] In addition, the network architecture and service scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions in the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the network architecture and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0064] For ease of understanding, relevant technical terms involved in this application are first explained.
[0065] The log specification refers to recording logs according to a preset standard.
[0066] A distributed system refers to a system composed of multiple independent node devices that are interconnected through a network and cooperate to complete a series of tasks. The biggest feature of a distributed system is its high scalability and reliability. New node devices can be easily added to the system to improve the system's performance and reliability.
[0067] Log data is usually used to record and provide key information during the operation of a software system to support various software operation and maintenance tasks, such as debugging, defect prediction, monitoring, and auditing. In addition to log data, other operation and maintenance data such as metric data and link data during the operation of a software system can also be used to support various software operation and maintenance tasks. Among them, metric data can visually present data in the form of numbers during the software operation and maintenance process, and link data can quickly locate problems during the software operation and maintenance process through the link. However, compared with other operation and maintenance data, log data can support more fine-grained root cause diagnosis of software system failures and can support the monitoring of the program execution logic of the software system, capturing program execution exceptions and performance exceptions across components and services.
[0068] Log data comes from log statements injected into the code. Log statements usually consist of a timestamp, a log level, log content, etc. The timestamp represents the exact time when the event corresponding to the log occurred and is used to record the order of event occurrence. The log level represents the severity or priority of the event, and common log levels include debugging, information, warning, error, and critical error, etc. The log content includes descriptive text about the event and is used to explain the detailed information of the event.
[0069] Intelligent operations (artificial intelligence for it operations, AIOps) refers to the application of artificial intelligence (AI) to the operation and maintenance field of software systems. Based on existing operation and maintenance data, problems that cannot be solved by traditional operation and maintenance are solved through AI. In the current application of intelligent operations, intelligent operations are mainly realized based on the above-mentioned metric data and link data. When AIOps based on metric data and link information is gradually reaching its bottleneck, log data has become an important breakthrough for improving the effect of intelligent operations.
[0070] Such as Figure 1As shown, the following log exception types can be used to detect exceptions in log statements in log data. The log exception types can include keywords, template count, template sequence, continuous variable (variable value), discrete variable (variable distribution), time interval, etc. Specifically, by monitoring the keywords in the log statement that characterize the error category, it is monitored whether the system has an exception. The keywords can include debugging, information, warning, error, and critical error, etc. By monitoring the change in the template count corresponding to the log statement, it is monitored whether the system has an exception. For example, when the template count suddenly increases or suddenly decreases, it is determined that there is a problem in the system when executing the code. By monitoring the template sequence corresponding to the log statement, it is monitored whether the system has an exception. For example, if the normal template sequence is ABC, and when it is monitored that the template sequence becomes ABD, it indicates that the system has an exception at this time. The continuous variable is some key metric values in certain logs, such as the response time. By extracting such variables as metrics for monitoring, when it is monitored that the metric value suddenly changes, it is determined that the system has an exception. The discrete variable is some discrete key metric values in certain logs, and these variables may have distributional exceptions. For example, a sudden increase in the proportion of 404 return codes indicates that the system has an exception. The log exception monitoring based on the time interval can be that if the number of logs generated by the system within a certain period of time significantly increases or significantly decreases (or the time interval becomes longer or shorter), it indicates that the system has an exception.
[0071] In the application scenario of realizing intelligent operation and maintenance based on log data, the quality of log data has a great impact on intelligent operation and maintenance. Here, the quality of log data includes but is not limited to the accuracy of the key information reflected by the log data during the operation of the software system, the proportion of valid data in the log data, and the consistency of the specifications of log data from different sources, etc. However, there are still several problems that need to be urgently solved in log data generation and log data analysis.
[0072] In a first aspect, the current method of generating log data results in unsatisfactory quality of log data in large-scale distributed systems. Specifically, taking large-scale distributed systems as an example, the injection of log statements mainly depends on preset log specifications and professional terms, as well as the professional capabilities, preferences, and development experience of developers. Specifically, node device manufacturers in different distributed systems have different log specifications and professional terms, and the log specifications are not unified. For example, the log specification of manufacturer A is that node devices generate logs in the order of log generation time, log level, and log content, while the log specification of manufacturer B is that node devices record logs in the order of log level, log content, and log time. In addition, there is a significant inconsistency between the log intentions of developers and the actual log statements generated in the code, resulting in a large amount of heterogeneous and useless information in most current log data. This phenomenon is more prominent in the development of distributed microservices in the DevOps mode, seriously affecting the subsequent analysis efficiency and effect of log data. In addition, from a large number of current practices, there is an obvious disconnection between log data generation and log data analysis. An effective closed loop cannot be formed from log statement generation to the final log data analysis, affecting the quality of log data and hindering the effective implementation of intelligent operation and maintenance. Therefore, although log data is considered to have advantages in capturing the dynamic behavior of programs, its effective application in AIOps is quite limited.
[0073] As Figure 2 shown, the sources of log data can include electronic devices in various communication systems, such as servers, terminal devices, gateway devices, and storage devices, which are used to reflect the key information during the operation of software operating systems, software modules, software services, and databases installed on various electronic devices. Among them, the storage devices include, but are not limited to, storage devices in a distributed file system (ceph). The gateway device can be a switch, a gateway, etc. The operating system can be Linux, Windows, etc. The software module can be MQ, Redis, etc. The software service can be Weblogic, WAS, etc. The database can be Oracle, DB2, etc.
[0074] There will be invalid information in the log data generated from the above sources depending on the professional capabilities, preferences, and development experience of developers. The invalid information can include duplicate information, useless information, etc. As Figure 3As shown, taking the log exception type as the keyword as an example, the repeated information in the log statement can be the redundant logs repeatedly generated by the system through different sequences due to the exception of the template sequence (templatesequence) in the log statement (for example, 09:59:42 INFO Receiving block). Each template is a set of words in a semi-structured text, and the template sequence is the sequence generated based on the template. The useless information in the log statement can be the noisy logs generated due to log changes in the log statement (for example, 09:59:38 INFO Receiving block from ip, 09:59:42 INFO Receiving block from ip). There may also be fault information in the log statement that can cause potential performance failures in the node devices in the distributed system (for example, 09:38:49 WARN Redundant addStoredBlock request, indicating that there is a redundant data node addition request in the computing device at this time, and the log level is WARN, which may cause potential performance failures).
[0075] In the second aspect, there are many defects in the current log data analysis method. In some related technologies, a log analysis method based on keywords and regular expressions is adopted to identify the abnormal logs in the logs for system troubleshooting. Taking the distributed system as an example, this method has the following disadvantages:
[0076] 1. Node device manufacturers in different distributed systems have different log specifications and professional terms. The log specifications are not unified, the log patterns are diverse, and the professional skills requirements for professionals are high. For example, the log specification of manufacturer A is that the node device generates logs in the order of the time when the log is generated, the log level, and the log content, and the log specification of manufacturer B is that the node device records logs in the order of the log level, the log content, and the log time. At this time, the node device can use the method based on keywords and regular expressions to identify the abnormal logs in the logs of manufacturer A. However, due to the different log specifications of manufacturer B and manufacturer A, the original method based on keywords and regular expressions cannot be directly used to identify the abnormal logs in the logs of manufacturer B, and professionals need to re-adjust the regular expressions, which requires relatively high professional skills for professionals.
[0077] 2. There are semi-structured or unstructured log statements in the log statements, and the rules are not obvious, resulting in it being difficult to extract the key information of the log statements through the preset regular expressions.
[0078] 3. If the service content of the distributed system changes rapidly, it may cause the log analysis rules set manually to become invalid. For example, the log analysis rule is that when the keyword a appears in the detected log statement, it is determined as type A exception. When the service content changes, the keyword a changes to keyword b, and at this time, the keyword b cannot be detected, resulting in the invalidation of the log analysis rule.
[0079] Figure 4 A log intelligent analysis method provided for some other related technologies, such as Figure 4 shown in the figure, this method defines each subsystem in the complex system as a managed object through the configuration and management module, the log files of each managed object, as well as the format of the log records and the error keywords, establishes a log correlation group, and at the same time designates the log file to be monitored as the trigger point; through the log record time calibration module, the date and time of each managed object are synchronized periodically based on the time reference; in the same log correlation group, when an error keyword appears, the log information collection module collects the log records generated in each managed object during the current period after the reference; through the log analysis output module, an analysis request is created, the collected log records are searched, sorted and analyzed, and a report for visual linkage display is generated. In this way, the logs of each subsystem in the complex distributed system can be associated to achieve intelligent linkage and improve the efficiency of problem troubleshooting. However, this method has the following disadvantages:
[0080] 1. This method will face the problem of generality. For example, for different distributed systems, especially medium and large-scale systems, the log types and styles are diverse. Before applying this method, it is necessary to perform adaptation operations on different distributed systems, such as adapting to the log format and error keywords, and the adaptation workload is large.
[0081] 2. This method only troubleshoots problems based on prior experience, and it is difficult to summarize experience and judge unknown problems.
[0082] Based on this, the present application provides a log statement recommendation method, and its basic principle is: running the first code in the first running scenario, and the first running scenario is any one of multiple preset code running scenarios of the distributed system. Based on the preset relationship between the code running scenario and the log specification information, the first log specification information corresponding to the first running scenario is determined. According to the generation rule of the log statement represented by the first specification information, the recommended log statement corresponding to the first code is generated.
[0083] The log statement recommendation method provided by the embodiments of this application can be applicable to different scenarios, such as different communication systems, etc. After running the first code in the first running scenario, based on the preset correspondence between the code running scenario and the log specification information, the first log specification information corresponding to the first running scenario is determined, and the recommended log statements corresponding to the first code are generated according to the generation rules of the log statements characterized by the first specification information. In this process, the log specification information can be refined based on the code running scenario, and log statements can be recommended for each code running scenario, avoiding differences in log specification information due to different code running scenarios.
[0084] The log statement recommendation method provided by the embodiments of this application can be applied to a communication system, which can be a distributed system, and a microservice architecture in the DevOps mode can be deployed on the distributed system. For example, Figure 5 As shown, the distributed system 500 can include a cloud service platform 501 and an infrastructure that provides multiple services. The infrastructure can include multiple cloud service centers 502. Each cloud service center 502 can include multiple node devices 503, and the node devices 503 can communicate with each other. The node devices 503 can communicate with the cloud service platform 501. Among them, the node device 503 can be a physical machine, or a virtual machine or a container, etc. The embodiments of this application do not limit the type of the node device 503. The node device 503 can be, for example but not limited to, a server, a web server (such as Nginx, Apache), an application server (such as Weblogic, Was), a switch, etc. It can be understood that Figure 5 The shown distributed system is only an example. In actual implementation, the multiple node devices 503 can also directly communicate with each other, or communicate through other devices. This application does not limit this.
[0085] The cloud service platform 501 can be used to provide access interfaces (such as an interface or an application program interface (API)). The user can operate the client 504 to remotely access the access interface and log in to the cloud service platform 501. After the cloud service platform 501 successfully authenticates the user, the user can further pay on the cloud service platform 501 to select and rent cloud services of a specific specification (processor, memory, disk). After the successful payment purchase, the cloud service platform 501 provides the cloud services rented by the user by scheduling the computing resources of one or more node devices 503 in one or more cloud service centers 502. The users of the cloud services can be individuals, enterprises, schools, hospitals, administrative agencies, etc.
[0086] Among the multiple node devices 503 in the cloud service center 502, a main node device and multiple slave node devices can be distinguished. The main node device can be used for monitoring and managing the slave node devices, etc. For example, the main node device can monitor the working status of each slave node device in the distributed system. The main node device can also be used to send data acquisition requests to the slave node devices. For example, the main node device sends a log data acquisition request to the slave node device to request the acquisition of the log data on the slave node device. After receiving the log data on the slave node device, the main node device quickly locates or anticipates in advance the faults occurring in the distributed system 500 according to the log statements in the log data.
[0087] The node device 503 is used to execute computing tasks and store data. The node device 503 generates log statements when executing computing tasks to record the key information of the node device 503 during operation.
[0088] In some embodiments, the log statement recommendation method provided by the embodiments of the present application can be executed by the node device 503. As Figure 6 shown, the node device 503 includes: a processor 601 and a memory 602.
[0089] Among them, the processor 601 can be a central processing unit (CPU), and the processor 601 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.
[0090] The memory 602 can be a volatile memory, such as a random-access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or a combination of the above types of memories, for storing application programs, configuration files, data information or other content that can implement the method of the present application.
[0091] The processor 601 performs the following functions by running or executing software programs and / or modules stored in the memory 602, and by invoking data stored in the memory 602:
[0092] The processor 601 runs the first code in the first operating scenario. Based on a preset correspondence between the code operating scenario and the log specification information, the first log specification information corresponding to the first operating scenario is determined, and a recommended log statement corresponding to the first code is generated according to the generation rule characterized by the first log specification information.
[0093] In some embodiments, as Figure 6 shown, the node device 503 provided in the embodiments of the present application may further include a display device 603 and an input device 604.
[0094] Among them, a communication connection may be established between the display device 603 and the processor 601. A communication connection may be established between the input device 604 and the processor 601.
[0095] The display device 603 may be used to display images, texts, videos, etc. of the computing device, and the user may view the software interface, files, web pages, etc. on the computing device 503 through the display device 603. For example, the display device 603 may be used to display the recommended log statement to the user after receiving the recommended log statement sent by the processor 601.
[0096] The input device 604 may be used to provide an interaction function for the user. The user may input data and information to the processor 601 in the computing device 503 through the input device 604. For example, the user may label a recommended tag for the log statement generated by the processor 601 to the processor 601 of the node device 503 through the input device 604, and the recommended tag is used to characterize the accuracy of the running condition of the code reflected by the log statement.
[0097] Figure 7 It is a schematic flowchart of a method for recommending log statements provided in the embodiments of the present application, and this method can be applied to Figure 6 the node device shown. As Figure 7 shown, this method may include the following steps:
[0098] S701, run the first code in the first operating scenario.
[0099] Among them, the first operating scenario is any one of multiple preset code operating scenarios of the distributed system.
[0100] In a possible implementation, the master node device can obtain the code running scenarios of each sub-node device in the distributed system. The master node device uses machine learning methods to generalize the code running scenarios with common attributes from multiple code running scenarios of the distributed system to form a classification directory of code running scenarios for the distributed system. This classification directory of code running scenarios can be a classification directory of multiple preset code running scenarios for the distributed system. The common attributes can be hardware failures, service overloads, etc. For example, the first running scenario can be scenarios such as cloud service hardware failure scenarios, cloud service overload scenarios, etc.
[0101] S702. Based on the preset correspondence between the code running scenario and the log specification information, determine the first log specification information corresponding to the first running scenario.
[0102] Among them, the log specification information corresponding to a code running scenario includes the characteristics of the historical log statements generated when any node device in the distributed system runs the historical code in this code running scenario. The log specification information is used to characterize the generation rules of the log statements.
[0103] The above-mentioned log statement characteristics may include one or more of the error category, statement format, and printing position of the log statement. For example, the error category of the log statement can be the log level, including debug, information, warning, error, and critical error, etc. The statement format can be the format of the log statement generated in the order of the log content. For example, the log statement is generated in the order of timestamp, log level, and message level.
[0104] In a possible implementation, as Figure 8 shown, the master node device can obtain the historical code run by each slave node device in the distributed system and the historical log statements corresponding to the historical code. The master node device extracts the log statement characteristics for the historical log statements corresponding to the historical code in each code running scenario in the code running scenario classification directory. Based on the log statement characteristics corresponding to each code running scenario, the master node device generates the log specification information corresponding to each code running scenario to obtain the preset correspondence.
[0105] In this possible implementation, the master node device uses the historical code run by the slave node device and the historical log statements corresponding to the historical code to extract the log statement characteristics for the historical log statements corresponding to the historical code in each code running scenario in the code running scenario classification directory. It can obtain the log specification information corresponding to each code running scenario based on the log statement characteristics corresponding to each code running scenario, so as to refine the log specification information based on the code running scenario, and avoid the difference in the log specification information caused by the difference in the code running scenario.
[0106] In a possible implementation, the historical log statements corresponding to the historical code have recommendation tags, where the recommendation tag is a positive tag or a negative tag opposite to the positive tag. The historical log statements are used to reflect the running situation of the corresponding historical code. The recommendation tag is used to characterize the accuracy of the running situation reflected by the historical log statement. For the target historical log statements corresponding to the historical code in each code running scenario, log statement features are extracted. The target historical log statement is a historical log statement with a positive recommendation tag.
[0107] Exemplarily, if the historical log statement can accurately reflect the running situation of the historical code, the recommendation tag of this historical log statement is a positive tag. If the historical log statement cannot accurately reflect the running situation of the historical code and needs to be modified before use, the recommendation tag of this historical log statement is a negative tag.
[0108] In this possible implementation, by annotating the recommendation tags for the historical log statements corresponding to the historical code, log specification information can be obtained based on the historical log statements with positive recommendation tags, and a preset correspondence between the code running scenario and the log specification information can be generated. Based on this preset correspondence, the quality of the recommended log statements can be continuously improved. Even if the service content of the distributed system changes, recommended code statements can be generated based on the newly generated preset correspondence.
[0109] S703, generate the recommended log statements for the first code according to the generation rule characterized by the first log specification information.
[0110] Exemplarily, the following is the code for the node device to run:
[0111]
[0112] After the node device executes the code "except KeyError as error:", it fails to obtain the latest data. At this time, the node device generates the corresponding recommended log statement "logger.error(\"try getting latest data,encountered%s\",error)" according to the error category and statement format corresponding to this code running scenario to reflect that the computing device fails to obtain the latest data at this time. After the node device executes the code "if len(new_data)<window+1:", the node device generates the corresponding recommended log statement "logger.info(\"not enough new data for detection\")" according to the error category and statement format corresponding to this code running scenario to indicate that there is not enough data for detection at this time. The generated recommended log statement can be printed after the code is executed.
[0113] Further, based on the recommended log statement corresponding to the first code, the node device determines the first log statement corresponding to the first code and prints the first log statement. The first log statement may be the recommended log statement or the first log statement generated after the user modifies the recommended log statement.
[0114] Among them, the printing position of the first log statement is determined according to the code in the context of the first code. The node device can print the first log statement corresponding to the first code statement at the printing position in the first log specification information.
[0115] Exemplarily, after the node device executes the code "except KeyError as error:", if it fails to obtain the latest data, the node device can, based on this code, print the log statement "logger.error(\"try getting latestdata,encountered%s\",error)" after the code "except KeyError as error:" to indicate that the node device fails to obtain the latest data at this time. After the node device executes the code "if len(new_data)<window+1:", it determines that there is not enough data for detection. The node device can, based on this code, print the log statement "logger.info(\"not enough new data for detection\")" after "if len(new_data)<window+1:" to print the first log statement according to the printing position.
[0116]
[0117] In the embodiment of the present application, the node device runs the first code in the first running scenario, determines the first log specification information corresponding to the first running scenario based on the preset correspondence between the code running scenario and the log specification information, and generates the recommended log statement corresponding to the first code according to the generation rule characterized by the first log specification information.
[0118] The log statement recommendation method provided in the embodiment of the present application is universal and can be applied to devices capable of printing logs. After the log statement recommendation method provided in the present application runs the first code in the first running scenario, it determines the first log specification information corresponding to the first running scenario based on the preset relationship between the code running scenario and the log specification information, and generates the recommended log statement corresponding to the first code according to the generation rule of the log statement characterized by the first specification information. In this process, the log specification information can be refined based on the code running scenario, and log statements can be recommended for each code running scenario, avoiding differences in log specification information due to different code running scenarios.
[0119] In a possible implementation, as Figure 9 shown, the method provided by the embodiments of the present application may further include S704 - S706. Exemplarily, S704 - S706 may be executed after S701.
[0120] S704, extract the semantic information of the first log statement from the first log statement corresponding to the first code.
[0121] For example, the node device extracts the semantic information of the first log statement as "open the file" from the first log statement corresponding to the first code.
[0122] S705, generate a first log event model based on the semantic information of at least one first log statement and the printing order between at least one first log statement.
[0123] Among them, the first log event model includes the structured information and unstructured information of the log events corresponding to at least one first log statement. The structured information may be the semantic order of the log statements. The unstructured information may be the semantic information of the log statements.
[0124] Exemplarily, after the node device executes S704 multiple times, it extracts the semantic information of log statement a as "open the file" from the log statement a corresponding to the first code. It extracts the semantic information of log statement b as "operate the file" from the log statement b corresponding to the first code. It extracts the semantic information of log statement c as "fail to operate the file" from the log statement c corresponding to the first code. The semantic information of log statement a, the semantic information of log statement b, and log statement c have a printing order, and the order is "open the file", "operate the file", "fail to operate the file". For example, the embodiments of the present application use the form of "->" to represent this order. The obtained first log event model is "open the file" -> "operate the file" -> "fail to operate the file".
[0125] In the embodiments of the present application, a first log event model is generated based on the semantic information of at least one first log statement and the printing order among at least one first log statement. The information of the log event can be automatically saved in the log event model with structured information and unstructured information, which is convenient for subsequent use of the log event model to detect faults in the running situation of the first code. For different node devices in a distributed system, the log specification information can be unified, and the key information in the log can be extracted using the unified log event model, improving the detection efficiency of detecting faults in the running situation of the first code and saving the time of analysts.
[0126] Meanwhile, if the service content in the distributed system changes rapidly, for example, the keywords in the log statement change, but the semantic information of the log statement does not change. At this time, a corresponding log event model can still be generated based on the semantic information of at least one log statement and the printing order among at least one log statement, and then the information of the log event is automatically saved in the log event model with structured information or unstructured information, and the log event model is used to detect faults in the running situation of the code.
[0127] S706, input the first log event model into a preset fault detection model based on a neural network.
[0128] Among them, the preset fault detection model is used to detect faults in the running situation of the first code, and the information output by the preset fault detection model includes the fault detection result of the running situation of the first code.
[0129] In a possible implementation manner, as Figure 10 shown, for each code running scenario, extract the semantic information of the historical log statements from the historical log statements corresponding to the historical code in the code running scenario. Based on the semantic information of at least one historical log statement and the printing order among at least one historical log statement in the code running scenario, generate a log event model, and the log event model includes the structured information and unstructured information of the log events corresponding to at least one historical log statement. According to the log event models in various code running scenarios, train a fault detection model based on a neural network to obtain a preset fault detection model.
[0130] Specifically, input the first log event model into a preset fault detection model based on a neural network, and match the first log event model with the historical event model corresponding to the first log event model to obtain the fault detection result of the running situation of the first code.
[0131] Exemplarily, the historical event model corresponding to the first log event model can be to extract the semantic information of the historical log statements from the historical log statements corresponding to the historical code as "open the file", "operate the file", and "close the file". Based on the semantic information of at least one historical log statement and the printing order between at least one historical log statement, where the printing order is to print in the order of "open the file", "operate the file", "close the file", a historical log event model is generated as "open the file" -> "operate the file" -> "close the file". The node device uses a preset fault detection model to match the historical log event model with the first event model to perform fault detection on the running situation of the first code, and determines that at this time the first event model does not match the historical log event model, that is, "fail to operate the file" is different from "close the file", and it can be determined that there is a problem with the steps of the node device for file operation, and it can be determined that there is a fault in the code corresponding to the file operation during operation.
[0132] The node device can also obtain the printing time corresponding to the recommended log statement and print the printing time so that the user can determine the fault generation time when there is a fault in the code corresponding to the file operation of the node device according to the printing time.
[0133] In the embodiment of the present application, the node device can extract the semantic information of the first log statement from the first log statement corresponding to the first code, and generate a first log event model based on the semantic information of at least one first log statement and the printing order between at least one first log statement. The first log event model is input into a preset fault detection model based on a neural network to obtain a fault detection result of the running situation of the first code, which can more efficiently detect faults during the code running process and locate the position and generation time of the faults.
[0134] In addition, the log statement recommendation method provided in the embodiment of the present application is interpretable. Using the preset fault detection model to match the historical log event model with the first event model to perform fault detection on the running situation of the first code can enable developers to better understand the fault detection process of the preset fault detection model for the first event model.
[0135] In a possible implementation manner, as Figure 9 shown, the method provided in the embodiment of the present application may further include S707. Exemplarily, S707 can be executed after S706.
[0136] S707, determine that the recommended label for at least one first log statement is the first recommended label according to the fault detection result of the operation of the first code output by the preset fault detection model.
[0137] Among them, the first recommended label is used to characterize the accuracy of the operation of the first code reflected by at least one first log statement.
[0138] In the embodiments of the present application, the node device can determine whether the recommended label for at least one first log statement is a positive label or a negative apology according to the fault detection result of the operation of the first code output by the preset fault detection model. If the recommended label for the first log statement corresponding to the first code is a positive label, extract the log statement features from the first log statement, and optimize the log specification information corresponding to the first operation scenario based on the log statement features extracted from the first log statement. For each code operation scenario, generate the log specification information corresponding to each code operation scenario based on the log statement features of the log statements with a positive recommended label, obtain the preset correspondence between the code operation scenario and the log specification information, and form an effective closed loop from log statement generation to log data analysis, which can continuously improve the quality of the recommended log statements.
[0139] The embodiments of the present application can divide the functional modules of the node device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0140] In the case of dividing each functional module corresponding to each function, Figure 11 shows a possible composition diagram of the log statement recommendation device involved in the above and embodiments. The log statement recommendation device 1100 is deployed on the node device. As Figure 11 shown, the log statement recommendation device 1100 may include: an operation module 111, a determination module 112, and a generation module 113.
[0141] Among them, the operation module 111 is used to support the log statement recommendation device 1300 to execute Figure 7 the S701 shown in the log statement recommendation method shown in Figure 9 or the S701 shown in
[0142] The determination module 112 is used to support the log statement recommendation device 1300 to execute Figure 7 the S702 shown in the log statement recommendation method shown in Figure 9The S707 shown.
[0143] A generation module 113, configured to support the log statement recommendation device 1300 to execute Figure 7 S703 in the log statement recommendation method shown or Figure 9 The S705 shown.
[0144] Furthermore, as Figure 12 shown, the log statement recommendation device 1100 provided by the embodiment of the present application may further include: an acquisition module 114, a classification module 115, an extraction module 116, an input module 117, and a training module 118.
[0145] The acquisition module 114 is configured to support the log statement recommendation device 1300 to execute the step of acquiring the historical code run by each node device in the distributed system and the historical log statements corresponding to the historical code in the log statement recommendation method.
[0146] The classification module 115 is configured to support the log statement recommendation device 1300 to execute the step of classifying the code running scenarios of the historical code to obtain multiple preset code running scenarios in the log statement recommendation method.
[0147] The extraction module 116 is configured to support the log statement recommendation device 1300 to execute Figure 9 S704 in the log statement recommendation method shown or the step of extracting log statement features for the historical log statements corresponding to the historical code under each code running scenario respectively.
[0148] The input module 117 is configured to support the log statement recommendation device 1300 to execute Figure 9 S705 in the log statement recommendation method shown.
[0149] The training module 118 is configured to support the log statement recommendation device 1300 to execute the step of training a neural network-based fault detection model according to the log event models under various code running scenarios to obtain a preset fault detection model in the log statement recommendation method.
[0150] It should be noted that all relevant contents of the steps involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.
[0151] Among them, the running module 111, the determining module 112, the generating module 113, the obtaining module 114, the classifying module 115, the extracting module 116, the input module 117, and the training module 118 can all be implemented by software or by hardware. Exemplarily, next, taking the running module 111 as an example, the implementation manner of the running module 111 will be introduced. Similarly, the implementation manners of the determining module 112, the generating module 113, the obtaining module 114, the classifying module 115, the extracting module 116, the input module 117, and the training module 118 can refer to the implementation manner of the running module 111.
[0152] As an example of a software functional unit, the running module 111 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (node device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the running module 111 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.
[0153] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0154] As an example of a hardware functional unit, the operation module 111 may include at least one node device, such as a server. Alternatively, the operation module 111 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0155] The multiple node devices included in the operation module 111 may be distributed in the same region or in different regions. The multiple node devices included in the operation module 111 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple node devices included in the operation module 111 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple node devices may be any combination of node devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0156] It should be noted that in other embodiments, the operation module 111 may be used to execute any step in the log statement recommendation method, the determination module 112 may be used to execute any step in the log statement recommendation method, the generation module 113 may be used to execute any step in the log statement recommendation method, the acquisition module 114 may be used to execute any step in the log statement recommendation method, the classification module 115 may be used to execute any step in the log statement recommendation method, the extraction module 116 may be used to execute any step in the log statement recommendation method, the input module 117 may be used to execute any step in the log statement recommendation method, and the training module 118 may be used to execute any step in the log statement recommendation method. The steps responsible for implementation by the operation module 111, the determination module 112, the generation module 113, the acquisition module 114, the classification module 115, the extraction module 116, the input module 117, and the training module 118 may be specified as needed. The entire function of the log statement recommendation device is realized by the operation module 111, the determination module 112, the generation module 113, the acquisition module 114, the classification module 115, the extraction module 116, the input module 117, and the training module 118 respectively implementing different steps in the log statement recommendation method.
[0157] The log statement recommendation device 1100 provided by the embodiments of the present application is used to execute the above-mentioned log statement recommendation method, and thus can achieve the same effect as the above-mentioned log statement recommendation method.
[0158] The present application further provides an electronic device, which may include a node device 1300. As Figure 13 shown, the node device 1300 includes: a bus 1301, a processor 1302, a memory 1303, and a communication interface 1304. The processor 1302, the memory 1303, and the communication interface 1304 communicate with each other through the bus 1301. The node device 100 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the node device 1300.
[0159] The bus 1301 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 13 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 1302 may include a path for transmitting information between various components of the node device 1300 (for example, the memory 1303, the processor 1302, and the communication interface 1304).
[0160] The processor 1302 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0161] The memory 1303 may include a volatile memory, such as a random access memory (RAM). The processor 1302 may further include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0162] The memory 1303 stores executable program codes, and the processor 1302 executes the executable program codes to respectively implement the functions of the aforementioned running module 111, determining module 112, generating module 113, obtaining module 114, classifying module 115, extracting module 116, input module 117, and training module 118, so as to implement the log statement recommendation method. That is, the memory 1303 stores instructions for executing the log statement recommendation method.
[0163] Alternatively, the memory 1303 stores executable codes, and the processor 1302 executes the executable codes to respectively implement the functions of the aforementioned log statement recommendation device, so as to implement the log statement recommendation method. That is, the memory 1303 stores instructions for executing the log statement recommendation method.
[0164] The communication interface 1304 uses a transceiver module such as, but not limited to, a network interface card and a transceiver to implement the communication between the node device 1300 and other devices or communication networks. An embodiment of the present application also provides a distributed system. The distributed system includes at least one node device. The node device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the node device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0165] As Figure 14 shown, the distributed system includes at least one node device 1300. The memory 1303 in one or more node devices 1300 in the distributed system may store the same instructions for executing the log statement recommendation method.
[0166] In some possible implementation manners, the memory 1303 of one or more node devices 1300 in the distributed system may also respectively store partial instructions for executing the log statement recommendation method. In other words, a combination of one or more node devices 1300 may jointly execute the instructions for executing the log statement recommendation method.
[0167] It should be noted that the memories 1303 in different node devices 1300 in the distributed system may store different instructions, respectively for implementing partial functions of the log statement recommendation device. That is, the instructions stored in the memories 1303 of different node devices 1300 may implement the functions of one or more of the running module 111, determining module 112, generating module 113, obtaining module 114, classifying module 115, extracting module 116, input module 117, and training module 118.
[0168] In some possible implementations, one or more node devices in a distributed system can be connected via a network. Among them, the network can be a wide area network or a local area network, etc. Figure 15 shows a possible implementation. As Figure 15 shown, two node devices 1300A and 1300B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each node device. In this type of possible implementation, the memory 1303 in node device 1300A stores instructions for executing the functions of the running module 111. At the same time, the memory 1303 in node device 1300B stores instructions for executing the functions of the determination module 112, generation module 113, acquisition module 114, classification module 115, extraction module 116, input module 117, and training module 118.
[0169] Figure 15 The connection method between the distributed systems shown can be considered that the log statement recommendation method provided in this application needs to determine the first log specification information corresponding to the first running scenario based on the preset correspondence between the code running scenario and the log specification information, generate the recommended log statement corresponding to the first code according to the generation rule characterized by the first log specification information, acquire the historical code run by each node device in the distributed system and the historical log statements corresponding to the historical code, classify the code running scenarios of the historical code to obtain various preset code running scenarios, extract the log statement features respectively for the historical log statements corresponding to the historical code under each code running scenario, input the first log event model into the preset fault detection model based on a neural network, and train the fault detection model based on a neural network according to the log event models under various code running scenarios to obtain the preset fault detection model. Therefore, it is considered to hand over the functions implemented by the determination module 112, generation module 113, acquisition module 114, classification module 115, extraction module 116, input module 117, and training module 118 to node device 1300B for execution.
[0170] It should be understood that Figure 15 the functions of node device 1300A shown in
[0171] can also be completed by multiple node devices 1300. Similarly, the functions of node device 1300B can also be completed by multiple node devices 1300. Figure 14 and Figure 15 the connection method of the distributed system described above. The difference is that the memory 1303 in one or more node devices in this distributed system can store the same instructions for executing the log statement recommendation method.
[0172] In some possible implementations, the memory 1303 of one or more node devices 1300 in the distributed system may also store some instructions for executing the log statement recommendation method respectively. In other words, the combination of one or more node devices 1300 can jointly execute the instructions for executing the log statement recommendation method.
[0173] It should be noted that the memories 1303 in different node devices 1300 in the distributed system may store different instructions for executing partial functions of the log statement recommendation system. That is to say, the instructions stored in the memories 1303 of different node devices 1300 can implement the functions of the log statement recommendation device.
Claims
1. A log statement recommendation method, characterized in that, Applied to node devices in a distributed system, the distributed system including a plurality of the node devices; Each of the node devices implements the functions corresponding to the code by running the code; The method includes: Running a first code in a first running scenario, the first running scenario being any one of a plurality of preset code running scenarios of the distributed system; Based on a preset correspondence between the code running scenario and the log specification information, determining first log specification information corresponding to the first running scenario; The log specification information corresponding to a code running scenario includes the characteristics of the historical log statements generated when any node device in the distributed system runs historical code in this code running scenario; The log specification information is used to characterize the generation rule of the log statements; Generating recommended log statements corresponding to the first code according to the generation rule characterized by the first log specification information.
2. The method according to claim 1, wherein The method further includes: Obtaining the historical code run by each of the node devices in the distributed system and the historical log statements corresponding to the historical code; Classifying the code running scenarios of the historical code to obtain the plurality of preset code running scenarios; Respectively extracting log statement characteristics for the historical log statements corresponding to the historical code in each of the code running scenarios; Generating log specification information corresponding to each of the code running scenarios based on the log statement characteristics corresponding to each of the code running scenarios, to obtain the preset correspondence.
3. The method according to claim 2, characterized in that, The historical log statements corresponding to the historical code have recommended tags, the recommended tags being positive tags or negative tags opposite to the positive tags; The historical log statements are used to reflect the running situation of the corresponding historical code; The recommended tags are used to characterize the accuracy of the running situation reflected by the historical log statements; The respectively extracting log statement characteristics for the historical log statements corresponding to the historical code in each of the code running scenarios includes: Respectively extracting log statement characteristics for the target historical log statements corresponding to the historical code in each of the code running scenarios; The target historical log statements are the historical log statements with the recommended tag being the positive tag.
4. The method according to any one of claims 1 to 3, characterized in that, The log statement characteristics include one or more of the error category, statement format, and printing position of the log statement.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: Based on the recommended log statements corresponding to the first code, determining a first log statement corresponding to the first code, and printing the first log statement.
6. The method according to claim 5, wherein The method further includes: Extracting semantic information of the first log statement from the first log statement corresponding to the first code; Generating a first log event model based on the semantic information of at least one first log statement and the printing order between the at least one first log statement, the first log event model including structured information and unstructured information of the log events corresponding to the at least one first log statement; Input the first log event model into a preset fault detection model based on a neural network. The preset fault detection model is used to detect faults in the running condition of the first code, and the information output by the preset fault detection model includes the fault detection result of the running condition of the first code.
7. The method according to claim 6, wherein The method further includes: According to the fault detection result of the running condition of the first code output by the preset fault detection model, determine that the recommended label for the at least one first log statement is the first recommended label; the first recommended label is used to characterize the accuracy of the running condition of the first code reflected by the at least one first log statement.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: For each of the code running scenarios, extract the semantic information of the historical log statements corresponding to the historical code under the code running scenario. Based on the semantic information of at least one historical log statement and the printing order between the at least one historical log statement under the code running scenario, generate a log event model. The log event model includes the structured information and unstructured information of the log events corresponding to the at least one historical log statement. According to the log event models under various code running scenarios, train a fault detection model based on a neural network to obtain a preset fault detection model.
9. A log statement recommendation device, characterized in that, Applied to the node devices in a distributed system; the device includes: A running module, configured to run a first code in a first running scenario, where the first running scenario is any one of multiple preset code running scenarios of the distributed system. A determination module, configured to determine the first log specification information corresponding to the first running scenario based on a preset correspondence between the code running scenario and the log specification information; the log specification information corresponding to one code running scenario includes the characteristics of the historical log statements generated when any node device in the distributed system runs the historical code under this code running scenario; the log specification information is used to characterize the generation rule of the log statements. A generation module, configured to generate the recommended log statements corresponding to the first code according to the generation rule characterized by the first log specification information.
10. The device according to claim 9, characterized in that, The device further includes: An acquisition module, configured to acquire the historical codes run by each node device in the distributed system and the historical log statements corresponding to the historical codes. A classification module, configured to classify the code running scenarios of the historical codes to obtain the multiple preset code running scenarios. An extraction module, configured to extract log statement features for the historical log statements corresponding to the historical codes under each code running scenario respectively. The generation module is further configured to generate the log specification information corresponding to each code running scenario based on the log statement features corresponding to each code running scenario, to obtain the preset correspondence.
11. The device according to claim 10, characterized in that, The historical log statement corresponding to the historical code has a recommendation tag, and the recommendation tag is a positive tag or a negative tag opposite to the positive tag; the historical log statement is used to reflect the running situation of the corresponding historical code; the recommendation tag is used to characterize the accuracy of the running situation reflected by the historical log statement. The extraction module is specifically configured to extract log statement features for the target historical log statements corresponding to the historical codes in each code running scenario, where the target historical log statements are the historical log statements with the positive tag as the recommendation tag.
12. The device according to any one of claims 9-11, characterized in that, The log statement features include one or more of the error category, statement format, and printing position of the log statement.
13. The device according to any one of claims 9-12, characterized in that The determination module is further configured to determine a first log statement corresponding to the first code based on the recommended log statement corresponding to the first code, and print the first log statement.
14. The device according to claim 13, characterized in that The extraction module is further configured to extract semantic information of the first log statement from the first log statement corresponding to the first code. The generation module is further configured to generate a first log event model based on the semantic information of at least one first log statement and the printing order between the at least one first log statement, where the first log event model includes structured information and unstructured information of the log events corresponding to the at least one first log statement. The device further includes: An input module, configured to input the first log event model into a preset fault detection model based on a neural network, where the preset fault detection model is used to perform fault detection on the running situation of the first code, and the information output by the preset fault detection model includes the fault detection result of the running situation of the first code.
15. The device according to claim 14, characterized in that The determination module is further configured to determine a first recommendation tag for the recommendation tag of the at least one first log statement according to the fault detection result of the running situation of the first code output by the preset fault detection model. The first recommendation tag is used to characterize the accuracy of the running situation of the first code reflected by the at least one first log statement.
16. The device according to any one of claims 9-15, characterized in that The extraction module is further configured to extract semantic information of the historical log statements from the historical log statements corresponding to the historical codes in each code running scenario. The generation module is further configured to generate a log event model based on the semantic information of at least one historical log statement in the code running scenario and the printing order between the at least one historical log statement, where the log event model includes structured information and unstructured information of the log events corresponding to the at least one historical log statement. The device further includes: A training module, configured to train a preset fault detection model based on a neural network according to the log event model in various code running scenarios.
17. An electronic device, characterized in that, Comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the at least one processor to execute the log statement recommendation method according to any one of claims 1-8.
18. A distributed system, characterized in that, Comprising at least one node device, the node device comprising the electronic device according to claim 17, each node device comprising a processor and a memory; The processor is configured to execute instructions stored in the memory of the node device, so that the distributed system executes the log statement recommendation method according to any one of claims 1-8.
19. A computer-readable storage medium, characterized in that, Comprising computer program instructions, when the computer program instructions are executed by a distributed system, the distributed system executes the log statement recommendation method according to any one of claims 1-8.