Artificial Intelligence-based Log Parsing Method, Device, Terminal Device and Medium

Through the log analysis method based on artificial intelligence, abnormal tasks in SQL database task logs are automatically identified and marked, which solves the problem of low parsing efficiency in the existing technology, and realizes efficient task log analysis and exception detection.

CN114201376BActive Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111524237.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-07-18
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

In the prior art, the parsing efficiency of SQL database task logs is low, especially in the testing of multiple database protocols, requiring a lot of labor costs and manual audits, resulting in low efficiency and quality.

Method used

Using an artificial intelligence-based log analysis method, we obtain the task log of the target database, identify the data of the operation stage and task execution process, and use preset data conditions to compare and mark abnormal tasks to achieve automated analysis.

Benefits of technology

The analysis efficiency of database task logs can be improved without manual auditing, and the abnormal tasks can be effectively identified, which improves the testing efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114201376B_ABST
    Figure CN114201376B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of artificial intelligence, and particularly relates to a log parsing method, device, terminal device and medium based on artificial intelligence. According to the task logs generated when the target database is driven to run, this method determines N running stages in the task logs, obtains the running logs corresponding to each running stage from the task logs, identifies each running log, determines the tasks and task execution process data corresponding to each running log, compares each task execution process data with the first preset data condition, determines the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, respectively marks the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently, obtains the first parsing result of the task logs, realizes the identification of the task logs based on artificial intelligence, and combines the preset conditions to obtain the parsing result, without manual review and positioning of the parsing result, effectively improving the parsing efficiency of the database task logs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence, and particularly relates to a log parsing method, device, terminal device, and medium based on artificial intelligence. Background Art

[0002] Currently, in daily development and maintenance, it is necessary to test the performance of a Structured Query Language (SQL) database and locate problems such as bottlenecks and production. The existing tests use a manual subjective observation method, that is, after a test, the parsing results are compared one by one to detect whether there are errors. When it is necessary to locate and test multiple types of SQL statements of multiple database protocols at one time, it often requires a large amount of human cost, and manual review and location are required. The efficiency and quality of testing and location are not high. Therefore, how to improve the parsing efficiency of database task logs has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, embodiments of this application provide a log parsing method, device, terminal device, and medium based on artificial intelligence to solve the problem of low parsing efficiency of database task logs in the prior art.

[0004] In a first aspect, embodiments of this application provide a log parsing method based on artificial intelligence. The log parsing method includes:

[0005] Obtain task logs generated when a target database is driven to run, and determine N running stages when a Job in the task logs is executed, where N is an integer greater than zero;

[0006] Obtain the running logs corresponding to each running stage from the task logs, identify each running log, and determine the task corresponding to each running log and the task execution process data;

[0007] Compare each task execution process data with a first preset data condition to determine the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, where the first comparison result is that the task execution process data meets the first preset data condition, and the second comparison result is that the task execution process data does not meet the first preset data condition;

[0008] Mark the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently to obtain the first parsing result of the task logs.

[0009] In a second aspect, embodiments of this application provide a log parsing device based on artificial intelligence. The log parsing device includes:

[0010] A phase determination module, configured to obtain task logs generated when a target database is driven to run, and determine N running phases when a Job in the task logs is executed, where N is an integer greater than zero;

[0011] A task determination module, configured to obtain running logs corresponding to each running phase from the task logs, identify each running log, and determine the task corresponding to each running log and the task execution process data;

[0012] A comparison module, configured to compare each task execution process data with a first preset data condition to determine the task corresponding to the first comparison result and the task corresponding to the second comparison result, where the first comparison result is that the task execution process data meets the first preset data condition, and the second comparison result is that the task execution process data does not meet the first preset data condition;

[0013] A first parsing module, configured to respectively mark the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently to obtain a first parsing result of the task logs.

[0014] In a third aspect, an embodiment of the present application provides a terminal device, where the terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the log parsing method described in the first aspect is implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the log parsing method described in the first aspect is implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute the log parsing method described in the first aspect above.

[0017] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: According to the task logs generated when the target database is driven to run, the present application determines N running stages in the task logs, obtains the running logs corresponding to each running stage from the task logs, identifies each running log, determines the tasks and task execution process data corresponding to each running log, compares each task execution process data with the first preset data condition, determines the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, marks the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently respectively, obtains the first parsing result of the task logs, realizes the identification of the task logs based on artificial intelligence, and obtains the parsing result in combination with the preset conditions, without manual review and positioning of the parsing result, effectively improving the parsing efficiency of the database task logs. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of a log parsing method based on artificial intelligence provided in Embodiment 1 of the present application;

[0020] Figure 2 It is a flowchart of a log parsing method based on artificial intelligence provided in Embodiment 2 of the present application;

[0021] Figure 3 It is a structural schematic diagram of a log parsing device based on artificial intelligence provided in Embodiment 3 of the present application;

[0022] Figure 4 It is a structural schematic diagram of a terminal device provided in Embodiment 4 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0024] It should be understood that, as used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations.

[0025] It should also be understood that the term "and / or" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0026] As used in the specification of this application and the appended claims, the term "if" can be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0027] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are used only for descriptive distinction and should not be construed as indicating or implying relative importance.

[0028] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0029] The terminal device in the embodiments of this application may be a palm computer, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud terminal device, a personal digital assistant (PDA), etc. The embodiments of this application do not impose any restrictions on the specific type of the terminal device.

[0030] Embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0031] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0032] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0033] To illustrate the technical solution of the present application, specific embodiments are used for illustration below.

[0034] See Figure 1 , which is a schematic flowchart of a log parsing method based on artificial intelligence provided by Embodiment 1 of the present application. The above log parsing method is applied to a terminal device, and the terminal device is connected to a target database through a preset application programming interface (API). When the target data is driven to run to execute corresponding tasks, corresponding task logs will be generated, and the above task logs can be collected through the API. As Figure 1 shown, the log parsing method may include the following steps:

[0035] Step S101, acquire the task logs generated when the target database is driven to run, and determine N running stages when a Job in the task logs is executed.

[0036] Among them, the target database can refer to a Structured Query Language (SQL) database, a MySQL database, etc. When the above target database is driven to run, task logs are generated. The task logs can be log files formed by natural language, programming language, etc. The log files contain process records generated when the target database executes corresponding tasks. The records include at least one Job and each running stage and each task under each running stage when the Job is executing. The target database is driven to run by an Application. In the application, each operation of an Action generates a Job. When a Job is executed, it needs to be divided into several stages, and this stage is the running stage.

[0037] In this application, the task logs obtained through the API can be a log file package composed of log files. After obtaining the log file package, it is also necessary to classify and identify the log file package to determine the target file as the task log, extract the specific log content of the task log, and use it for subsequent identification of the task log. The above task logs contain at least one running stage, that is, the above N is an integer greater than zero.

[0038] In this application, a trained Natural Language Processing (NLP) model is used to identify the task logs. The NLP model needs to be trained with a large number of labeled texts. A large number of training sets are sent into the NLP model to obtain deviations, adjust weights and thresholds. The NLP model can be constructed based on traditional machine learning methods or deep learning methods. After training, a model that can extract keywords and perform semantic analysis on texts is obtained.

[0039] The above task logs are input into the NLP model to obtain corresponding keywords or key sentences, where the keywords or key sentences are used to represent the running stages in the above task logs. For example, a running stage is matched through the "stage" keyword.

[0040] In an implementation manner, the above terminal device is connected to other storage databases. The other storage databases are used to store the task logs generated when the target database is driven to run. According to the acquisition instruction of the terminal device, the other storage databases send the corresponding task logs to the terminal device.

[0041] Step S102: Obtain the running logs corresponding to each running stage from the task logs, identify each running log, and determine the tasks and task execution process data corresponding to each running log.

[0042] Among them, each running stage in the task log corresponds to the running log of that running stage, and the running log of each running stage is extracted. For example, the log content between a running stage and its next running stage is the information generated during the running process of that running stage, from which the running log corresponding to that running stage can be extracted, and this running log is used as the specific running feature of the first running stage. Among them, the running log may include the start time, end time, running duration, the number of reducer tasks, and the Uniform Resource Locator (URL) corresponding to the task.

[0043] In this application, matching fields are preset in advance, such as time matching fields, map task matching fields, URL matching fields, etc. Based on the matching fields, data corresponding to the fields can be matched from the running log. The terminal device compares the running log with each matching field to determine the running data of each running log, including the start time, end time, running duration, and the URL corresponding to the task, etc. Then, the log content of each task is extracted, and the task execution process data of the task is determined through NLP model recognition, including the start time, end time, running duration, and the amount of input data, number of bytes, amount of output data, number of bytes, etc. For example, the log content between a task and its next task is the information generated during the running process of that task, and this log content is the specific running feature of that task.

[0044] For example, two times are obtained under the task. By comparing the order of the two times, the start time and end time are determined. After obtaining the start time and end time, the corresponding running duration is calculated by computing the two.

[0045] Optionally, determining the N running stages during the execution of a Job in the task log includes:

[0046] Using the characters representing the running stage to match the corresponding running stage from the task log;

[0047] Correspondingly, obtaining the running log corresponding to each running stage from the task log includes:

[0048] According to the positions of each running stage matched in the task log, the log between the positions of two adjacent running stages is extracted, and it is determined that this log is the running log corresponding to the running stage with the earlier position among the two adjacent running stages.

[0049] Among them, the character representation can be a string, regular expression, etc. that can uniquely characterize the running stage. For example, when the running stage in the task log is represented by "stage-N", the string corresponding to "stage" is used to match the running stage, and the running log corresponding to the running stage of "stage-1" is the log between the positions of "stage-1" and "stage-2".

[0050] For example, in the task log, "stage-1" and "stage-2" are matched through the "stage" string, and the log content between "stage-1" and "stage-2" is used as the log content of "stage-1". The log content of "stage-1" includes:

[0051] starting job=job_1606375577593_12689991,

[0052] tracking URL=http: / / xxx.com.cn / proxy / application_1606375577593_12689991 / ;

[0053] Hadoop job information for stage-1: number of mappers: 1066, number ofreducers: 287;

[0054] The corresponding URL can be extracted through the tracking URL, the number of map tasks can be extracted through the number of mappers as 1066, and the number of reducers tasks can be extracted through the number of reducers as 287.

[0055] Optionally, each running log is identified, and the tasks and task execution process data corresponding to each running log include:

[0056] Using the character representation that characterizes the task, the corresponding task is matched from each running log;

[0057] The log content under each task is extracted, and the trained natural language processing model is used to identify the log content under each task to determine the task execution process data of each task.

[0058] Among them, the character representation can be a string, regular expression, etc. that can uniquely characterize the task. For example, when represented by "task-N" in the running log, the string corresponding to "task" is used to match the task, and the URL under "task-1" is used to obtain the corresponding running file as the log content of the task. The NPL model is used to identify the above running file to determine the task execution process data of the task, including start time, end time, running duration, input data volume, number of bytes, output data volume, number of bytes, etc.

[0059] For example, through the log generated by the hive sql driver, it is obtained that the hive sql has generated a total of 2 stages (running stages), namely stage1 and stage1; among them, the running information of stage1:

[0060] Start time t11, end time t21, running duration t31, number of map tasks s11 during running, number of reducer tasks s21, and URL:

[0061] http: / / xxx.com.cn / proxy / application_1606375577593_12689991 / ;

[0062] According to the URL corresponding to task1 in stage1, obtain the task execution process data of task1:

[0063] Start time t41, end time t51, running duration t61, and input data volume s31, number of bytes s41, output data volume s51 of each task;

[0064] The running information of stage2: start time t12, end time t22, running duration t32, number of map tasks s1 during running, number of reducer tasks s22, and URL:

[0065] http: / / xxx.com.cn / proxy / application_1606375577593_11119981 / ;

[0066] According to the URL corresponding to task2 in stage2, obtain the task execution process data of task2:

[0067] Start time t42, end time t52, running duration t62, and input data volume s32, number of bytes s42, output data volume s52 of each task.

[0068] Step S103: Compare each task execution process data with the first preset data condition to determine the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result.

[0069] Among them, the first comparison result is that the task execution process data meets the first preset data condition, and the second comparison result is that the task execution process data does not meet the first preset data condition. The first preset data condition can be a data condition set for any data in the task execution process data. Extract the corresponding data in each task execution process data and compare it with this data condition to determine that the task corresponding to the task execution process data that meets the condition is the task corresponding to the first comparison result, and determine that the task corresponding to the task execution process data that does not meet the condition is the task corresponding to the second comparison result.

[0070] For example, the first preset data condition is that the running duration is less than t. The running duration of task1 in stage1 is t61, and the running duration of task1 in stage1 is t62. Among them, t61 of task1 in stage1 is less than t, and the corresponding comparison result is the first comparison result. t62 of task1 in stage2 is greater than t, and the corresponding comparison result is the second comparison result.

[0071] Step S104: Mark the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently to obtain the first parsing result of the task log.

[0072] Among them, the first preset data condition is a pre-established judgment criterion, which can be used to determine whether there is an abnormal operation situation of the task, formulate a judgment criterion for performance anomalies, judge the task operation process, and mark the abnormally running tasks and the normally running tasks differently according to the judgment results. For example, display each task in each running stage, and mark the abnormally running tasks as highlighted compared with the normally running tasks, so as to attract the user's attention.

[0073] The parsing of the task log in this application correspondingly obtains a parsing result that differentiates the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result.

[0074] For example, through the logs generated by the driver of Hive SQL, it is obtained that a total of 2 stages are generated for this Hive SQL, namely stage1 and stage1; among them, the running logs of stage1 include: start time t11, end time t21, running duration t31, the number of map tasks s11 and the number of reducer tasks s21 during operation. Based on the running logs corresponding to stage1, determine the task execution process data of stage1: start time t41, end time t51, running duration t61, and the input data volume s31, byte count s41, and output data volume s51 of each task; the running information of stage2 includes: start time t12, end time t22, running duration t32, the number of map tasks s12 and the number of reducer tasks s22 during operation. Based on the running logs corresponding to stage2, determine the task execution process data of stage2: start time t42, end time t52, running duration t62, and the input data volume s32, byte count s42, and output data volume s52 of each task; the content to be displayed is: start time t41, end time t51, running duration t61, and the input data volume s31, byte count s41, and output data volume s51 of each task, and start time t42, end time t52, running duration t62, and the input data volume s32, byte count s42, and output data volume s52 of each task. The first preset data condition is the start time threshold. If t41 is greater than the start time threshold, then the first preset data condition is not met. If t42 is less than the start time threshold, then the first preset data condition is met. Therefore, when displaying, the tasks of stage1 are displayed in red font and the tasks of stage2 are displayed in black font.

[0075] Optionally, after obtaining the running logs corresponding to each running stage from the task logs and identifying each running log, it further includes:

[0076] Determine the running data corresponding to each running stage, where the running data is the data in the task logs other than the tasks and the task execution process data;

[0077] Compare the running data with the second preset data condition to determine the running stage corresponding to the third comparison result and the running stage corresponding to the fourth comparison result, where the third comparison result is that the running data meets the second preset data condition, and the fourth comparison result is that the running data does not meet the second preset data condition;

[0078] Mark the running stage corresponding to the third comparison result and the running stage corresponding to the fourth comparison result differently to obtain the third parsing result of the task logs.

[0079] Among them, the second preset data condition is a data condition set for any data in the operation data corresponding to the operation stage. Extract the corresponding data in each operation data and compare it with this data condition to determine that the operation stage corresponding to the operation data that meets the condition is the operation stage corresponding to the third comparison result, and determine that the operation stage corresponding to the operation data that does not meet the condition is the operation stage corresponding to the fourth comparison result.

[0080] The second preset data condition is also a pre-established judgment criterion, which can be used to determine whether there is an abnormal operation situation in the operation stage, formulate a judgment criterion for performance anomalies, judge the operation stage, and mark the operation stages with abnormal operations and normal operations separately according to the judgment results. For example, display each operation stage, and mark the operation stages with abnormal operations as highlighted compared with the operation stages with normal operations, so as to attract the user's attention.

[0081] Optionally, after comparing each task execution process data with the first preset data condition to determine the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, it further includes:

[0082] Extract the uniform resource locator corresponding to each task in the tasks corresponding to the second comparison result;

[0083] Access each uniform resource locator to obtain the corresponding task execution information, and determine the task execution information corresponding to each task in the tasks corresponding to the second comparison result as the fourth parsing result of the task log.

[0084] Among them, for the tasks that do not meet the first preset data condition, access the corresponding address through the URL of the task to obtain the task execution information of the task. The task execution information includes the stability of the Hadoop Distributed File System (HDFS) used by the task, and can also include the performance of the machine on which the task is executed, and whether a retry mechanism has been started, so as to judge the reasons for the above non-compliance with the first preset data condition.

[0085] Optionally, the task execution process data includes the start time when the task starts to execute and the processing duration of the task. After determining the tasks corresponding to each operation log and the task execution process data, it further includes:

[0086] For any operation stage, determine the number of all tasks in this operation stage, as well as the earliest start time, the latest start time, and the longest processing duration among all tasks;

[0087] If the product of the number of all tasks and the longest processing duration is less than the duration between the earliest start time and the latest start time, it is determined that there is a lack of resources during the task execution, and the insufficient resources are used as the fifth parsing result of the task log.

[0088] Among them, for the tasks in the same stage, it is determined whether the poor timeliness is caused by insufficient resources according to the earliest start time and the latest start time. For example: the longest processing time of the task only needs 3 minutes, while the difference between the earliest start time and the latest start time of the task is 1 hour, which means that the resources are the bottleneck at this moment. Theoretically, if the resources are sufficient, this stage should be able to run to completion within 5 minutes. Therefore, the solution is to increase resources or reduce the parallelism of this stage to let each task process more data to solve this problem.

[0089] When the above parsing results are displayed, the above operation logs and the tasks and task execution process data under each operation log are displayed in the form of a list. For example, the first column displays the application URL, the second column displays the number of map tasks, the third column displays the number of reducer tasks, the fourth column displays the start time, the fifth column displays the end time, and the sixth column displays the overall running time.

[0090] This application can also compare the input data volume and output data volume of the task with those of the median task to confirm whether the task is skewed and use it to optimize SQL.

[0091] According to the task log generated when the target database is driven to run, the embodiments of this application determine N running stages in the task log, obtain the running logs corresponding to each running stage from the task log, identify each running log, determine the tasks and task execution process data corresponding to each running log, compare each task execution process data with the first preset data condition, determine the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, and mark the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently to obtain the first parsing result of the task log. Based on artificial intelligence, the task log is identified, and the parsing result is obtained by combining the preset conditions, without manual review and positioning of the parsing result, effectively improving the parsing efficiency of the database task log.

[0092] See Figure 2 , which is a schematic flowchart of a log parsing method based on artificial intelligence provided by the second embodiment of this application. As Figure 2 shown, the log parsing method may include the following steps:

[0093] Step S201: Obtain the task log generated when the target database is driven to run, and determine N running stages when a Job is executed in the task log.

[0094] Step S202: Obtain the running log corresponding to each running stage from the task log, identify each running log, and determine the task and task execution process data corresponding to each running log.

[0095] Step S203: Compare each task execution process data with the first preset data condition to determine the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result.

[0096] Among them, the content of Step S201 to Step S203 is the same as that of the above-mentioned Step S101 to Step S103. For the description, please refer to Step S101 to Step S103 and will not be elaborated here.

[0097] Step S204: Determine the first type of running stage and the second type of running stage according to the running stage to which each task belongs in the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result.

[0098] Among them, the first type of running stage is the running stage that contains the tasks corresponding to the first comparison result, and the second type of running stage is the running stage that contains the tasks corresponding to the second comparison result. That is, the running stage to which the task that meets the conditions belongs is the first type of running stage, and the running stage to which the task that does not meet the conditions belongs is determined as the second type of running stage.

[0099] Step S205: Mark the first type of running stage and the second type of running stage differently to obtain the second parsing result of the task log.

[0100] Among them, the first preset data condition is a pre-established judgment criterion, which can also be used to determine whether there is an abnormal running situation in the running stage. According to the judgment result, the running stages with abnormal running and the running stages with normal running are marked differently. For example, each running stage is displayed, and the running stage with abnormal running is marked as highlighted compared with the running stage with normal running, so as to attract the user's attention.

[0101] The parsing of the task log in this application corresponds to the parsing result that differentiates the running stage corresponding to the first comparison result from the running stage corresponding to the second comparison result.

[0102] According to the task log generated when the target database is driven to run, the embodiments of the present application determine N running stages in the task log, obtain the running log corresponding to each running stage from the task log, identify each running log, determine the task corresponding to each running log and the task execution process data, compare each task execution process data with the first preset data condition, determine the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, determine the first type of running stage and the second type of running stage according to the running stage to which each task belongs in the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, respectively mark the first type of running stage and the second type of running stage, obtain the second parsing result of the task log, realize the identification of the task log based on artificial intelligence, obtain the parsing result in combination with the preset conditions, and generate the parsing result of the running stage.

[0103] Corresponding to the log parsing method in the above embodiment, Figure 3 The structural block diagram of the log parsing device based on artificial intelligence provided in the third embodiment of the present application is shown. The above log parsing device is applied to a terminal device, and the terminal device is connected to the target database through a preset application program interface. When the target database is driven to run to execute corresponding tasks, corresponding task logs will be generated, and the above task logs can be collected through the API. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0104] See Figure 3 , the log parsing device includes:

[0105] The stage determination module 31 is configured to obtain the task log generated when the target database is driven to run, and determine N running stages during the execution of a Job in the task log, where N is an integer greater than zero;

[0106] The task determination module 32 is configured to obtain the running log corresponding to each running stage from the task log, identify each running log, and determine the task corresponding to each running log and the task execution process data;

[0107] The comparison module 33 is configured to compare each task execution process data with the first preset data condition, and determine the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, where the first comparison result is that the task execution process data meets the first preset data condition, and the second comparison result is that the task execution process data does not meet the first preset data condition;

[0108] The first parsing module 34 is configured to respectively mark the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently, and obtain the first parsing result of the task log.

[0109] Optionally, the above log parsing device further includes:

[0110] A phase classification module, which is used to compare each piece of task execution process data with a first preset data condition, and after determining the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, determine a first type of running phase and a second type of running phase according to the running phases to which each task in the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result belongs, wherein the first type of running phase is the running phase containing the tasks corresponding to the first comparison result, and the second type of running phase is the running phase containing the tasks corresponding to the second comparison result;

[0111] A second parsing module, which is used to perform different markings on the first type of running phase and the second type of running phase respectively to obtain a second parsing result of the task log.

[0112] Optionally, the above log parsing device further includes:

[0113] An operation data determination module, which is used to obtain the operation log corresponding to each running phase from the task log, and after identifying each operation log, determine the operation data corresponding to each running phase, where the operation data is the data in the task log other than the tasks and the task execution process data;

[0114] A running phase determination module, which is used to compare the operation data with a second preset data condition to determine the running phase corresponding to the third comparison result and the running phase corresponding to the fourth comparison result, wherein the third comparison result is that the operation data meets the second preset data condition, and the fourth comparison result is that the operation data does not meet the second preset data condition;

[0115] A third parsing module, which is used to perform different markings on the running phase corresponding to the third comparison result and the running phase corresponding to the fourth comparison result respectively to obtain a third parsing result of the task log.

[0116] Optionally, the above log parsing device further includes:

[0117] An extraction module, which is used to compare each piece of task execution process data with a first preset data condition, and after determining the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result, extract the uniform resource locators corresponding to each task in the tasks corresponding to the second comparison result;

[0118] A fourth parsing module, which is used to access each uniform resource locator to obtain the corresponding task execution information, and determine that the task execution information corresponding to each task in the tasks corresponding to the second comparison result is the fourth parsing result of the task log.

[0119] Optionally, the above log parsing device further includes:

[0120] A task content determination module, which is used for the task execution process data including the start time when the task starts to execute and the processing duration of the task. After determining the task corresponding to each running log and the task execution process data, for any running stage, determine the number of all tasks in this running stage, as well as the earliest start time, the latest start time, and the longest processing duration among all tasks;

[0121] A fifth parsing module, which is used for determining that there is a lack of resources during the task execution if the product of the number of all tasks and the longest processing duration is less than the duration between the earliest start time and the latest start time, and taking the insufficient resources as the fifth parsing result of the task log.

[0122] Optionally, the above-mentioned task determination module 32 includes:

[0123] A first matching unit, which is used for matching the corresponding task from each running log by using the character representation characterizing the task;

[0124] A first recognition unit, which is used for extracting the log content under each task, and using the trained natural language processing model to recognize the log content under each task to determine the task execution process data of each task.

[0125] Optionally, the above-mentioned stage determination module 31 includes:

[0126] A first matching unit, which is used for matching the corresponding running stage from the task log by using the character representation characterizing the running stage;

[0127] Correspondingly, the above-mentioned task determination module 32 includes:

[0128] A log determination unit, which is used for extracting the log between the positions of two adjacent running stages according to the positions of each running stage matched in the task log, and determining that this log is the running log corresponding to the running stage with the earlier position among the two adjacent running stages.

[0129] It should be noted that the information interaction, execution process, etc. between the above-mentioned modules, because they are based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.

[0130] Figure 4 This is a schematic structural diagram of a terminal device provided in Embodiment 4 of the present application. As Figure 4 shown, the terminal device 4 in this embodiment includes: at least one processor 40 ( Figure 4 only one is shown here), a memory 41, and a computer program 42 stored in the memory 41 and operable on at least one processor 40. When the processor 40 executes the computer program 42, the steps in any of the above-mentioned method embodiments of the log parsing method are implemented.

[0131] The terminal device 4 may include, but is not limited to, a processor 40 and a memory 41. Those skilled in the art can understand that Figure 4 This is merely an example of the terminal device 4 and does not constitute a limitation on the terminal device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0132] The so-called processor 40 may be a CPU. The processor 40 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0133] The memory 41 may be an internal storage unit of the terminal device 4 in some embodiments, such as the hard disk or memory of the terminal device 4. The memory 41 may also be an external storage device of the terminal device 4 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 4. Further, the memory 41 may also include both the internal storage unit and the external storage device of the terminal device 4. The memory 41 is used to store an operating system, application programs, a boot loader, data, and other programs, such as program codes of computer programs. The memory 41 may also be used to temporarily store data that has been output or will be output.

[0134] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0135] To implement all or part of the processes in the above method embodiments of this application, it can also be completed by a computer program product. When the computer program product runs on a terminal device, the terminal device can be made to execute the steps in the above method embodiments when executed.

[0136] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0137] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0138] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0139] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0140] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included in the protection scope of this application.

Claims

1. An artificial intelligence-based log parsing method, characterized in that, The log parsing method includes: Obtain the task log generated when the target database is driven to run, use the trained NLP model to identify the N running stages when a Job in the task log is executed, input the task log into the NLP model to obtain corresponding keywords or key sentences, where the keywords or key sentences are used to characterize the running stages in the above task log, and N is an integer greater than zero; Obtain the running log corresponding to each running stage from the task log, identify each running log, and determine the task and task execution process data corresponding to each running log; wherein, the log file contains the process record generated when the target database executes the corresponding task, and the record includes at least one Job, each running stage when it is executed, and each task under each running stage; Compare each task execution process data with the first preset data condition to determine the task corresponding to the first comparison result and the task corresponding to the second comparison result, where the first comparison result is that the task execution process data meets the first preset data condition, and the second comparison result is that the task execution process data does not meet the first preset data condition; Mark the tasks corresponding to the first comparison result and the tasks corresponding to the second comparison result differently to obtain the first parsing result of the task log; Wherein, the determining the N running stages when a Job in the task log is executed includes: Using the character representation of the running stage, match the corresponding running stage from the task log; Correspondingly, the obtaining the running log corresponding to each running stage from the task log includes: According to the positions of each running stage matched in the task log, extract the log between the positions of two adjacent running stages, and determine that the log is the running log corresponding to the running stage with the earlier position among the two adjacent running stages.

2. The log parsing method according to claim 1, wherein After comparing each task execution process data with the first preset data condition to determine the task corresponding to the first comparison result and the task corresponding to the second comparison result, it further includes: Determine the first type of running stage and the second type of running stage according to the running stages to which each task belongs in the task corresponding to the first comparison result and the task corresponding to the second comparison result, where the first type of running stage is the running stage containing the task corresponding to the first comparison result, and the second type of running stage is the running stage containing the task corresponding to the second comparison result; Mark the first type of running stage and the second type of running stage differently to obtain the second parsing result of the task log.

3. The log parsing method according to claim 1, wherein After obtaining the running log corresponding to each running stage from the task log and identifying each running log, it further includes: Determine the running data corresponding to each running stage, where the running data is the data in the task log other than the task and the task execution process data; Compare the operation data with the second preset data condition to determine the operation stage corresponding to the third comparison result and the operation stage corresponding to the fourth comparison result, where the third comparison result is that the operation data meets the second preset data condition, and the fourth comparison result is that the operation data does not meet the second preset data condition; Mark the operation stage corresponding to the third comparison result and the operation stage corresponding to the fourth comparison result differently to obtain the third parsing result of the task log.

4. The log parsing method according to claim 1, wherein After comparing each task execution process data with the first preset data condition to determine the task corresponding to the first comparison result and the task corresponding to the second comparison result, it further includes: Extract the uniform resource locator corresponding to each task in the task corresponding to the second comparison result; Access each uniform resource locator to obtain the corresponding task execution information, and determine that the task execution information corresponding to each task in the task corresponding to the second comparison result is the fourth parsing result of the task log.

5. The log parsing method according to claim 1, wherein The task execution process data includes the start time when the task starts to execute and the processing duration of the task. After determining the task corresponding to each operation log and the task execution process data, it further includes: For any operation stage, determine the number of all tasks in this operation stage, as well as the earliest start time, the latest start time, and the longest processing duration among all tasks; If the product of the number of all tasks and the longest processing duration is less than the duration between the earliest start time and the latest start time, it is determined that there is a lack of resources during the task execution, and the insufficient resources are used as the fifth parsing result of the task log.

6. The log parsing method according to claim 1, wherein The identifying each operation log and determining the task corresponding to each operation log and the task execution process data includes: Use the character representation representing the task to match the corresponding task from each operation log; Extract the log content under each task, and use the trained natural language processing model to identify the log content under each task to determine the task execution process data of each task.

7. An artificial intelligence-based log parsing device, characterized in that, The log parsing device includes: A stage determination module, configured to obtain a task log generated when a target database is driven to run, use a trained NLP model to identify the task log to determine N operation stages when a Job in the task log is executed, input the task log into the NLP model to obtain corresponding keywords or key sentences, where the keywords or key sentences are used to represent the operation stages in the above task log, and N is an integer greater than zero; A task determination module, configured to obtain the operation log corresponding to each operation stage from the task log, identify each operation log, and determine the task corresponding to each operation log and the task execution process data; where the log file contains the process records generated when the target database executes the corresponding task, and the records include at least one Job, each operation stage when it is executed, and each task under each operation stage; A comparison module, configured to compare each task execution process data with a first preset data condition, and determine a task corresponding to a first comparison result and a task corresponding to a second comparison result, where the first comparison result is that the task execution process data meets the first preset data condition, and the second comparison result is that the task execution process data does not meet the first preset data condition; A first parsing module, configured to perform different markings on the task corresponding to the first comparison result and the task corresponding to the second comparison result respectively, to obtain a first parsing result of the task log; Wherein, the phase determination module includes: A first matching unit, configured to match a corresponding running phase from the task log by using a character representation of the running phase; The task determination module includes: A log determination unit, configured to extract the log between the positions of two adjacent running phases according to the positions of each running phase matched in the task log, and determine that the log is the running log corresponding to the running phase with the earlier position among the two adjacent running phases.

8. A terminal device, characterized in that, The terminal device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the log parsing method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the log parsing method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • General log analysis method, terminal equipment and storage medium

    CN111581057A

  • Task running log processing method and device, equipment and storage medium

    CN111611127A

  • Link tracking method and system

    CN113746883A