Log file processing method and device, equipment and storage medium
By filtering and up-down line completion processing in the log file, log blocks are generated, and a model is constructed to enter prompt words to analyze the cause of the failure, the problem of reduced fault diagnosis efficiency caused by the increase in the number of log files is solved, and automated fault diagnosis and root cause analysis are realized.
Patent Information
- Application Number
- CN202510192433.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
With the continuous expansion and iterative update of business system functions, the frequency of failure events increases, resulting in an increase in the number of log files, thereby reducing the efficiency of failure event diagnosis.
By obtaining the log lines in the pending log file corresponding to the target failure event, a log line collection is formed, and a log block is generated based on the log keyword filtering and up-down line completion processing. Then, use the log content in the log block to build a model and enter the prompt word, and enter the target model to generate the root cause analysis results of the failure event.
It realizes automated generation of root cause analysis results including the causes of failures, improves fault diagnosis efficiency, and reduces the time and energy of manual analysis.
Smart Images

Figure CN120045418A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular, to a method, apparatus, device, and storage medium for processing log files. Background Art
[0002] Log files are used to record various operation information generated during the operation of a business system. When a fault event occurs during the operation of the business system, technicians can check and locate the cause of the fault event by viewing the log files.
[0003] Currently, with the continuous expansion and iterative update of the functions of business systems, the occurrence frequency of fault events is increasing day by day, and the number of corresponding log files generated is also increasing.
[0004] Therefore, how to improve the diagnostic efficiency of fault events in order to timely understand the cause of the fault has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and storage medium for processing log files.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for processing log files, the method including:
[0007] Obtain log lines in a to-be-processed log file corresponding to a target fault event to form a first log line set; wherein, the log line includes a log keyword, and the log keyword is used to characterize the association degree between the log line and the target fault event;
[0008] Filter the log lines in the first log line set according to the log keyword to obtain a second log line set, and perform upper and lower line completion processing on each log line in the second log line set to generate a log block corresponding to each log line;
[0009] Construct a model input prompt word using the log content in the log block, and input the model input prompt word into a target model, and after being processed by the target model, output a root cause analysis result corresponding to the target fault event; wherein, the root cause analysis result includes the cause of the target fault event.
[0010] In an optional implementation manner, after performing upper and lower line completion processing on each log line in the second log line set to generate a log block corresponding to each log line, it further includes:
[0011] Determine the weight corresponding to each log line in the log block; wherein, the weight is determined based on the log keyword of the log line;
[0012] Perform a weighted operation on each log line in the log block to obtain the weight of the log block; wherein, the weight of the log block is used to characterize the degree of association between the log content in the log block and the target fault event;
[0013] Correspondingly, the constructing the model input prompt word using the log content in the log block includes:
[0014] If the total character length of each log block is greater than the character limit length of the model input prompt word, then construct the model input prompt word using the log content in the log block with a weight greater than the preset weight threshold.
[0015] In an alternative implementation, the if the total character length of each log block is greater than the character limit length of the model input prompt word, then construct the model input prompt word using the log content in the log block with a weight greater than the preset weight threshold includes:
[0016] If the total character length of each log block is greater than the character limit length of the model input prompt word, then construct the model input prompt word using the log content in the log block with a weight greater than the preset weight threshold and the log lines at the tail in the second set of log lines.
[0017] In an alternative implementation, the filtering the log lines in the first set of log lines according to the log keywords to obtain a second set of log lines includes:
[0018] Match the log keywords of each log line in the first set of log lines with a preset keyword set respectively, and construct a second set of log lines based on the log lines corresponding to the log keywords that match successfully; wherein, the preset keyword set includes log keywords whose degree of association with the target fault event meets a preset association condition.
[0019] In an alternative implementation, before performing the up and down line completion processing on the log lines in the second set of log lines to generate the log block corresponding to the log line, it further includes:
[0020] Match each log line in the second set of log lines with each log template in the log template set respectively, and construct a third set of log lines based on the log lines that do not match any log template; wherein, the log templates in the log template set are obtained by analyzing a standard log file, and the standard log file is a log file that has a corresponding relationship with the log file to be processed and does not have the target fault event;
[0021] Correspondingly, the performing the up and down line completion processing on the log lines in the second set of log lines to generate the log block corresponding to the log line includes:
[0022] Perform context completion processing on the log lines in the third set of log lines respectively to generate log blocks corresponding to the log lines.
[0023] In an alternative implementation, the step of matching each log line in the second set of log lines with each log template in the log template set respectively includes:
[0024] Determine the edit distance between the first log line in the second set of log lines and the first log template in the log template set, and determine that the first log line and the first log template match successfully when the ratio between the edit distance and the total character length of the first log template is less than a preset first threshold;
[0025] And / or,
[0026] Obtain the non-overlapping character length between the first log line in the second set of log lines and the first log template in the log template set, and determine that the first log line and the first log template match successfully when the ratio between the non-overlapping character length and the total character length of the first log template is less than a preset second threshold; wherein, the non-overlapping character length is the character length in the first log line that does not match successfully with the first log template.
[0027] In an alternative implementation, before obtaining the non-overlapping character length between the first log line in the second set of log lines and the first log template in the log template set, it further includes:
[0028] Perform word segmentation processing on the first log line in the second set of log lines and the first log template in the log template set respectively to obtain the word-segmented first log line and the word-segmented first log template;
[0029] Correspondingly, the step of obtaining the non-overlapping character length between the first log line in the second set of log lines and the first log template in the log template set includes:
[0030] Perform character matching on the word-segmented first log line and the word-segmented first log template to obtain the non-overlapping character length.
[0031] In an alternative implementation, the step of performing context completion processing on the log lines in the third set of log lines respectively to generate log blocks corresponding to the log lines includes:
[0032] Determine the position of the log line in the third set of log lines in the log file to be processed, and obtain the adjacent first N log lines and the adjacent last M log lines of the log line based on the position; M and N are preset integers;
[0033] Generate a log block for the log line based on the log line and the adjacent N previous lines and M subsequent lines of the log.
[0034] In an alternative embodiment, the root cause analysis result further includes the log line information corresponding to the fault cause, and the log line information is used to obtain the log content corresponding to the fault cause in the to-be-processed log file.
[0035] In a second aspect, the present disclosure provides a log file processing device, which includes:
[0036] An acquisition module, configured to acquire log lines in a to-be-processed log file corresponding to a target fault event to form a first set of log lines; wherein, the log lines contain log keywords, and the log keywords are used to characterize the association degree between the log lines and the target fault event;
[0037] A filtering and complementing module, configured to filter the log lines in the first set of log lines according to the log keywords to obtain a second set of log lines, and perform up and down line complementing processing on each log line in the second set of log lines to generate log blocks corresponding to each log line respectively;
[0038] A construction module, configured to construct a model input prompt word by using the log content in the log block, and input the model input prompt word into a target model, and output a root cause analysis result corresponding to the target fault event after being processed by the target model; wherein, the root cause analysis result includes the fault cause of the target fault event.
[0039] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the log file processing method provided by the embodiment of the present disclosure.
[0040] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the log file processing method provided by the embodiment of the present disclosure.
[0041] In a fifth aspect, the present disclosure provides a computer program product, which includes a computer program / instruction, and when the computer program / instruction is executed by a processor, the above method is implemented.
[0042] The technical solution provided by the embodiment of the present disclosure has the following advantages compared with the prior art:
[0043] In the log file processing method provided by the embodiments of the present disclosure, first, log lines in the to-be-processed log file corresponding to the target fault event are obtained to form a first set of log lines. Among them, the log lines contain log keywords, and the log keywords are used to characterize the degree of association between the log lines and the target fault event. Then, the log lines in the first set of log lines are filtered according to the log keywords to obtain a second set of log lines, and up and down line completion processing is performed on each log line in the second set of log lines to generate log blocks corresponding to each log line. Subsequently, model input prompt words are constructed using the log content in the log blocks, and the model input prompt words are input into the target model. After being processed by the target model, a root cause analysis result corresponding to the target fault event is output, where the root cause analysis result includes the fault cause of the target fault event.
[0044] After the embodiments of the present disclosure form a first set of log lines based on the to-be-processed log file corresponding to the target fault event, the log lines in the first set of log lines are filtered, up and down line completion processing is performed to generate log blocks, and model input prompt words are constructed using the log blocks to input into the target model to generate a root cause analysis result corresponding to the target fault event. It can be seen that the embodiments of the present disclosure can automatically generate a root cause analysis result including the fault cause of the target fault event based on the to-be-processed log file, improving the fault diagnosis efficiency of the target fault event so that users can timely understand the fault cause. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In combination with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.
[0046] Figure 1 It is a flowchart of a log file processing method provided by the embodiments of the present disclosure;
[0047] Figure 2 It is a flowchart of another log file processing method provided by the embodiments of the present disclosure;
[0048] Figure 3 It is a flowchart of another log file processing method provided by the embodiments of the present disclosure;
[0049] Figure 4 It is a schematic structural diagram of a log file processing device provided by the embodiments of the present disclosure;
[0050] Figure 5 It is a schematic structural diagram of a log file processing device provided by the embodiments of the present disclosure. Detailed Embodiments
[0051] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0052] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0053] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0054] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0055] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".
[0056] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0057] In the related art, when a fault event occurs during the operation of a business system or device, relevant personnel usually analyze and locate based on the log file corresponding to the fault event, and then find out the cause of the fault event.
[0058] However, with the continuous expansion and iterative update of the functions of business systems, the number of log files corresponding to fault events is huge. Facing the huge number of log files, technicians need to spend a lot of time locating the cause of the fault.
[0059] Therefore, how to improve the efficiency of fault diagnosis is a technical problem that urgently needs to be solved at present.
[0060] To this end, embodiments of the present disclosure provide a method for processing log files. Specifically, first, log lines in a to-be-processed log file corresponding to a target fault event are obtained to form a first set of log lines. Among them, the log lines contain log keywords, and the log keywords are used to characterize the degree of association between the log lines and the target fault event. Then, the log lines in the first set of log lines are filtered according to the log keywords to obtain a second set of log lines, and up and down line completion processing is performed on each log line in the second set of log lines to generate log blocks corresponding to the respective log lines. Subsequently, a model input prompt is constructed using the log content in the log blocks, and the model input prompt is input into a target model. After being processed by the target model, a root cause analysis result corresponding to the target fault event is output, where the root cause analysis result includes the fault cause of the target fault event.
[0061] After embodiments of the present disclosure form a first set of log lines based on a to-be-processed log file corresponding to a target fault event, the log lines in the first set of log lines are filtered, up and down line completion processing is performed to generate log blocks, and a model input prompt is constructed using the log blocks to input into a target model to generate a root cause analysis result corresponding to the target fault event. It can be seen that embodiments of the present disclosure can automatically generate a root cause analysis result including the fault cause of the target fault event based on the to-be-processed log file, improving the fault diagnosis efficiency of the target fault event so that users can timely understand the fault cause.
[0062] Based on this, embodiments of the present disclosure provide a method for processing log files, which will be introduced below in combination with specific embodiments.
[0063] Figure 1 FIG. is a schematic flowchart of a method for processing log files provided by embodiments of the present disclosure. This method can be executed by a log file processing device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device.
[0064] As Figure 1 shown, the method includes:
[0065] S101: Obtain log lines in a to-be-processed log file corresponding to a target fault event to form a first set of log lines.
[0066] Among them, the log lines contain log keywords, and the log keywords are used to characterize the degree of association between the log lines and the target fault event.
[0067] The target fault event in the embodiments of the present disclosure can be the fault event that needs to be solved currently. This fault event can refer to an event where abnormal conditions such as operation failure or unexpected results occur during the operation of a business system, device, etc. The to-be-processed log file corresponding to the target fault event can refer to the file that records the operation information of the target fault event. Specifically, the to-be-processed log file can be the abnormal log file of the target fault event, and the to-be-processed log file can include various operation operations and event information of the business system.
[0068] In the embodiments of the present disclosure, after obtaining the to-be-processed log file corresponding to the target fault event, based on the to-be-processed log file, the to-be-processed file is split into log lines in the form of rows, and the obtained log lines form a first log line set, that is, the first log line set includes the log lines in the to-be-processed log file.
[0069] The log lines in the first log line set may contain log keywords. The log keywords can be specific words, such as error, fail, warn, exception, etc. The log keywords are related to the target fault event, and they can represent the degree of association between the log content in the log line and the target fault event. That is, the higher the level of the log keywords included in the log line, the higher the degree of association between the log line and the target fault event.
[0070] The log file processing method provided by the embodiments of the present disclosure is applied to the client. Specifically, the client can be deployed on terminals such as smart phones, tablet computers, and desktop computers.
[0071] S102: Filter the log lines in the first log line set according to the log keywords to obtain a second log line set, and perform up and down line complement processing on each log line in the second log line set to generate log blocks corresponding to each log line respectively.
[0072] In practical applications, since the first log line set contains a large number of log lines, including log lines that are irrelevant to the current target fault event, by using log keywords for filtering, irrelevant log lines can be quickly filtered, reducing the amount of log line data and improving the efficiency and accuracy of subsequent fault diagnosis.
[0073] In an optional implementation manner, the log lines in the first log line set can be filtered based on a preset keyword set to construct a second log line set. Specifically, identify the log keywords included in each log line in the first log line set, match the above log keywords with the keywords in the preset keyword set, and if there is a log keyword that matches successfully with the keyword in the preset keyword set, add the log line where the log keyword is located to the second log line set.
[0074] Among them, the preset keyword set is a set pre-set for log keywords. The log keywords included in the preset keyword set satisfy the preset association condition with the target fault event. For example, the preset keyword set may include "error" and "fail". The preset association condition is a condition preset for judging the association degree between the log keyword and the target fault event. If a log keyword meets the preset association condition, the log keyword is added to the preset keyword set.
[0075] The second set of log lines is a subset obtained by filtering the log lines in the first set of log lines using the preset keyword set. The log keywords included in the log lines in the second set of log lines are a subset of the log keywords in the preset keyword set. That is to say, one or more log keywords in the preset keyword set are included in the log lines in the second set of log lines.
[0076] Based on the above embodiments, since a log line usually only records information at a specific moment, to ensure information integrity, in the embodiments of the present disclosure, after obtaining the second set of log lines, the up and down line completion processing is performed on each log line in the second set of log lines to obtain log blocks corresponding to each log line. Among them, the up and down line completion processing may refer to an operation of perfecting and supplementing the log line.
[0077] In an optional implementation manner, the way to perform the up and down line completion processing on each log line in the second set of log lines may be to obtain the positions of each log line in the log file to be processed in the second set of log lines, and based on the positions, obtain the adjacent first N log lines and the adjacent last M log lines of each log line, and then splice each log line and the adjacent first N log lines and the adjacent last M log lines to obtain the log block corresponding to each log line. Among them, M and N are preset integers. For example, M is 5 and N is 5.
[0078] By performing the up and down line completion processing on the log line, the event information corresponding to the log line can be made more complete, which helps to understand the context information and improve the accuracy of subsequent fault diagnosis.
[0079] In the embodiments of the present disclosure, after performing the up and down line completion processing on each log line in the second set of log lines, a log block is obtained. The log block includes multiple log lines. The log block may be a log unit jointly composed of the log lines in the second set of log lines and the upper and lower log lines adjacent to the log line, and the upper and lower log lines adjacent to the log line can be obtained from the log file to be processed.
[0080] In practical applications, when performing the up-and-down line completion process on each log line in the second set of log lines, if there is an overlapping part between the adjacent upper and lower log lines of a certain log line and the adjacent upper and lower log lines of another log line, then the embodiments of the present disclosure can merge the above log lines into a log block. That is, if there is an overlapping part in the upper and lower log lines to be supplemented for at least two log lines, then the at least two log lines and the overlapping log lines can be merged together to form a log block. In this way, it is possible to effectively avoid duplicate log line content in the log block.
[0081] For example, one log line is the 6th line, and the upper and lower log lines to be completed are the 1st to 5th lines and the 7th to 10th lines; another log line is the 13th line, and the upper and lower log lines to be completed are the 8th to 12th lines and the 14th to 18th lines. Obviously, the 8th to 10th lines are the overlapping part, so the log lines from the 1st to 18th lines can form a log block.
[0082] S103: Construct a model input prompt word using the log content in the log block, and input the model input prompt word into the target model. After being processed by the target model, output the root cause analysis result corresponding to the target fault event.
[0083] Among them, the root cause analysis result includes the fault cause of the target fault event.
[0084] After generating the log block by performing the up-and-down line completion process on each log line in the second set of log lines as described above, construct a model input prompt word based on each log block. This model input prompt word can be an input parameter of the target model. Specifically, each log block can be spliced to form the model input prompt word.
[0085] In practical applications, on the basis of constructing the model input prompt word based on the log block, to ensure the accuracy of the output of the target model, the model constraint conditions can also be used to construct the model input prompt word. Among them, the model constraint conditions can include the model input prompt word construction template, operation configuration information, model output format requirements, etc.
[0086] Among them, after obtaining the model input prompt word, input the model input prompt word into the target model. After being processed by the target model, output the root cause analysis result corresponding to the target fault event. Among them, the target model is used to generate the root cause analysis result of the target fault from the model input prompt word. This root cause analysis result is an output parameter of the target model. The root cause analysis result is the cause conclusion summarized after comprehensively diagnosing the target fault problem.
[0087] The root cause analysis result may include the cause of the target fault event, where the cause of the fault may refer to the root cause that leads to the occurrence of the target fault event, and the cause of the fault can be used to subsequently resolve the target fault event.
[0088] In practical applications, to verify the accuracy of the cause of the fault in the root cause analysis result, the root cause analysis result output by the target model in the embodiments of the present disclosure may further include log line information corresponding to the cause of the fault, and the log line information is used to obtain the log content corresponding to the cause of the fault in the log file to be processed.
[0089] The log line information includes information related to the log content. Specifically, the log line may include information such as the location and line number of the log content.
[0090] The log file processing method provided by the embodiments of the present disclosure. Specifically, first, obtain the log lines in the log file to be processed corresponding to the target fault event to form a first log line set, where the log lines contain log keywords, and the log keywords are used to characterize the association degree between the log lines and the target fault event. Then, filter the log lines in the first log line set according to the log keywords to obtain a second log line set, and perform up and down line completion processing on each log line in the second log line set to generate log blocks corresponding to each log line. Subsequently, use the log content in the log blocks to construct a model input prompt, and input the model input prompt into the target model. After being processed by the target model, the root cause analysis result corresponding to the target fault event is output, where the root cause analysis result includes the cause of the target fault event.
[0091] After the embodiments of the present disclosure form the first log line set based on the log file to be processed corresponding to the target fault event, filter and perform up and down line completion processing on the log lines in the first log line set to generate log blocks, and use the log blocks to construct a model input prompt to input the target model to generate the root cause analysis result corresponding to the target fault event. It can be seen that the embodiments of the present disclosure can automatically generate a root cause analysis result including the cause of the target fault event based on the log file to be processed, improving the fault diagnosis efficiency of the target fault event so that the user can timely understand the cause of the fault.
[0092] In practical applications, after performing up and down line completion processing on each log line in the second log line set to generate log blocks corresponding to each log line, the order of constructing the model input words may also be determined based on the weights of the log blocks.
[0093] In an alternative implementation, after obtaining the log blocks, the weights corresponding to each log line in each log block are determined. Then, a weighted operation is performed on the weights of each log line to obtain the weights of each log block. Among them, the weight corresponding to a log line is determined based on the log keywords included in the log line. The weight of a log block can represent the degree of association between the log content in the log block and the target fault event, that is, the higher the weight of the log block, the higher the degree of association between the log content in the log block and the target fault event.
[0094] In an alternative implementation, the method for determining the weight of a log line based on the log keywords included in the log line can specifically determine the weight of the log line based on the level of the log keywords. Specifically, the log keywords included in the log line are determined, and then the weight of the log line is determined based on the level of the log keyword. Among them, the level of the log keyword can be set based on manual experience and is not limited here. The weights corresponding to different levels of log keywords can be set in advance.
[0095] Among them, the higher the level of the log keyword, the higher the weight of the log line where the log keyword is located. The weight of a log line can represent the degree of association between the log content of the log line and the target fault event. For example, a certain log line contains the log keywords "error" and "warn". Assuming that the level of "error" is higher than that of "warn", the weight of "error" is 5, and the weight of "warn" is 3, then the weight of this log line can be 5 + 3 = 8.
[0096] In an alternative implementation, the method for performing a weighted operation on the weights of each log line in a log block to obtain the weights of each log block can include performing a weighted sum on the weights of each log line in the log block to obtain a total weight value, and determining the number of log lines included in the log block. Based on this number of lines and the total weight value, the average weight is calculated as the weight of the log block. For example, assume that there are 2 log lines in log block A, and the weights of each log line are calculated as W1 and W1 based on the log keywords. The total weight value obtained by weighted summation is W1 + W2, and then the weight of log block A is (W1 + W2) / 2.
[0097] After determining the weights of the log blocks, the log content of the log blocks can be concatenated in order of the weights of the log blocks from high to low to form the model input prompt.
[0098] Based on the above embodiments, since the model input prompt has a character limit length, when constructing the model input prompt using the log content in the log block in the embodiments of the present disclosure, it is also necessary to determine whether the total character length of the log block is greater than the character limit length of the model input prompt. Among them, the character limit length is the maximum number of characters for constructing the model input prompt, that is, the maximum number of characters that the target model can receive.
[0099] Specifically, first, count the total character length of the log block. When constructing the model input prompt using the log content in the log block, determine whether the total character length of the log block is less than or equal to the character limit length of the model input prompt. If the total character length of the log block is not greater than this character limit length, then construct the model input prompt using the log content in the log block.
[0100] If the total character length of the log block is greater than this character limit length, then preferentially construct the model input prompt using the log content in the log block with a weight greater than the preset weight threshold. Here, the preset weight threshold is a pre-set weight value. Here, the preset weight threshold is a pre-set weight value.
[0101] In an optional implementation manner, on the basis of determining that the total character length of the log block is greater than the character limit length, it is also necessary to continue to determine whether the total character length of the log block with a weight greater than the preset weight threshold is greater than the character limit length. If the total character length of the log block with a weight greater than the preset weight threshold is greater than the character limit length, then construct the model input word based on the log content in the log block with a weight greater than the preset weight threshold and the log lines at the tail in the second log line set.
[0102] Specifically, on the basis that the total character length of the log block with a weight greater than the preset weight threshold is greater than the character limit length, splice the model input prompt in the order of the log block weights, and determine whether the remaining character length of the currently constructed model input prompt is sufficient to splice the log content of the next log block. If it is determined that the remaining character length of the currently constructed model input prompt is not sufficient to splice the log content of the next log block, then obtain the log lines at the tail position from the second log line set for splicing the model input prompt until the character length of the currently constructed model input prompt reaches the character limit length. Here, the remaining character length is the difference between the character length of the currently constructed model input prompt and the character limit length.
[0103] For example, assume that the character limit length of the model input prompt is 10,000 characters, the total character length of log block A is 5,000, the total character length of log block B is 1,000, the total character length of log block C is 3,000, and the total character length of log block D is 2,000. The weight levels among log block A, log block B, log block C, and D are such that the weight of log block A is greater than the weight of log block B, the weight of log block B is greater than the weight of log block C, and the weight of log block C is greater than the weight of log block D. Then, based on the weight levels, log block A, log block B, and log block C are preferentially used to construct the model output word. At this time, the remaining character length of the current model input word is 1,000 (10,000 - 5,000 - 1,000 - 3,000). It can be seen that the remaining character length of 1,000 is not sufficient to splice the log content of log block D. Therefore, the log line at the tail position is obtained from the second log line set and spliced into the model input prompt.
[0104] In the embodiments of the present disclosure, when constructing the model input word using the log content in the log block, it is determined whether the remaining character length of the current model input prompt can accommodate the next log block. When the remaining character length of the model input prompt is not sufficient to splice the log block, the log line at the tail in the second log line set is used for splicing, so that the target model receives more complete information, improving the diagnostic efficiency and accuracy of the target fault event.
[0105] In an optional implementation, after forming the first log line set from the log lines in the log file to be processed corresponding to the target fault event, each log line in the first log line set can also be matched with each log template in the log template set to obtain a third log line set, so as to filter out irrelevant log lines, reduce the data volume, and improve the efficiency of subsequent fault diagnosis.
[0106] Specifically, each log line in the first log line set is respectively matched with each log template in the log template set, and a third log line set is constructed based on the log lines that do not match any log template. Among them, the log templates in the log template set are generated based on the standard log file, and the log templates can be used to represent log lines with the same structure in the log file to be processed. The standard log file is a log file that has a corresponding relationship with the log file to be processed and does not contain the target fault event. For example, the standard log file can be a successful log file.
[0107] In an optional implementation, the method for generating the log template set for the standard log file can be to analyze each log line in the standard log file, determine the structure of each log line, and determine the log lines with the same structure as the same log template, thereby forming the log template set. That is, log lines with the same structure are represented by one log template. The specific algorithm implementation for generating the log template is not limited in the embodiments of the present disclosure.
[0108] In practical applications, before generating a log template for a standard log file, it is also possible to replace the target string in the log lines included in the standard log file to improve the generation efficiency of the log template and enhance the generality of the log template. The target string can be a common variable, such as a date variable, an IP address, etc. Specifically, identify the target string in the log line, replace the target string with an identifier to obtain the replaced standard log file, and then generate a log template for the replaced standard log file.
[0109] Among them, the matching of each log line in the first log line set with each log template in the log template set can mean that each log line in the first log line set is sequentially matched with each log template in the log template set. In practical applications, to improve the matching efficiency, a parallel processing method can be adopted to simultaneously perform matching operations on each log line in the first log line set and each log template in the log template set.
[0110] In an optional implementation manner, the method for matching each log line in the first log line set with each log template in the log template set can be to determine the edit distance between the first log line in the first log line set and the first log template in the log template set, and calculate the ratio of the edit distance to the total character length of the first log template. If the ratio is less than a preset first threshold, it is determined that the first log line and the first log template match successfully. Among them, the first log line is any log line in the first log set, and the first log template is any log template in the log template set.
[0111] Among them, the edit distance, also known as the Levenshtein distance, refers to the minimum number of edit operations required to convert one character into another between two characters. The edit operations can include replacement, insertion, and deletion. For example, the edit distance for converting abcd to acg is 3, and the edit operations are: deleting b, deleting d, and adding g at the end. The total character length of the first log template can be the total number of characters included in the first log template. The preset first threshold is a pre-set value, and this preset first threshold is used to determine whether the first log line matches the first log template successfully. Specifically, if the ratio of the edit distance between the first log line and the first log template to the total character length of the first log template is less than the preset first threshold, it indicates that the first log line and the first log template match successfully; otherwise, it is the opposite.
[0112] If the first log line matches the first log template successfully, it can be indicated that the first log line appears in the standard log file and has a low degree of association with the target fault event. If the first log line does not match the first log template successfully, it indicates that the first log line does not appear in the standard log file and has a high degree of association with the target fault event.
[0113] The ratio between the edit distance and the total character length of the first log template being less than a preset first threshold can be expressed by the following formula:
[0114]
[0115] Among them, Levinsten(abnormal_log_line,normal_template) is used to calculate the edit distance, where abnormal_log_line represents the first log line, normal_template represents the log template, length(normal_template) is used to calculate the total character length of the log template, and threshold represents the preset first threshold. For example: Assume that the total character length of the first log template is 20, the edit distance between the first log template and the first log line is 3, and the preset first threshold is 0.2. Then the ratio between the edit distance and the total character length of the first log template is 3 / 20 = 0.15. Since this ratio 0.15 is less than the preset first threshold 0.2, it is determined that the first log line matches the first log template successfully.
[0116] In another alternative implementation, it is also possible to determine whether the first log line matches the first log template based on the ratio between the non-overlapping character length between the first log line in the first log line set and the first log template in the log template set and the total character length of the first log template and a preset second threshold.
[0117] Specifically, obtain the non-overlapping character length between the first log line and the first log template, as well as the total character length of the first log template, calculate the ratio between the non-overlapping character length and the total character length, and determine whether this ratio is less than the preset second threshold. If this ratio is less than the preset second threshold, it is determined that the first log line matches the first log template successfully. Otherwise, it is the opposite.
[0118] Among them, the non-overlapping character length between the first log line and the first log template can refer to the length of the different characters in the first log line and the first log template. For example, if the log line is abcde and the log template is abxyz, then the non-overlapping character length between the log line and the log template is 3. The preset second threshold is a preset value, and the preset first threshold and the preset second threshold can be the same or different.
[0119] The ratio between the overlapping character length between the first log line and the first log template and the total character length of the first log template is less than a preset second threshold, which can be expressed by the following formula:
[0120]
[0121] Where abnormal_logline represents the first log line, normal_template represents the log template, length(normal_temolate) is used to calculate the total character length of the log template, and threshold represents the preset second threshold. is a set comprehension, indicating that character elements not belonging to the intersection of abnormal_logline and normal_template are selected from abnormal_logline. is used to calculate the length of the character elements selected from abnormal_logline that do not belong to the intersection of abnormal_logline and normal_template.
[0122] For example: Assume that the total character length of the first log template is 20, the length of the character elements in the first log line that do not belong to the intersection of the first log template and the character elements of the first log line is 2, and the preset second threshold is 0.15. Then the ratio between the length of the character elements that do not belong to the intersection of the first log template and the character elements of the first log line and the total character length of the first log template is 2 / 20 = 0.1. Since this ratio 0.1 is less than the preset second threshold 0.15, it is determined that the first log line and the first log template match successfully.
[0123] When using the above two methods to determine whether the first log line and the first log template are successful, as long as one of the above two formulas is satisfied, it can be determined that the first log line and the first log template match successfully.
[0124] In addition, to improve the matching efficiency, embodiments of the present disclosure may adopt a parallel processing method to simultaneously use the above two methods to determine whether each log line in the first log line set matches each log template in the log template set.
[0125] In an alternative embodiment, to improve the matching accuracy, before determining whether the first log line in the first log set matches the first log template in the log template set, the first log line and the first log template may be tokenized to obtain the tokenized first log line and the tokenized first log template, and then the tokenized first log line and the tokenized first log template are matched. Among them, tokenization may refer to splitting a continuous string into multiple strings according to a preset rule.
[0126] For example: The log content of the log line is "2000-01-15 10:00:00 Zhang San logged in to the system", and through tokenization, the log content of the tokenized log line can be obtained, such as ["date", "user", "logged in", "system"].
[0127] After the above matching of the log templates in the log template set with each log line in the first log line set is completed, the log lines that have not been successfully matched with any log template are used to construct a third log line set, and then the up and down lines of each log line in the third log line set are complemented to obtain the log blocks corresponding to each log line. Furthermore, the log content in the log blocks is used to construct model input prompt words and input them into the target model to obtain the root cause analysis result of the target fault event.
[0128] To facilitate the understanding of the content of the present disclosure, the present disclosure also provides another log file processing method. Refer to Figure 2 , which is a schematic diagram of another log file processing method provided by the embodiments of the present disclosure.
[0129] S201: Obtain the log lines in the to-be-processed log file corresponding to the target fault event to form a first log line set.
[0130] Among them, the log line contains a log keyword, and the log keyword is used to characterize the degree of association between the log line and the target fault event.
[0131] S202: Filter the log lines in the first log line set according to the log keyword to obtain a second log line set.
[0132] S203: Match each log line in the second log line set with each log template in the log template set, and construct a third log line set based on the log lines that have not been successfully matched with any log template.
[0133] Among them, the log templates in the log template set are obtained by analyzing a standard log file, and the standard log file is a log file that has a corresponding relationship with the to-be-processed log file and does not have the target fault event.
[0134] Among them, the content of S201 - S203 can be understood with reference to the content of the above embodiments, which will not be elaborated here.
[0135] In practical applications, the execution order of S202 and S203 can be unrestricted.
[0136] S204: Perform context completion processing on the log lines in the third log line set respectively to generate log blocks corresponding to the log lines.
[0137] In an optional implementation manner, based on the log lines in the third log line set, determine the position of each log line in the log file to be processed, and obtain the adjacent first N lines and adjacent next M lines of each log line based on this position. Then, based on each log line and the adjacent first N lines and adjacent next M lines, generate the log block of each log line. Wherein, M and N are preset integers.
[0138] S205: Use the log content in the log block to construct a model input prompt, and input the model input prompt into the target model. After being processed by the target model, output the root cause analysis result corresponding to the target fault event. Wherein, the root cause analysis result includes the fault cause of the target fault event.
[0139] In practical applications, after outputting the root cause analysis result, the embodiments of the present disclosure can also display the root cause analysis result on the front - end interface, so that users can obtain a solution to handle the target fault event based on the root cause analysis result. In addition, during the process of displaying the root cause analysis result on the interface, the embodiments of the present disclosure can also support user feedback on the root cause analysis result. Specifically, interactive controls can be displayed on the front - end interface. When receiving the information output by the user, answer based on this information. The interactive control is used to input feedback information for the root cause analysis result or to ask questions about the root cause analysis.
[0140] To facilitate the understanding of the content of the above embodiments, the embodiments of the present disclosure also provide a data schematic diagram for log file processing. Refer to Figure 3 .
[0141] In an optional implementation manner, obtain the log file to be processed corresponding to the target fault event, split the log lines in the log file to be processed, and construct a first log line set from the split log lines. Among them, the log lines in the first log line set contain log keywords.
[0142] In an alternative implementation, after obtaining the first set of log lines, to filter out the log lines that are not relevant to the target fault event, each log line in the first set of log lines is respectively matched against a preset keyword set according to the log keywords of the log lines, and a second set of log lines is constructed based on the log lines corresponding to the successfully matched log keywords. Among them, the preset keyword set includes log keywords whose degree of association with the target fault event meets a preset association condition.
[0143] In an alternative implementation, based on the above embodiment, the log lines in the second set of log lines are filtered to further retrieve the number of log lines and improve the subsequent model diagnosis efficiency. Specifically, each log line in the second set of log lines is respectively matched against each log template in the log template set, and a third set of log lines is constructed based on the log lines that are not successfully matched with any log template. Among them, the log templates in the log template set are generated based on standard log files.
[0144] In an alternative implementation, based on the above-obtained third set of log lines, the positions of the log lines in the third set of log lines in the log file to be processed are determined, and the adjacent first N lines and adjacent last M lines of the log lines are obtained based on the positions. Then, a log block corresponding to the log line is generated based on the log line and the adjacent first N lines and adjacent last M lines of the log lines.
[0145] After obtaining the log block corresponding to the log line, the weights corresponding to each log line in the log block are determined, and a weighted operation is performed on each log line to obtain the weight of the log block. The weight of the log block is used to represent the degree of association between the log content in the log block and the target fault event.
[0146] In an alternative implementation, after obtaining the log block weight, a model input prompt is constructed using the log content in the log block, and the model input prompt is input into the target model. After being processed by the target model, a root cause analysis result corresponding to the template fault event is output, and the root cause analysis result includes the fault cause of the target fault event.
[0147] To implement the above embodiment, the present disclosure also proposes a log file processing device. Figure 4 The following is a schematic structural diagram of a log file processing device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 4 shown, the device includes:
[0148] An acquisition module 401, configured to acquire log lines in a to-be-processed log file corresponding to a target fault event, and form a first set of log lines; wherein, the log lines include log keywords, and the log keywords are used to characterize the association degree between the log lines and the target fault event;
[0149] A filtering and complementing module 402, configured to filter the log lines in the first set of log lines according to the log keywords to obtain a second set of log lines, and perform upper and lower line complementing processing on each log line in the second set of log lines to generate a log block corresponding to each log line;
[0150] A construction module 403, configured to construct a model input prompt word by using the log content in the log block, and input the model input prompt word into a target model, and output a root cause analysis result corresponding to the target fault event after being processed by the target model; wherein, the root cause analysis result includes the fault cause of the target fault event.
[0151] In an optional implementation manner, the device further includes:
[0152] A determination module, configured to determine the weight corresponding to each log line in the log block; wherein, the weight is determined based on the log keyword of the log line;
[0153] A weighted processing module, configured to perform weighted operation processing on each log line in the log block to obtain the weight of the log block; wherein, the weight of the log block is used to characterize the association degree between the log content in the log block and the target fault event;
[0154] Correspondingly, the construction module is specifically configured to:
[0155] If the total character length of each log block is greater than the character limit length of the model input prompt word, then construct the model input prompt word by using the log content in the log block with a weight greater than a preset weight threshold.
[0156] In an optional implementation manner, the construction module is specifically configured to:
[0157] If the total character length of each log block is greater than the character limit length of the model input prompt word, then construct the model input prompt word by using the log content in the log block with a weight greater than a preset weight threshold and the log lines at the tail in the second set of log lines.
[0158] In an optional implementation manner, the filtering and complementing module is specifically configured to:
[0159] Match the log keywords of each log line in the first set of log lines with a preset set of keywords respectively, and construct a second set of log lines based on the log lines corresponding to the successfully matched log keywords; wherein, the preset set of keywords includes log keywords whose degree of association with the target fault event meets a preset association condition.
[0160] In an optional implementation manner, the apparatus further includes:
[0161] A matching module, configured to match each log line in the second set of log lines with each log template in a set of log templates respectively, and construct a third set of log lines based on the log lines that fail to match any log template; wherein, the log templates in the set of log templates are obtained by analyzing a standard log file, and the standard log file is a log file that has a corresponding relationship with the log file to be processed and does not have the target fault event.
[0162] Correspondingly, the filtering and completion module is specifically configured to:
[0163] Perform context completion processing on the log lines in the third set of log lines respectively to generate log blocks corresponding to the log lines.
[0164] In an optional implementation manner, the matching module includes:
[0165] A first determination sub-module, configured to determine the edit distance between a first log line in the second set of log lines and a first log template in the set of log templates, and determine that the first log line matches the first log template successfully when the ratio between the edit distance and the total character length of the first log template is less than a preset first threshold.
[0166] And / or,
[0167] An acquisition sub-module, configured to acquire the non-overlapping character length between a first log line in the second set of log lines and a first log template in the set of log templates, and determine that the first log line matches the first log template successfully when the ratio between the non-overlapping character length and the total character length of the first log template is less than a preset second threshold; wherein, the non-overlapping character length is the character length of the first log line that fails to match the first log template.
[0168] In an optional implementation manner, the apparatus further includes:
[0169] A word segmentation processing module, configured to perform word segmentation processing on a first log line in the second set of log lines and a first log template in the set of log templates respectively to obtain the word-segmented first log line and the word-segmented first log template.
[0170] Correspondingly, the obtaining sub-module is specifically configured to:
[0171] Perform character matching on the tokenized first log line and the tokenized first log template to obtain the non-overlapping character length.
[0172] In an optional implementation manner, the filtering and complementing module includes:
[0173] A second determination sub-module, configured to determine the position of the log line in the third log line set in the to-be-processed log file, and obtain the adjacent first N log lines and the adjacent last M log lines of the log line based on the position; M and N are preset integers;
[0174] A generation sub-module, configured to generate a log block of the log line based on the log line and the adjacent first N log lines and the adjacent last M log lines.
[0175] In an optional implementation manner, the root cause analysis result further includes the log line information corresponding to the fault cause, and the log line information is used to obtain the log content corresponding to the fault cause in the to-be-processed log file.
[0176] In the log file processing device provided by the embodiments of the present disclosure, specifically, first, obtain the log lines in the to-be-processed log file corresponding to the target fault event to form a first log line set, where the log line contains a log keyword, and the log keyword is used to characterize the association degree between the log line and the target fault event. Then, filter the log lines in the first log line set according to the log keyword to obtain a second log line set, and perform up-and-down line complementing processing on each log line in the second log line set to generate a log block corresponding to each log line. Subsequently, use the log content in the log block to construct a model input prompt, and input the model input prompt into the target model. After being processed by the target model, the root cause analysis result corresponding to the target fault event is output, where the root cause analysis result includes the fault cause of the target fault event.
[0177] After the embodiments of the present disclosure form the first log line set based on the to-be-processed log file corresponding to the target fault event, filter the log lines in the first log line set based on the log keyword, perform up-and-down line complementing processing to generate a log block, and use the log block to construct a model input prompt to input the target model to generate the root cause analysis result corresponding to the target fault event. It can be seen that the embodiments of the present disclosure can automatically generate the root cause analysis result including the fault cause of the target fault event based on the to-be-processed log file, improving the fault diagnosis efficiency of the target fault event so that the user can timely understand the fault cause.
[0178] The log file processing device provided by the embodiments of the present disclosure can execute the log file processing method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0179] To implement the above embodiments, the present disclosure also proposes a computer program product, including a computer program / instructions, which when executed by a processor, implement the log file processing method in the above embodiments.
[0180] In addition to the above methods and devices, the embodiments of the present disclosure also provide a computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device implements the log file processing method described in the embodiments of the present disclosure.
[0181] In addition, the embodiments of the present disclosure also provide a log file processing device, as shown in Figure 5 It may include:
[0182] A processor 501, a memory 502, an input device 503, and an output device 504. The number of processors 501 in the log file processing device may be one or more. Figure 5 Taking one processor as an example. In some embodiments of the present disclosure, the processor 501, the memory 502, the input device 503, and the output device 504 may be connected by a bus or other means. Among them, Figure 5 Taking the connection by bus as an example.
[0183] The memory 502 can be used to store software programs and modules. The processor 501 runs the software programs and modules stored in the memory 502, thereby executing various functional applications and data processing of the log file processing device. The memory 502 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The input device 503 can be used to receive input digital or character information, and generate signal inputs related to the user settings and function controls of the log file processing device.
[0184] Specifically in this embodiment, the processor 501 will load the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 will run the application programs stored in the memory 502, thereby implementing various functions of the above log file processing device.
[0185] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0186] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A log file processing method, characterized in that: include: Obtaining log lines in a to-be-processed log file corresponding to a target fault event to form a first log line set; wherein the log lines contain log keywords, and the log keywords are used to characterize the degree of association between the log lines and the target fault event; Filtering the log lines in the first log line set according to the log keyword to obtain a second log line set, and performing uplink and downlink completion processing on each log line in the second log line set to generate log blocks corresponding to each log line; A model input prompt word is constructed using the log content in the log block, and the model input prompt word is input into the target model. After being processed by the target model, a root cause analysis result corresponding to the target fault event is output; wherein the root cause analysis result includes the fault cause of the target fault event.
2. The method according to claim 1, characterized in that After performing upstream and downstream completion processing on each log line in the second log line set to generate log blocks corresponding to each log line, the method further includes: Determine a weight corresponding to each log line in the log block; wherein the weight is determined based on a log keyword of the log line; Performing weighted calculation processing on each log line in the log block to obtain the weight of the log block; wherein the weight of the log block is used to represent the degree of association between the log content in the log block and the target fault event; Accordingly, the step of constructing a model input prompt word by using the log content in the log block includes: If the total character length of each log block is greater than the character limit length of the model input prompt word, the model input prompt word is constructed using the log content in the log block whose weight is greater than the preset weight threshold.
3. The method according to claim 2, characterized in that If the total character length of each log block is greater than the character limit length of the model input prompt word, the log content in the log block with a weight greater than a preset weight threshold is used to construct the model input prompt word, including: If the total character length of each log block is greater than the character limit length of the model input prompt word, the model input prompt word is constructed using the log content in the log block whose weight is greater than the preset weight threshold and the log line at the tail of the second log line set.
4. The method according to claim 1, characterized in that: The filtering the log lines in the first log line set according to the log keyword to obtain the second log line set includes: The log keywords of each log line in the first log line set are matched with a preset keyword set respectively, and a second log line set is constructed based on the log lines corresponding to the successfully matched log keywords; wherein the preset keyword set includes log keywords whose correlation with the target fault event meets preset correlation conditions.
5. The method according to claim 1, characterized in that Before performing upstream and downstream completion processing on the log lines in the second log line set to generate log blocks corresponding to the log lines, the method further includes: Match each log line in the second log line set with each log template in the log template set, and construct a third log line set based on the log lines that are not successfully matched with any log template; wherein the log templates in the log template set are obtained by analyzing a standard log file, and the standard log file is a log file that has a corresponding relationship with the log file to be processed and does not contain the target fault event; Correspondingly, performing upstream and downstream completion processing on the log lines in the second log line set to generate log blocks corresponding to the log lines includes: Context completion processing is performed on the log lines in the third log line set to generate log blocks corresponding to the log lines.
6. The method according to claim 5, characterized in that The matching each log line in the second log line set with each log template in the log template set includes: Determine an edit distance between a first log line in the second log line set and a first log template in the log template set, and determine that the first log line successfully matches the first log template when a ratio between the edit distance and a total character length of the first log template is less than a preset first threshold; and / or, Obtain the non-overlapping character length between the first log line in the second log line set and the first log template in the log template set, and determine that the first log line successfully matches the first log template when the ratio between the non-overlapping character length and the total character length of the first log template is less than a preset second threshold; wherein the non-overlapping character length is the character length in the first log line that did not successfully match the first log template.
7. The method according to claim 6, characterized in that Before obtaining the non-overlapping character length between the first log line in the second log line set and the first log template in the log template set, the method further includes: Performing word segmentation processing on the first log line in the second log line set and the first log template in the log template set respectively, to obtain a first log line after word segmentation and a first log template after word segmentation; Correspondingly, the step of obtaining the non-overlapping character length between the first log line in the second log line set and the first log template in the log template set includes: Character matching is performed on the first log line after the word segmentation and the first log template after the word segmentation to obtain non-overlapping character lengths.
8. The method according to claim 5, characterized in that The performing context completion processing on the log lines in the third log line set respectively to generate log blocks corresponding to the log lines includes: Determine the position of the log line in the third log line set in the to-be-processed log file, and obtain the log lines of the adjacent N preceding and M following the log line based on the position; M and N are preset integers; A log block of the log line is generated based on the log line and the adjacent first N lines and the adjacent last M lines.
9. The method according to claim 1, characterized in that: The root cause analysis result also includes log line information corresponding to the fault cause, and the log line information is used to obtain log content corresponding to the fault cause in the log file to be processed.
10. A log file processing device, characterized in that: The device comprises: An acquisition module, used to acquire log lines in a to-be-processed log file corresponding to a target fault event, to form a first log line set; wherein the log lines contain log keywords, and the log keywords are used to characterize the degree of association between the log lines and the target fault event; a filtering and completion module, configured to filter the log lines in the first log line set according to the log keyword to obtain a second log line set, and perform uplink and downlink completion processing on each log line in the second log line set to generate a log block corresponding to each log line; A construction module is used to construct a model input prompt word using the log content in the log block, and input the model input prompt word into the target model, and output the root cause analysis result corresponding to the target fault event after being processed by the target model; wherein the root cause analysis result includes the fault cause of the target fault event.
11. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the log file processing method described in any one of claims 1-9 above.
12. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the log file processing method described in any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product comprises a computer program / instruction, and when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 9 is implemented.