Log analysis rule determination method and related equipment

The semantic vectors of the log are extracted through the language model, and compared similar example logs, and log analysis rules are derived, which solves the problem that some logs cannot be parsed in the existing technology, and improves the efficiency and accuracy of log parsing.

CN120011529APending Publication Date: 2025-05-16BEIJING HONGTENG INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311508990.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Direct use of built-in parsing rules packages in the prior art may not be able to parse part of the log, resulting in a decrease in log parsing efficiency and accuracy.

Method used

By inputting the sample log library into the language model, the semantic vector library is extracted, and the log to be detected is input into the language model, and the target semantic vector is extracted. Then, by comparing the semantic vector library and the target semantic vector, similar sample logs are found, and feasible log parsing rules are derived.

Benefits of technology

Improves the efficiency, success rate and accuracy of log parsing, and can parse logs that cannot be processed by built-in parsing rule packages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011529A_ABST
    Figure CN120011529A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a log analysis rule determination method and related equipment. The log analysis rule determination method comprises the following steps: inputting a sample log library into a language model to obtain a semantic vector library; inputting the to-be-detected log into the language model to obtain a target semantic vector; according to the semantic vector library and the target semantic vector, a similar sample result is obtained, and the similar sample result comprises similar samples from the sample log library; and obtaining a log analysis rule result according to the similar sample result. According to the technical scheme, after the similar sample logs are found to form the similar sample result, the feasible log analysis rule can be deduced and determined according to the similar sample in the similar sample result, so that the problem that part of logs cannot be analyzed by directly using a built-in analysis rule packet is solved, the log analysis efficiency is improved, and the user experience is improved. And the success rate and accuracy of log analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer and communication technology, and in particular to a log parsing rule determination method and related equipment. Background Art

[0002] With the gradual deepening of digital industrialization, industrial digitization, and digital twins, the world is moving towards a digital future defined by software, networked, and data-driven. The types and number of various basic security facilities in enterprises continue to grow, but the increasingly urgent information system audits and internal controls, as well as the increasing business continuity needs of enterprises, also require more and more products and devices that generate logs, and more and more log deformation and packaging. Based on the diversity and complexity of these on-site environments, directly using the built-in parsing rule package may not be able to parse some logs. Summary of the invention

[0003] The embodiments of the present application provide a log parsing rule determination method and related devices, thereby overcoming the problem in the prior art that directly using a built-in parsing rule package may fail to parse part of the logs, at least to a certain extent.

[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0005] According to one aspect of an embodiment of the present application, a method for determining a log parsing rule is provided, comprising: inputting a sample log library into a language model to obtain a semantic vector library; inputting a log to be detected into a language model to obtain a target semantic vector; obtaining similar sample results based on the semantic vector library and the target semantic vector, wherein the similar sample results include similar samples from the sample log library; and obtaining a log parsing rule result based on the similar sample results.

[0006] In an embodiment of the present application, the step of inputting the log to be detected into a language model to obtain a target semantic vector specifically includes: preprocessing the log to be detected to obtain a preprocessed text; and inputting the preprocessed text into the language model to obtain a target semantic vector.

[0007] In an embodiment of the present application, the preprocessing of the log to be detected to obtain a preprocessed text specifically includes: converting the log to be detected into a text in a natural language form to obtain a preliminary text; and generalizing the content of the preliminary text to obtain a preprocessed text.

[0008] In an embodiment of the present application, the inputting the preprocessed text into the language model to obtain the target semantic vector specifically includes: if the length of the preprocessed text is greater than a predetermined length, dividing the preprocessed text into multiple text segments according to a predetermined window size and a predetermined moving step, and the predetermined window size is greater than the predetermined moving step; inputting the multiple text segments into the language model one by one to obtain the target semantic vector.

[0009] In an embodiment of the present application, obtaining similar sample results based on the semantic vector library and the target semantic vector specifically includes: comparing the target semantic vector with the sample log vectors in the semantic vector library one by one to obtain a similarity result; and obtaining similar sample results based on the similarity result.

[0010] In an embodiment of the present application, obtaining a similar sample result according to the similarity result specifically includes: sorting the sample log vectors according to the similarity result to obtain a sorting result; and obtaining the similar sample result according to the sorting result.

[0011] In an embodiment of the present application, the log parsing rule determination method also includes: obtaining a set of log sample pairs, the set of log sample pairs including multiple log sample pairs, each of the log sample pairs including two log samples, and marked with a label of similarity; inputting the log sample pairs into the language model to obtain sample semantic vectors corresponding to the two log samples respectively; obtaining a result of similarity based on the sample semantic vectors corresponding to the two log samples; and updating the parameters of the language model based on the result of similarity output by the language model and the label of similarity marked until a predetermined condition is met, stopping training, and obtaining a trained language model.

[0012] According to one aspect of an embodiment of the present application, a log parsing rule determination device is provided, and the log parsing rule determination device includes: a first input module, used to input a sample log library into a language model to obtain a semantic vector library; a second input module, used to input a log to be detected into a language model to obtain a target semantic vector; a vector comparison module, used to obtain similar sample results based on the semantic vector library and the target semantic vector, and the similar sample results include similar samples from the sample log library; a result determination module, used to obtain a log parsing rule result based on the similar sample results.

[0013] In an embodiment of the present application, the second input module specifically includes: a preprocessing submodule, used to preprocess the log to be detected to obtain a preprocessed text; and a semantic vector submodule, used to input the preprocessed text into the language model to obtain a target semantic vector.

[0014] In an embodiment of the present application, the preprocessing submodule specifically includes: a preliminary text unit, which is used to convert the log to be detected into a text in a natural language form to obtain a preliminary text; and a preprocessing text unit, which is used to generalize the content of the preliminary text to obtain a preprocessed text.

[0015] In an embodiment of the present application, the semantic vector submodule specifically includes: a sliding window segmentation unit, which is used to segment the preprocessed text into multiple text segments according to a predetermined window size and a predetermined moving step if the length of the preprocessed text is greater than a predetermined length, and the predetermined window size is greater than the predetermined moving step; a one-by-one input unit, which is used to input the multiple text segments into the language model one by one to obtain a target semantic vector.

[0016] In an embodiment of the present application, the vector comparison module specifically includes: a similarity determination submodule, used to compare the target semantic vector with the sample log vectors in the semantic vector library one by one to obtain a similarity result; a similar sample submodule, used to obtain a similar sample result based on the similarity result.

[0017] In an embodiment of the present application, the similar sample middle module specifically includes: a sorting result unit, used to sort the sample log vector according to the similarity result to obtain a sorting result; a similar sample unit, used to obtain the similar sample result according to the sorting result.

[0018] In an embodiment of the present application, the log parsing rule determination device further includes:

[0019] The sample pair acquisition module is used to obtain a set of log sample pairs, wherein the set of log sample pairs includes multiple log sample pairs, each of which includes two log samples and is marked with a label of similarity; the sample input module is used to input the log sample pairs into the language model to obtain sample semantic vectors corresponding to the two log samples respectively; the result judgment module is used to obtain a result of similarity based on the sample semantic vectors corresponding to the two log samples; the parameter updating module is used to update the parameters of the language model based on the result of similarity output by the language model and the label of similarity marked, until a predetermined condition is met, and the training is stopped to obtain a trained language model.

[0020] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the log parsing rule determination method as described in the above embodiment is implemented.

[0021] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the log parsing rule determination method as described in the above embodiments.

[0022] In the technical solutions provided in some embodiments of the present application, the semantic vector of each sample log in the configured sample log library is extracted through a language model to form a semantic vector library, and then the semantic vector of the log to be detected, that is, the target semantic vector, is extracted through the above-mentioned language model. By comparing the target semantic vector with the semantic vector of each sample log in the sample log library, similar sample logs are found to form similar sample results. Based on the similar samples in the similar sample results, feasible log parsing rules can be derived and determined, thereby solving the problem that directly using the built-in parsing rule package may not be able to parse some logs, improving the efficiency of log parsing, and improving the success rate and accuracy of log parsing.

[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0025] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown.

[0026] Figure 2 A flow chart of a log parsing rule determination method provided in an embodiment of the present application is shown.

[0027] Figure 3 Shown according to Figure 2 A specific implementation flowchart of step S200 in the log parsing rule determination method shown in the corresponding embodiment.

[0028] Figure 4 Shown according to Figure 2 A specific implementation flowchart of step S300 in the log parsing rule determination method shown in the corresponding embodiment.

[0029] Figure 5A schematic diagram of the structure of a log parsing rule determination device provided in an embodiment of the present application is shown.

[0030] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0032] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the present application.

[0033] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0034] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0035] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown.

[0036] like Figure 1 As shown, the system architecture may include terminal devices (such as Figure 1 The embodiment of the present invention is a schematic diagram of a mobile phone 101, a tablet computer 102, a portable computer 103, a desktop computer, etc., a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as a wired communication link, a wireless communication link, etc.

[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.

[0038] The user can use the terminal device to interact with the server 105 through the network 104 to receive or send messages, etc. The server 105 can be a server that provides various services. For example, the user uses the terminal device 103 (or the terminal device 101 or 102) to upload the sample log library and the log to be detected to the server 105. The server 105 can input the sample log library into the language model to obtain a semantic vector library; input the log to be detected into the language model to obtain a target semantic vector; obtain similar sample results based on the semantic vector library and the target semantic vector, and the similar sample results include similar samples from the sample log library; obtain log parsing rule results based on the similar sample results.

[0039] It should be noted that the log parsing rule determination method provided in the embodiment of the present application is generally executed by the server 105, and accordingly, the log parsing rule determination device is generally set in the server 105. However, in other embodiments of the present application, the terminal device may also have similar functions as the server, so as to execute the log parsing rule determination scheme provided in the embodiment of the present application.

[0040] The implementation details of the technical solution of the embodiment of the present application are described in detail below:

[0041] Figure 2 A flow chart of a method for determining a log parsing rule according to an embodiment of the present application is shown. The method for determining a log parsing rule can be executed by a server, which can be Figure 1 Refer to the server shown in Figure 2 As shown, the log parsing rule determination method at least includes:

[0042] In step S100, the sample log library is input into the language model to obtain a semantic vector library.

[0043] In step S200, the log to be detected is input into the language model to obtain a target semantic vector.

[0044] In step S300, similar sample results are obtained according to the semantic vector library and the target semantic vector, and the similar sample results include similar samples from the sample log library.

[0045] In step S400, a log parsing rule result is obtained according to the similar sample result.

[0046] In an embodiment of the present application, a semantic vector of each sample log in a configured sample log library is extracted through a language model to form a semantic vector library, and then a semantic vector of a log to be detected, i.e., a target semantic vector, is extracted through the above-mentioned language model. By comparing the target semantic vector with the semantic vector of each sample log in the sample log library, similar sample logs are found to form similar sample results. Based on similar samples in the similar sample results, feasible log parsing rules can be derived and determined, thereby solving the problem that some logs may not be parsed when the built-in parsing rule package is used directly, thereby improving the efficiency of log parsing, and improving the success rate and accuracy of log parsing.

[0047] In step S100, the sample log corresponding to each parsing rule is obtained and stored in association to form a sample log library. The sample logs in the sample log library are input into the language model one by one to obtain the corresponding semantic vector, i.e., the sample log vector. Each sample log vector corresponds to the parsing rule of the corresponding sample log, and the semantic vectors of all sample logs together form a semantic vector library.

[0048] The language model processes the sample logs in the sample log library in the same way as the logs to be detected, so please refer to the detailed description of step S200 below. The specific processing steps of the language model will be detailed in step S200.

[0049] In step S200, the log to be detected is parsed by a language model to obtain a target semantic vector, and the text of the log to be detected is vectorized to find similar examples through the vectorized target semantic vector, thereby determining the corresponding log parsing rules.

[0050] Specifically, in some embodiments, the specific implementation of step S200 can be found in Figure 3 . Figure 3 is based on Figure 2 The detailed description of step S200 in the log parsing rule determination method shown in the corresponding embodiment, in the log parsing rule determination method, step S200 may include the following steps:

[0051] Step S210, preprocessing the log to be detected to obtain a preprocessed text.

[0052] Step S220: input the preprocessed text into the language model to obtain a target semantic vector.

[0053] In an embodiment of the present application, before inputting the log text to be detected into the language model, the log to be detected needs to be preprocessed, converted into preprocessed text in natural language form, and then input into the language model to obtain a target semantic vector.

[0054] In step S210, the log to be detected is preprocessed and converted into preprocessed text in natural language form for language model processing.

[0055] Specifically, in some embodiments, the specific implementation of step S210 can refer to the following embodiments. Figure 3 The detailed description of step S210 in the log parsing rule determination method shown in the corresponding embodiment, in the log parsing rule determination method, step S210 may include the following steps:

[0056] The log to be detected is converted into a text in a natural language form to obtain a preliminary text.

[0057] The preliminary text is generalized to obtain a preprocessed text.

[0058] In this embodiment, logs are usually structured or semi-structured texts. Structured log text refers to log text that strictly complies with standard formats such as json and xml. Semi-structured log text refers to log text that does not fully comply with the standard (that is, it cannot be directly parsed) but still has structural information, such as log text that contains some natural language descriptions and structured information.

[0059] In this embodiment, the preprocessing of the log to be detected includes natural language form conversion and general content generalization processing.

[0060] Specifically, the specific steps of natural language form conversion include: for the multi-layer nested relationship of JSON, load it into an ordered dictionary to try to eliminate the structural information that seems different in form but is essentially the same (such as the order change of multiple keys at the same level), thereby eliminating the impact of form changes on semantics. At the same time, it can also retain the valid structural information in the log, such as the parent-child relationship.

[0061] For all structured texts, the structured dictionary is reorganized into natural language, such as key=value text is converted to key is value; {a:[b,c]} text is converted to a contains b and c.

[0062] The generalization of common content can specifically include generalizing information that can be ignored in practice (i.e. useless information) such as hash, IP, and time strings that are usually found in logs into fixed strings using regular expression technology. For example, replace "2017-07-17T18:35:43.683+0000" with "datetime", "192.168.0.1" with "ip", etc., so that the model can better focus on truly useful semantic information in subsequent semantic matching.

[0063] It can be understood that in step S100, the sample log library also needs to be preprocessed as described above before being input into the language model.

[0064] In step S220, the preprocessed log to be detected is input into the language model to obtain a target semantic vector.

[0065] In this embodiment, the language model can be obtained based on the best transformer model currently used for natural language processing tasks (such as BERT, RoBERTa, DistilBERT, ALBERT, XLNet, etc.). That is, in this embodiment, the transformer model is first selected for pre-training on a large amount of public corpus, so that it can extract the semantic vector that performs best in various natural language processing tasks such as text similarity. Then, it is used on a large number of labeled (similar or dissimilar) sentence pairs (sentence-pairs) data sets, and fine-tuned for the semantic similarity of the sentences to extract the semantic vector that currently performs best in the semantic similarity of the sentences.

[0066] Specifically, in some embodiments of the present application, the above-mentioned semantic model training method may include:

[0067] A log sample pair set is obtained, where the log sample pair set includes a plurality of log sample pairs, each of which includes two log samples and is marked with a label indicating whether they are similar.

[0068] The log sample pair is input into the language model to obtain sample semantic vectors corresponding to the two log samples respectively.

[0069] A result of whether the two log samples are similar is obtained based on the sample semantic vectors corresponding to the two log samples.

[0070] According to the similarity results output by the language model and the similarity labels marked, the parameters of the language model are updated until a predetermined condition is met, and the training is stopped to obtain a trained language model.

[0071] In this embodiment, in order to analyze and learn the log text, the convolutional neural network can well capture the semantic features contained in the log text itself, mine the correlation of the text in time and space, and has strong generalization ability and robustness. When training, the conditions for stopping training, that is, the conditions for model training, can be multiple, and the specific embodiments can be referred to as follows.

[0072] Specifically, in some embodiments, the updating of parameters of the language model according to the similarity results output by the language model and the similarity labels marked, until a predetermined condition is met, the training is stopped, and a trained language model is obtained, specifically including:

[0073] If, in the set of log sample pairs, there are less than a predetermined number of log sample pairs input into the language model and the classification results outputted are consistent with the classification labels, the parameters of the language model are updated.

[0074] If, in the log sample pair set, more than a predetermined number of log sample pairs are input into the language model and the similarity results outputted are consistent with the similarity labels, then the predetermined end condition is met, the training is ended, and a trained language model is obtained.

[0075] Specifically, in some other embodiments, the updating of parameters of the language model according to the similarity results output by the language model and the similarity labels marked, until a predetermined condition is met, the training is stopped, and a trained language model is obtained, specifically including:

[0076] A loss function is determined according to labels of similarity of the log sample pairs and corresponding similarity results.

[0077] The parameters of the language model are updated according to the loss function until a predetermined end condition is reached, and the training is terminated to obtain a trained language model.

[0078] In this embodiment, the loss function reaching a predetermined end condition may be that the loss function converges or the loss function is less than a predetermined loss (eg, 0.001).

[0079] In the above embodiment, log text recognition is performed through a neural network model obtained through multiple trainings to obtain corresponding judgment results. The neural network model is obtained through multiple trainings. The more samples it trains, the more accurate the results obtained. It basically does not require maintenance during operation, which also reduces maintenance costs, improves the efficiency and accuracy of traffic identification, and reduces the false alarm rate.

[0080] Specifically, in some embodiments, the specific implementation of step S220 can refer to the following embodiments. Figure 3 The detailed description of step S220 in the log parsing rule determination method shown in the corresponding embodiment, in the log parsing rule determination method, step S220 may include the following steps:

[0081] If the length of the preprocessed text is greater than a predetermined length, the preprocessed text is divided into a plurality of text segments according to a predetermined window size and a predetermined moving step, and the predetermined window size is greater than the predetermined moving step.

[0082] The multiple text segments are input into the language model one by one to obtain a target semantic vector.

[0083] In the embodiments of the present application, the sliding window segmentation method plus the similarity average algorithm effectively solves the problem that the pre-processed text is too long to be processed, or even if it is processed, the accuracy is not high.

[0084] In some usage scenarios, the on-site log may be forwarded multiple times, and each forwarding adds additional information, resulting in the log text being too long. However, there is a limit on the length of the text input processed by the model, for example, token_max_length = 512. If a log text exceeds 512 in length, it will be truncated. At the same time, because forwarding usually adds information at the front of the original log, the valid information (that is, the information that can hit the sample log) will be gradually moved back, causing the model to fail to hit the sample log.

[0085] Therefore, in this embodiment, if the preprocessed log text to be detected is too long and exceeds the predetermined length, the sliding window segmentation method is enabled to segment the original preprocessed text according to the window size and the moving step length to obtain multiple text segments, and each text segment directly has an overlapping part (realized by the predetermined window size being greater than the predetermined moving step length) to retain semantic information to the greatest extent through the overlapping part, thereby solving the problem of being unable to determine the best segmentation position, and the best segmentation position is the position that can best retain semantic information. Then, multiple text segments are input into the language model one by one to obtain the target semantic vector.

[0086] For example, in one embodiment, for the text '0123456789', its text length is input len =10, assuming the preset window size window size =5, the predetermined moving step length is move step =3, the segmentation results are '01234', '34567', '6789'. Then the text fragments '01234', '34567', '6789' are input into the language model one by one to obtain the target semantic vector In step S300, Compare with the vectors in the semantic vector library to obtain similar sample results.

[0087] In the above embodiment, when the length of the last remaining segment is less than the window size, if the absolute value of the difference between the length of the last remaining segment and the window size is less than a predetermined value (for example, less than a predetermined number of characters or less than a predetermined ratio), the remaining segment is directly used as the last text segment.

[0088] For example, in another embodiment, for the text 'qwertyuiopasdfghjkl', the text length is input len =19, assuming the preset window size window size =11, preset moving step length move step =6, the segmentation results are 'qwertyuiopa', 'uiopasdfghj', 'opasdfghjkl'. Then the text fragments 'qwertyuiopa', 'uiopasdfghj', 'opasdfghjkl' are input into the language model one by one to obtain the target semantic vector In step S300, Compare with the vectors in the semantic vector library to obtain similar sample results.

[0089] In the above embodiment, when the length of the last remaining fragment is less than the window size, if the absolute value of the difference between it and the window size exceeds a predetermined value (for example, exceeds a predetermined number of characters or is less than a predetermined ratio), the last character of the entire preprocessed text is taken as the starting point, and characters with the same position as the number of bits of the window size are taken forward to form the last text fragment, and the character order of the text fragment is consistent with the original preprocessed text.

[0090] In step S300, by comparing the target semantic vector with the semantic vector of each sample log in the sample log library, similar sample logs are found to form similar sample results.

[0091] Specifically, in some embodiments, the specific implementation of step S300 can be found in Figure 4 . Figure 4 is based on Figure 2 The detailed description of step S300 in the log parsing rule determination method shown in the corresponding embodiment, in the log parsing rule determination method, step S300 may include the following steps:

[0092] Step S310: Compare the target semantic vector with the sample log vectors in the semantic vector library one by one to obtain a similarity result.

[0093] Step S320: obtaining similar sample results according to the similarity results.

[0094] In this embodiment, the target semantic vector is first compared with the sample log vectors in the semantic vector library one by one to obtain corresponding similarity results, and then based on the similarity results, similar sample results are obtained.

[0095] In step S310, the target semantic vector is compared with the sample log vectors in the semantic vector library one by one, and the similarity between the target semantic vector and each sample log vector is obtained, thereby obtaining a similarity result. The calculated similarity result can be cosine similarity, that is, the target semantic vector is extracted. Then, traverse each sample log vector in the semantic vector library Calculate the cosine similarity between the two and get the similarity score between them

[0096] In some embodiments, to improve computing performance, all sample log vectors in the semantic vector library can be packaged into a matrix and then calculated at one time. If the semantic vector library is loaded as a matrix M, the matrix score can be obtained. Matrix score Score matrix is a matrix containing the target semantic vector The similarity score between the log vectors and all the sample log vectors in the semantic vector library.

[0097] In some other embodiments, the similarity result may also be a similarity level, which is determined according to the magnitude of the cosine similarity. The similarity level is different in different value intervals of the cosine similarity.

[0098] When the log text to be detected is too long, the embodiment of sliding window segmentation and similarity averaging is adopted. After obtaining the semantic vector of each segmented text segment After that, we first calculate the similarity between the semantic vector of each text segment and the sample log vector, and get the following formula:

[0099]

[0100] Then, according to the similarity between the semantic vector of each text segment and the sample log vector, the average similarity Score is calculated. avg =(Score1+Score2+…+Score n ) / n, as the similarity between the original log to be detected and the sample log.

[0101] When all sample log vectors in the semantic vector library are packaged into a matrix and calculated at one time, the similarity between the semantic vector of each text segment and the matrix M is calculated separately, that is, the following formula is obtained:

[0102]

[0103] Then, according to the similarity between the semantic vectors of the above text fragments and the matrix M, the average similarity is calculated. As the similarity between the original log to be detected and each sample log in the sample log library.

[0104] In step S320, similar sample logs are screened in the sample log library according to the similarity result, so as to obtain similar sample results.

[0105] Specifically, in some embodiments, the specific implementation of step S320 can refer to the following embodiments. Figure 4 The detailed description of step S320 in the log parsing rule determination method shown in the corresponding embodiment, in the log parsing rule determination method, step S320 may include the following steps:

[0106] The sample log vectors are sorted according to the similarity result to obtain a sorting result.

[0107] According to the sorting result, the similar sample result is obtained.

[0108] In the embodiment of the present application, the sample log vectors are first sorted according to the similarity results, and after the sorting results are obtained, the most similar pre-positioned sample logs are selected according to the sorting results as similar samples to form similar sample results. For example, the top 1, top 3, top 5, and top 10 sample logs are selected as similar samples to form similar sample results.

[0109] In step S400, the parsing rules corresponding to the similar samples contained in the similar sample results are very likely to be the parsing rules required by the staff. The parsing rules corresponding to these similar sample results are output to the user as log parsing rule results so that the user can refer to and select appropriate parsing rules. This can solve the problem that part of the logs may not be parsed when the built-in parsing rule package is used directly, thereby improving the efficiency of log parsing and improving the success rate and accuracy of log parsing.

[0110] In other embodiments, the parsing rules corresponding to all similar examples included in the similar example results may be used one by one to parse the log to be detected, and if the log to be detected is successfully parsed, the parsing rule is output as the log parsing rule result. Specifically, the parsing may be performed one by one starting from the highest similarity according to the sorting result until the log parsing rule result is obtained.

[0111] The following describes an apparatus embodiment of the present application, which can be used to execute the log parsing rule determination method in the above-mentioned embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above-mentioned embodiment of the log parsing rule determination method of the present application.

[0112] Figure 5 A block diagram of a log parsing rule determination device according to an embodiment of the present application is shown.

[0113] Reference Figure 5 As shown, according to an embodiment of the present application, a log parsing rule determination device 500 includes: a first input module 510, a second input module 520, a vector comparison module 530, and a result determination module 540. Among them, the first input module 510 is used to input the sample log library into the language model to obtain a semantic vector library; the second input module 520 is used to input the log to be detected into the language model to obtain a target semantic vector; the vector comparison module 530 is used to obtain similar sample results based on the semantic vector library and the target semantic vector, and the similar sample results include similar samples from the sample log library; the result determination module 540 is used to obtain the log parsing rule results based on the similar sample results.

[0114] In an embodiment of the present application, the second input module specifically includes: a preprocessing submodule, used to preprocess the log to be detected to obtain a preprocessed text; and a semantic vector submodule, used to input the preprocessed text into the language model to obtain a target semantic vector.

[0115] In an embodiment of the present application, the preprocessing submodule specifically includes: a preliminary text unit, which is used to convert the log to be detected into a text in a natural language form to obtain a preliminary text; and a preprocessing text unit, which is used to generalize the content of the preliminary text to obtain a preprocessed text.

[0116] In an embodiment of the present application, the semantic vector submodule specifically includes: a sliding window segmentation unit, which is used to segment the preprocessed text into multiple text segments according to a predetermined window size and a predetermined moving step if the length of the preprocessed text is greater than a predetermined length, and the predetermined window size is greater than the predetermined moving step; a one-by-one input unit, which is used to input the multiple text segments into the language model one by one to obtain a target semantic vector.

[0117] In an embodiment of the present application, the vector comparison module specifically includes: a similarity determination submodule, used to compare the target semantic vector with the sample log vectors in the semantic vector library one by one to obtain a similarity result; a similar sample submodule, used to obtain a similar sample result based on the similarity result.

[0118] In an embodiment of the present application, the similar sample middle module specifically includes: a sorting result unit, used to sort the sample log vector according to the similarity result to obtain a sorting result; a similar sample unit, used to obtain the similar sample result according to the sorting result.

[0119] In an embodiment of the present application, the log parsing rule determination device further includes:

[0120] The sample pair acquisition module is used to obtain a set of log sample pairs, wherein the set of log sample pairs includes multiple log sample pairs, each of which includes two log samples and is marked with a label of similarity; the sample input module is used to input the log sample pairs into the language model to obtain sample semantic vectors corresponding to the two log samples respectively; the result judgment module is used to obtain a result of similarity based on the sample semantic vectors corresponding to the two log samples; the parameter updating module is used to update the parameters of the language model based on the result of similarity output by the language model and the label of similarity marked, until a predetermined condition is met, and the training is stopped to obtain a trained language model.

[0121] In the technical solutions provided in some embodiments of the present application, the semantic vector of each sample log in the configured sample log library is extracted through a language model to form a semantic vector library, and then the semantic vector of the log to be detected, that is, the target semantic vector, is extracted through the above-mentioned language model. By comparing the target semantic vector with the semantic vector of each sample log in the sample log library, similar sample logs are found to form similar sample results. Based on the similar samples in the similar sample results, feasible log parsing rules can be derived and determined, thereby solving the problem that directly using the built-in parsing rule package may not be able to parse some logs, improving the efficiency of log parsing, and improving the success rate and accuracy of log parsing.

[0122] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.

[0123] It should be noted that Figure 6 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0124] like Figure 6As shown, the computer system includes a central processing unit (CPU) 1801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1802 or the program loaded from the storage part 1808 to the random access memory (RAM) 1803, such as executing the method described in the above embodiment. In RAM 1803, various programs and data required for system operation are also stored. CPU 1801, ROM 1802 and RAM 1803 are connected to each other through bus 1804. Input / output (I / O) interface 1805 is also connected to bus 1804.

[0125] The following components are connected to the I / O interface 1805: an input section 1806 including a keyboard, a mouse, etc.; an output section 1807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to the I / O interface 1805 as needed. A removable medium 1811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1810 as needed so that a computer program read therefrom is installed into the storage section 1808 as needed.

[0126] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1809, and / or installed from a removable medium 1811. When the computer program is executed by a central processing unit (CPU) 1801, various functions defined in the system of the present application are executed.

[0127] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0128] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0129] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not constitute limitations on the units themselves in some cases.

[0130] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.

[0131] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0132] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.

[0133] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.

[0134] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for determining log parsing rules, characterized in that: The log parsing rule determination method comprises: Input the sample log library into the language model to obtain the semantic vector library; Input the log to be detected into the language model to obtain the target semantic vector; Obtaining similar sample results according to the semantic vector library and the target semantic vector, wherein the similar sample results include similar samples from the sample log library; According to the similar sample results, the log parsing rule results are obtained.

2. The log parsing rule determination method according to claim 1, characterized in that: The step of inputting the log to be detected into the language model to obtain the target semantic vector specifically includes: Preprocessing the log to be detected to obtain a preprocessed text; The preprocessed text is input into the language model to obtain a target semantic vector.

3. The log parsing rule determination method according to claim 2, characterized in that: The preprocessing of the log to be detected to obtain a preprocessed text specifically includes: Convert the log to be detected into text in natural language to obtain a preliminary text; The preliminary text is generalized to obtain a preprocessed text.

4. The log parsing rule determination method according to claim 2, characterized in that: The step of inputting the preprocessed text into the language model to obtain a target semantic vector specifically includes: If the length of the preprocessed text is greater than a predetermined length, segmenting the preprocessed text into a plurality of text segments according to a predetermined window size and a predetermined moving step length, wherein the predetermined window size is greater than the predetermined moving step length; The multiple text segments are input into the language model one by one to obtain a target semantic vector.

5. The log parsing rule determination method according to claim 1, characterized in that: The obtaining of similar sample results according to the semantic vector library and the target semantic vector specifically includes: Compare the target semantic vector with the sample log vectors in the semantic vector library one by one to obtain a similarity result; According to the similarity result, a similar sample result is obtained.

6. The log parsing rule determination method according to claim 5, characterized in that: Obtaining similar sample results according to the similarity results specifically includes: According to the similarity result, the sample log vectors are sorted to obtain a sorting result; According to the sorting result, the similar sample result is obtained.

7. The log parsing rule determination method according to claim 1, characterized in that: The log parsing rule determination method also includes: Obtain a log sample pair set, the log sample pair set comprising a plurality of log sample pairs, each of the log sample pairs comprising two log samples and marked with labels of similarity; Inputting the log sample pair into the language model to obtain sample semantic vectors corresponding to the two log samples respectively; Obtaining a similarity result based on the sample semantic vectors corresponding to the two log samples; According to the similarity results output by the language model and the similarity labels marked, the parameters of the language model are updated until a predetermined condition is met, and the training is stopped to obtain a trained language model.

8. A log parsing rule determination device, characterized in that: The log parsing rule determination device comprises: The first input module is used to input the sample log library into the language model to obtain a semantic vector library; The second input module is used to input the log to be detected into the language model to obtain the target semantic vector; A vector comparison module, used for obtaining similar sample results according to the semantic vector library and the target semantic vector, wherein the similar sample results include similar samples from the sample log library; The result determination module is used to obtain the log parsing rule result according to the similar sample result.

9. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the log parsing rule determination method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: one or more processors; A storage device, used to store one or more programs, when the one or more programs are executed by the one or more processors, enables the one or more processors to implement the log parsing rule determination method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Cross-system AI log gateway intercommunication method and system

    CN121658451A