A method, device, medium and equipment for processing log data
By merging aggregation model based on cluster center and density, identifying and merging log data, the problems of high complexity and error rate of log data management in the prior art are solved, and efficient and accurate log data processing is achieved.
Patent Information
- Application Number
- CN202111272620.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-10-29
AI Technical Summary
When processing log data from multiple sources and formats, the existing log management platform has high management complexity and the probability of error increases, making it difficult to achieve efficient and accurate log data processing.
The log merging model formed by combining cluster center-based aggregation model and density-based aggregation model is used to identify the log data, determine the log type as a single-row log or multi-row log, and perform merging processing to improve the processing accuracy of log data.
Through this method, the applicability of log data processing is enhanced, the management and maintenance complexity is reduced, and the accuracy of log data processing is improved.
Smart Images

Figure CN114020715B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer and artificial intelligence technology, and in particular, to a method, device, medium and equipment for processing log data. Background Art
[0002] With the rapid development of the information age, various types of systems have emerged. Log data, which can accurately record key operating information in the system, has become an important indicator data for various systems to monitor whether errors occur during operation.
[0003] Since log data may be a single line or multiple lines, and the basis for division is often different, the error information in the log data needs to be correctly configured to merge into one line, which is convenient for developers to view and locate problems. Many log management platforms use different log merging configurations to manage log data. The current sources and formats of logs are increasing, which brings a surge in the complexity of management and maintenance of the corresponding log management platform, and the probability of log data management errors increases. Therefore, how to save and manage log data more efficiently and accurately has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] The embodiments of the present application provide a method, apparatus, medium, and device for processing log data, which can enhance the applicability of log data processing without complicated management and maintenance, and can also improve the accuracy of log data processing.
[0005] In a first aspect, an embodiment of the present application provides a method for processing log data, the method comprising:
[0006] Get the log data to be processed on the target platform;
[0007] Extract characteristic characters from each line of the log data to be processed, and construct a characteristic code of the log data to be processed;
[0008] Inputting the feature code into a log merging model pre-merged by a cluster center-based aggregation model and a density-based aggregation model;
[0009] Determine the log type of each line of the log data to be processed according to the output result of the log merging model; wherein the log type includes a single-line log type and a multi-line log type;
[0010] Multiple lines of log data are merged to obtain a processing result of the log data to be processed.
[0011] Further, multiple lines of log data are merged to obtain the processing result of the log data to be processed, including:
[0012] If it is a single-line log type, determine that the log data of the current line is complete data;
[0013] If it is a multi-line log type, the log data of the current line is merged with the log data of the multi-line log type of the adjacent line to obtain complete data.
[0014] Furthermore, the training process of the log merging model includes:
[0015] Collecting log data of at least one platform, and dividing a training set from the log data;
[0016] Extracting characteristic characters from each line of log data in the training set;
[0017] Constructing a characteristic code for each line of log data according to the characteristic characters;
[0018] Clustering the feature codes using a log merging model formed by merging a cluster center-based aggregation model and a density-based aggregation model to obtain a training feature set;
[0019] The log merging model is trained according to the recognition results of the single-line log feature set and the multi-line log feature set of the training feature set, and the pre-labeling results of the log data in the training set, so as to obtain a training result.
[0020] Furthermore, the construction process of the log merging model includes:
[0021] Determining an initial merging weight of the cluster center-based aggregation model and the density-based aggregation model;
[0022] According to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, the initial merging weight of the log merging model is adjusted to obtain the final merging weight.
[0023] Further, according to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, the initial merging weight of the log merging model is adjusted to obtain the final merging weight, including:
[0024] According to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, respectively determine a first accuracy of the cluster center-based aggregation model and a second accuracy of the density-based aggregation model;
[0025] According to the magnitude relationship between the first accuracy and the second accuracy, the initial merging weight of the log merging model is adjusted to obtain a final merging weight.
[0026] Furthermore, after obtaining the training result, the method further includes:
[0027] Dividing a test set from the log data;
[0028] Extracting characteristic characters from the log data of the test set;
[0029] Constructing a feature code of the log data according to the feature character;
[0030] Inputting the feature code of the log data of the test set into the log merging model to obtain a test feature set;
[0031] Determining whether the log merging model meets the accuracy requirement according to the recognition results of the single-line log feature set and the multi-line log feature set of the test feature set and the pre-labeling results of the log data in the training set;
[0032] If so, it is determined that the training result is available.
[0033] Furthermore, feature character extraction is performed on each line of log data in the training set, including:
[0034] According to the predetermined number of at least two characteristic characters, starting from the first characteristic character of each line of log data, a target character is determined;
[0035] Determine the target character as a characteristic character, and extract the characteristic characters of each characteristic character quantity;
[0036] Correspondingly, after obtaining the training results of the number of each characteristic character, the method further includes:
[0037] The target number of characteristic characters is determined according to the accuracy of the training results of the number of each characteristic character.
[0038] Furthermore, before determining the target number of characteristic characters according to the accuracy of the training results of the number of characteristic characters, the method further includes:
[0039] Obtain the training time factor and training computing power factor of each characteristic character quantity;
[0040] According to the accuracy of the training results of the number of each characteristic character, the target number of characteristic characters is determined, including:
[0041] The target number of characteristic characters is determined according to the accuracy of the training results of the number of characteristic characters, the training time factor and the training computing power factor.
[0042] In a second aspect, an embodiment of the present application provides a log data processing device, the device comprising:
[0043] The module for obtaining the log data to be processed is used to obtain the log data to be processed of the target platform;
[0044] A feature code construction module, used to extract feature characters from each line of the log data to be processed, and construct a feature code for the log data to be processed;
[0045] A feature code input module, used for inputting the feature code into a log merging model pre-merged by a cluster center-based aggregation model and a density-based aggregation model;
[0046] A log type output module, used to determine the log type of each line of the log data to be processed according to the output result of the log merging model; wherein the log type includes a single-line log type and a multi-line log type;
[0047] The log data processing module is used to merge and process multiple lines of log data to obtain the processing result of the log data to be processed.
[0048] Furthermore, the log data processing module is specifically used to:
[0049] If it is a single-line log type, determine that the log data of the current line is complete data;
[0050] If it is a multi-line log type, the log data of the current line is merged with the log data of the multi-line log type of the adjacent line to obtain complete data.
[0051] Furthermore, the device also includes a log merging model training module, which is used to:
[0052] Collecting log data of at least one platform, and dividing a training set from the log data;
[0053] Extracting characteristic characters from each line of log data in the training set;
[0054] Constructing a characteristic code for each line of log data according to the characteristic characters;
[0055] Clustering the feature codes using a log merging model formed by merging a cluster center-based aggregation model and a density-based aggregation model to obtain a training feature set;
[0056] The log merging model is trained according to the recognition results of the single-line log feature set and the multi-line log feature set of the training feature set, and the pre-labeling results of the log data in the training set, so as to obtain a training result.
[0057] Furthermore, the device also includes a log merging model building module, including:
[0058] An initial merging weight determining unit, used to determine the initial merging weights of the cluster center-based aggregation model and the density-based aggregation model;
[0059] The final merging weight calculation unit is used to adjust the initial merging weight of the log merging model to obtain the final merging weight according to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set.
[0060] Furthermore, the final combined weight calculation unit is specifically used for:
[0061] According to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, respectively determine a first accuracy of the cluster center-based aggregation model and a second accuracy of the density-based aggregation model;
[0062] According to the magnitude relationship between the first accuracy and the second accuracy, the initial merging weight of the log merging model is adjusted to obtain a final merging weight.
[0063] Furthermore, the device also includes a log merging model testing module, which is used to:
[0064] Dividing a test set from the log data;
[0065] Extracting characteristic characters from the log data of the test set;
[0066] Constructing a feature code of the log data according to the feature character;
[0067] Inputting the feature code of the log data of the test set into the log merging model to obtain a test feature set;
[0068] Determining whether the log merging model meets the accuracy requirement according to the recognition results of the single-line log feature set and the multi-line log feature set of the test feature set and the pre-labeling results of the log data in the training set;
[0069] If so, it is determined that the training result is available.
[0070] Furthermore, the log merging model training module is also used for:
[0071] According to the predetermined number of at least two characteristic characters, starting from the first characteristic character of each line of log data, a target character is determined;
[0072] Determine the target character as a characteristic character, and extract the characteristic characters of each characteristic character quantity;
[0073] Correspondingly, after obtaining the training results of the number of each characteristic character, the method further includes:
[0074] The target number of characteristic characters is determined according to the accuracy of the training results of the number of each characteristic character.
[0075] Furthermore, the device also includes a feature character quantity selection module, which is used to:
[0076] Obtain the training time factor and training computing power factor of each characteristic character quantity;
[0077] According to the accuracy of the training results of the number of each characteristic character, the target number of characteristic characters is determined, including:
[0078] The target number of characteristic characters is determined according to the accuracy of the training results of the number of characteristic characters, the training time factor and the training computing power factor.
[0079] In a third aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for processing log data as described in the embodiment of the present application is implemented.
[0080] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for processing log data as described in the embodiment of the present application is implemented.
[0081] The technical solution provided in the embodiment of the present application uses a log merging model formed by merging two aggregation models to identify each line of log, and can accurately determine whether the log type is a single-line log or a multi-line log, and then merge the multi-line logs to obtain the final log processing result. By implementing this solution, the applicability of log data processing can be enhanced without complex management and maintenance, and the accuracy of log data processing can also be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is a flow chart of the method for processing log data provided in Example 1 of the present application;
[0083] Figure 2 is a flow chart of the training process of the log merging model provided in the second embodiment of the present invention;
[0084] Figure 3 is a structural block diagram of a log data processing device provided by Embodiment 3 of the present invention;
[0085] Figure 4 It is a structural schematic diagram of an electronic device provided in Example 5 of the present application. DETAILED DESCRIPTION
[0086] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only the parts related to the present application, rather than all structures, are shown in the accompanying drawings.
[0087] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0088] Embodiment 1
[0089] Figure 1 It is a flowchart of the log data processing method provided in Example 1 of the present application. This embodiment can be applicable to the scenario of processing log data. The method can be executed by the log data processing device provided in the embodiment of the present application. The device can be implemented by software and / or hardware and can be integrated into an electronic device for managing a connection pool or managing a database.
[0090] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0091] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0092] like Figure 1 As shown, the method for processing log data includes:
[0093] S110, obtaining log data to be processed of the target platform.
[0094] In this solution, the target platform can be any platform that generates log data, and the log data to be processed can be multiple log data generated by the platform. Since there are many types of log data, and the content of each log data record may be different, it is necessary to process a large amount of log data generated by the target platform and determine the single-line logs and multi-line logs therein to form each complete log data. Here, constructing complete log data is the basis for managing log data. Otherwise, it is easy to mistakenly identify a single-line log as a multi-line log and merge it with other multi-line logs, which will cause the content of the log data to be repeated and the information of the log data cannot be accurately read. If a multi-line log is mistakenly identified as a single-line log, the content of the log data will be lost, and the information of the log data will not be accurately read.
[0095] In this solution, the acquisition method may be through batch transmission, real-time synchronization, etc. Specifically, the log data may be transmitted through a data transmission interface provided by the target platform.
[0096] It is understandable that the execution subject of this solution can be any electronic device with log processing capability. The electronic device can be part or all of the computing devices of the log management platform. The electronic device can be used to obtain and process log data to obtain the processing results of the log data.
[0097] S120, extracting characteristic characters from each line of the log data to be processed, and constructing a characteristic code of the log data to be processed.
[0098] The log data is recorded in the form of lines. If the log data is a single-line log data, one line records a complete log data. If the log data is a multi-line log data, a complete log data needs to be recorded in multiple lines.
[0099] Feature characters are extracted from each line of data, which may be all characters of a line of log data, or part of the characters, such as the first 20 characters, the first 50 characters, or all characters. The characters of the log data here may include numbers, letters, and other characters. Constructing the feature code of the log data to be processed may be to convert the feature characters to obtain the feature code. Here, the log data can be converted into a feature code of binary characters, and the first 20 characters can be converted to obtain a 20-bit feature code consisting of "0" and "1" for each line of log data.
[0100] S130: Input the feature code into a log merging model that is pre-merged by merging a cluster center-based aggregation model and a density-based aggregation model.
[0101] In this scheme, the log merging model can be a combination of the cluster center-based aggregation model and the density-based aggregation model. The two can be merged based on their respective weights. The cluster center algorithm is a type of KMeans clustering algorithm. Given the K value and K initial cluster center points, each point (that is, data record) is assigned to the cluster represented by the cluster center point closest to it. After all points are assigned, the center point of the cluster is recalculated based on all points in a cluster (average value is taken), and then the steps of assigning points and updating cluster center points are iterated until the change of the cluster center point is small or the specified number of iterations is reached. The main goal of the density-based clustering algorithm is to find high-density areas separated by low-density areas. Unlike the distance-based clustering algorithm, the clustering result of the distance-based clustering algorithm is a spherical cluster, while the density-based clustering algorithm can find clusters of any shape, which plays an important role for data with noise points.
[0102] S140, determining the log type of each line of the log data to be processed according to the output result of the log merging model; wherein the log type includes a single-line log type and a multi-line log type.
[0103] In this solution, the clustering result of each row of data can be determined according to the output result of the log merging model, and the log type of each current cluster can be obtained. For example, four clusters are output, two of which are single-row log types and two are multi-row log types.
[0104] S150, merging multiple lines of log data to obtain a processing result of the log data to be processed.
[0105] In this solution, since both single-line logs and multi-line logs are generated and recorded one by one, if the current line of data is a single-line log, it can be directly determined as a complete log data. If the current line is a multi-line log, it can be determined whether the previous line is multi-line. If so, it is merged with the previous line. If not, it continues to determine whether the next line is a multi-line log. If so, the next line is merged into the current line.
[0106] In one case, there may be a situation where the previous line is a single-line log, the current line is a multi-line log, and the next line is a single-line log. In this case, it can be considered that it is caused by a recognition error, because a multi-line log will occupy two or more consecutive lines. If only one line is a multi-line log, and the upper and lower adjacent lines are single-line logs, it is determined that the multi-line logs cannot be merged, resulting in a log recognition error, which can generate an error message.
[0107] In this solution, optionally, multiple lines of log data are merged to obtain a processing result of the log data to be processed, including:
[0108] If it is a single-line log type, determine that the log data of the current line is complete data;
[0109] If it is a multi-line log type, the log data of the current line is merged with the log data of the multi-line log type of the adjacent line to obtain complete data.
[0110] It is understandable that if it is a single-line log type, the log data of the current line is determined to be complete data for use in subsequent log data content identification, etc. If it is a multi-line log type, the log data of the current line is merged with the log data of the multi-line log type of the adjacent line to obtain complete data.
[0111] In this solution, by identifying single-line logs and multi-line logs, each complete log data can be determined for use in subsequent content identification, error log extraction, and operation log division, thereby improving the management efficiency of log data.
[0112] The technical solution provided by the embodiment of the present application obtains the log data to be processed of the target platform; extracts characteristic characters from each line of the log data to be processed, and constructs the characteristic code of the log data to be processed; inputs the characteristic code into a log merging model pre-formed by merging a cluster center-based aggregation model and a density-based aggregation model; determines the log type of each line of the log data to be processed according to the output result of the log merging model; wherein the log type includes a single-line log type and a multi-line log type; merges the data of the multi-line log type to obtain the processing result of the log data to be processed. By executing this solution, the applicability of log data processing can be enhanced without complicated management and maintenance, and the accuracy of log data processing can also be improved.
[0113] Embodiment 2
[0114] Figure 2 It is a flow chart of the training process of the log merging model provided in the second embodiment of the present invention, and this embodiment is optimized based on the above embodiment. The specific optimization is: the training process of the log merging model includes: collecting log data from at least one platform, and dividing the training set from the log data; extracting feature characters for each line of log data in the training set; constructing feature codes for each line of log data based on the feature characters; using a log merging model formed by merging a cluster center-based aggregation model and a density-based aggregation model to cluster the feature codes to obtain a training feature set; based on the recognition results of the single-line log feature set and the multi-line log feature set of the training feature set, and the pre-labeling results of the log data in the training set, the log merging model is trained to obtain a training result.
[0115] like Figure 2 As shown, the method of this embodiment specifically includes the following steps:
[0116] S210, collecting log data of at least one platform, and dividing a training set from the log data.
[0117] Among them, after collection, the log data can be divided into training sets.
[0118] Specifically, the training type can be supervised training or unsupervised training. If it is supervised training, each log data needs to be labeled, for example, labeled as a single-line log label and a multi-line log label. The labeling process can be performed by a staff member, who can enter the data in the electronic device of this solution. The electronic device can divide the labeled log data into a training set based on the input results.
[0119] In addition to dividing the training set, this solution can also be divided into a test set. The test set and the training set can be used to train and test the model respectively. Moreover, the testing process can be similar to the training process, and the accuracy of the model recognition can be determined by comparing the obtained results with the annotated labels, thereby determining whether the model is usable.
[0120] S220, extracting characteristic characters from each line of log data in the training set.
[0121] The proposed method of extracting characteristic characters is similar to that of the above-mentioned embodiment and will not be described in detail.
[0122] S230: construct a feature code for each line of log data according to the feature character.
[0123] In this solution, the construction of the signature code is based on the conversion of the signature characters. Specifically, if the numbers and letters in the log are different but do not affect the format of the log, the numbers in the first N characters are replaced with 0, the letters are replaced with 1, the special symbols are replaced with the corresponding ASCII codes, and the Chinese characters are replaced with 2. After the replacement, the signature codes corresponding to the first N characters are obtained.
[0124] S240, clustering the feature codes using a log merging model formed by merging a cluster center-based aggregation model and a density-based aggregation model to obtain a training feature set.
[0125] In this solution, the log data of the training set can be input into the log merging model and clustered to obtain a training feature set. The training feature set here can be the recognition result obtained by the log merging model based on the log data of the training set, such as the cluster division result and the log type recognition result of each cluster.
[0126] In this scheme, the internal parameters of the cluster center-based aggregation model and the density-based aggregation model can be initial values, and the weighted coefficients of the cluster center-based aggregation model and the density-based aggregation model can also be initial values. After training, the internal initial values and weighted initial values of the cluster center-based aggregation model and the density-based aggregation model can be iterated to obtain a final internal initial value and weighted initial value, so that the log classification results of the log data in the training set correspond to the labels.
[0127] In this technical solution, the construction process of the log merging model includes:
[0128] Determining an initial merging weight of the cluster center-based aggregation model and the density-based aggregation model;
[0129] According to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, the initial merging weight of the log merging model is adjusted to obtain the final merging weight.
[0130] The initial merging weight may be the above-mentioned initial weight value, which can be used to determine the influence ratio between the cluster center-based aggregation model and the density-based aggregation model on the output result.
[0131] According to the recognition result of each line of log data in the training set, the accuracy of the log data recognition result corresponding to each merging weight can be determined in the iteration process. Specifically, the accuracy can be determined by comparing the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set. It can be understood that the one with the highest accuracy can be used as the final selected merging weight, that is, the final merging weight.
[0132] The advantage of this setting is that the merging weight of the two can be used as a parameter in the training iteration process and trained together, avoiding the problem of affecting the accuracy of the model due to inappropriate weight setting, improving training efficiency, and removing the influence of subjective factors to obtain a more stable log merging model.
[0133] In the above technical solution, optionally, according to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, the initial merging weight of the log merging model is adjusted to obtain the final merging weight, including:
[0134] According to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, respectively determine a first accuracy of the cluster center-based aggregation model and a second accuracy of the density-based aggregation model;
[0135] According to the magnitude relationship between the first accuracy and the second accuracy, the initial merging weight of the log merging model is adjusted to obtain a final merging weight.
[0136] In this solution, the recognition result of each row of log data in the training set can be compared with the label of the log data, and the first recognition result can be obtained by using the cluster center-based aggregation model, and the second recognition result can be obtained by using the density-based aggregation model. The accuracy of each aggregation model can be obtained by comparing the first recognition result and the second recognition result with the label data. On this basis, how to adjust the initial merging weight can be determined according to the accuracy of the two.
[0137] The benefit of this setting is that it can avoid the problem of insufficient tolerance of noise in the cluster center-based aggregation model, and can also avoid the density-based aggregation model being biased by noise, thereby performing weighted merging to obtain a more robust log merging model.
[0138] S250, training the log merging model according to the recognition results of the single-line log feature set and the multi-line log feature set of the training feature set and the pre-labeling results of the log data in the training set to obtain a training result.
[0139] In this solution, a training feature set can be output through a log merging model, wherein the training feature set is output in the form of a set or a cluster, and for each training feature set, the log merging model can mark the set as a single-line log feature set or a multi-line log feature set.
[0140] Correspondingly, the log merging model may be trained according to the pre-labeled results of the log data in the training set to obtain a training result.
[0141] Based on the above technical solutions, optionally, after obtaining the training result, the method further includes:
[0142] Dividing a test set from the log data;
[0143] Extracting characteristic characters from the log data of the test set;
[0144] Constructing a feature code of the log data according to the feature character;
[0145] Inputting the feature code of the log data of the test set into the log merging model to obtain a test feature set;
[0146] Determining whether the log merging model meets the accuracy requirement according to the recognition results of the single-line log feature set and the multi-line log feature set of the test feature set and the pre-labeling results of the log data in the training set;
[0147] If so, it is determined that the training result is available.
[0148] The training results can be tested based on the pre-divided test set to determine whether the trained log merging model can meet the accuracy requirement. Since the process performed in the testing phase is the same as that in the training phase, it will not be repeated here.
[0149] Through such a setting, this solution can improve the accuracy of the training results obtained during the training process, and after testing, it can be determined whether the training results can be put online for use.
[0150] Based on the above embodiments, this embodiment provides a training method and a testing method for the log merging model. By adopting the specific method provided by this solution, the training effect of the log merging model can be effectively improved, and the compatibility of the log merging model can be improved and deviations caused by subjective factors can be avoided.
[0151] On the basis of the above embodiments, optionally, feature character extraction is performed on each line of log data in the training set, including:
[0152] According to the predetermined number of at least two characteristic characters, starting from the first characteristic character of each line of log data, a target character is determined;
[0153] Determine the target character as a characteristic character, and extract the characteristic characters of each characteristic character quantity;
[0154] Correspondingly, after obtaining the training results of the number of each characteristic character, the method further includes:
[0155] The target number of characteristic characters is determined according to the accuracy of the training results of the number of each characteristic character.
[0156] Among them, at least two numbers of characteristic characters can be 10, 20, 30, 50 and 100. That is, there are five numbers, and training is performed for each number separately. That is, the first 10, first 20, first 30, first 50 and first 100 in each line of log are extracted respectively. After training, the accuracy of the training results corresponding to each number is obtained, and the one with the highest accuracy is used as the final target number of characteristic characters. For example, among the first 10, first 20, first 30, first 50 and first 100, the result obtained by extracting the first 50 characters and outputting the log merging model is the most accurate, then the method of extracting the characteristic characters is determined to be extracting the first 50.
[0157] Through such a setting, this solution can extract the number of feature characters in a scientific way, and can extract the number of feature characters that can accurately reflect the log type of each line of log, thereby improving the accuracy of the log merging model.
[0158] Based on the above technical solution, optionally, before determining the target number of characteristic characters according to the accuracy of the training results of the number of characteristic characters, the method further includes:
[0159] Obtain the training time factor and training computing power factor of each characteristic character quantity;
[0160] According to the accuracy of the training results of the number of each characteristic character, the target number of characteristic characters is determined, including:
[0161] The target number of characteristic characters is determined according to the accuracy of the training results of the number of characteristic characters, the training time factor and the training computing power factor.
[0162] Among them, due to the different number of feature characters, the training time will definitely be different. For example, the first 10 feature characters take the shortest time, the first 30 feature characters take the middle time, and the first 100 feature characters take the longest time. The training computing power factor is the same. If the number of feature characters is smaller, the computing power used in the training process is also less. In the actual training process, if you only pursue the accuracy of the results, there may be a phenomenon that the accuracy increases with the increase in the number. However, the increase in the number will cause a sharp increase in the training time and training computing power, while the increase in accuracy is not obvious. In this case, you can increase the reference training time factor and training computing power factor to obtain a more efficient log merging model.
[0163] In order to facilitate those skilled in the art to understand the present solution, this embodiment also provides a preferred implementation scheme.
[0164] This method is based on an automated alarm prediction method based on artificial intelligence. It uses machine learning and artificial intelligence algorithms to perform cluster analysis on log data for massive operation and maintenance data, extract regular information from logs, and use similarity measurement to merge logs.
[0165] The specific steps are as follows:
[0166] Preprocessing: Data cleaning, removing syslog, kafka and other log information that does not need to be merged according to the log type.
[0167] Step 1: Collect historical log data as a training set. Each log format is fixed, so whether each log line is a single-line log can be quickly determined based on the first N characters of each log line. Based on the historical log data set, extract the first 10, 20, 30, 50, and 100 characters of each log line as the basis for determining whether the log is a single line.
[0168] The numbers and letters in the log are different but do not affect the format of the log. Replace the numbers in the first N characters with 0, the letters with 1, the special symbols with the corresponding ASCII codes, and the Chinese characters with 2. After replacement, the feature codes corresponding to the first N characters are obtained.
[0169] Step 2: Use the K-means algorithm to cluster the feature codes corresponding to the first N characters to obtain the feature sets corresponding to the first N characters.
[0170] Step 3: Use the cluster center-based aggregation model to divide the training log feature set into two categories: single-line data and non-single-line data, and save the model of the corresponding relationship between the feature set and the single-line data. According to the eigenvalue and accuracy statistics, the clustering based on the cluster center has outliers and the noise data processing is not ideal.
[0171] Step 4: Use the density-based aggregation model to divide the training log feature set into two categories: single-row data and non-single-row data, and save the model of the correspondence between the feature set and the single-row data.
[0172] Step 5: Merge the cluster center-based aggregation model and the density-based aggregation model to generate a log merge model. After merging, it is found that clusters of arbitrary shapes can be found and are not sensitive to noisy data.
[0173] Step 6: Add production data, extract feature codes, and verify the efficiency of the model. Comparing the accuracy of the five models, we found that the calculation efficiency and accuracy of the first 30 characters were just right. The first 10 and 20 characters had too little data and low accuracy. The first 50 and 100 characters had a large amount of calculation, slow speed, and little improvement in accuracy.
[0174] Through the above method, logs can be automatically merged to improve the visibility of logs, so that development and operation and maintenance personnel can better troubleshoot, diagnose and analyze the logs of equipment or services, thereby reducing the economic losses caused by failures for companies or enterprises; reducing the difficulty of platform management, thereby reducing labor costs for companies or enterprises.
[0175] Embodiment 3
[0176] Figure 3It is a structural block diagram of a log data processing device provided in Embodiment 3 of the present invention. The device can execute the log data processing method provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0177] like Figure 3 As shown, the device comprises:
[0178] The to-be-processed log data acquisition module 310 is used to acquire the to-be-processed log data of the target platform;
[0179] A feature code construction module 320 is used to extract feature characters from each line of the log data to be processed, and construct a feature code of the log data to be processed;
[0180] A feature code input module 330, used for inputting the feature code into a log merging model pre-merged by a cluster center-based aggregation model and a density-based aggregation model;
[0181] A log type output module 340, configured to determine the log type of each line of the to-be-processed log data according to the output result of the log merging model; wherein the log type includes a single-line log type and a multi-line log type;
[0182] The log data processing module 350 is used to merge multiple lines of log data to obtain a processing result of the log data to be processed.
[0183] Furthermore, the log data processing module is specifically used to:
[0184] If it is a single-line log type, determine that the log data of the current line is complete data;
[0185] If it is a multi-line log type, the log data of the current line is merged with the log data of the multi-line log type of the adjacent line to obtain complete data.
[0186] Furthermore, the device also includes a log merging model training module, which is used to:
[0187] Collecting log data of at least one platform, and dividing a training set from the log data;
[0188] Extracting characteristic characters from each line of log data in the training set;
[0189] Constructing a characteristic code for each line of log data according to the characteristic characters;
[0190] Clustering the feature codes using a log merging model formed by merging a cluster center-based aggregation model and a density-based aggregation model to obtain a training feature set;
[0191] The log merging model is trained according to the recognition results of the single-line log feature set and the multi-line log feature set of the training feature set, and the pre-labeling results of the log data in the training set, so as to obtain a training result.
[0192] Furthermore, the device also includes a log merging model building module, including:
[0193] An initial merging weight determining unit, used to determine the initial merging weights of the cluster center-based aggregation model and the density-based aggregation model;
[0194] The final merging weight calculation unit is used to adjust the initial merging weight of the log merging model to obtain the final merging weight according to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set.
[0195] Furthermore, the final combined weight calculation unit is specifically used for:
[0196] According to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, respectively determine a first accuracy of the cluster center-based aggregation model and a second accuracy of the density-based aggregation model;
[0197] According to the magnitude relationship between the first accuracy and the second accuracy, the initial merging weight of the log merging model is adjusted to obtain a final merging weight.
[0198] Furthermore, the device also includes a log merging model testing module, which is used to:
[0199] Dividing a test set from the log data;
[0200] Extracting characteristic characters from the log data of the test set;
[0201] Constructing a feature code of the log data according to the feature character;
[0202] Inputting the feature code of the log data of the test set into the log merging model to obtain a test feature set;
[0203] Determining whether the log merging model meets the accuracy requirement according to the recognition results of the single-line log feature set and the multi-line log feature set of the test feature set and the pre-labeling results of the log data in the training set;
[0204] If so, it is determined that the training result is available.
[0205] Furthermore, the log merging model training module is also used for:
[0206] According to the predetermined number of at least two characteristic characters, starting from the first characteristic character of each line of log data, a target character is determined;
[0207] Determine the target character as a characteristic character, and extract the characteristic characters of each characteristic character quantity;
[0208] Correspondingly, after obtaining the training results of the number of each characteristic character, the method further includes:
[0209] The target number of characteristic characters is determined according to the accuracy of the training results of the number of each characteristic character.
[0210] Furthermore, the device also includes a feature character quantity selection module, which is used to:
[0211] Obtain the training time factor and training computing power factor of each characteristic character quantity;
[0212] According to the accuracy of the training results of the number of each characteristic character, the target number of characteristic characters is determined, including:
[0213] The target number of characteristic characters is determined according to the accuracy of the training results of the number of characteristic characters, the training time factor and the training computing power factor.
[0214] The above-mentioned product can execute the log data processing method provided in the embodiment of the present application, and has the corresponding functional modules and beneficial effects of the execution method.
[0215] Embodiment 4
[0216] Embodiment 4 of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for processing log data provided in all embodiments of the present invention is implemented:
[0217] Get the log data to be processed of the target platform;
[0218] Extract characteristic characters from each line of the log data to be processed, and construct a characteristic code of the log data to be processed;
[0219] Inputting the feature code into a log merging model pre-merged by a cluster center-based aggregation model and a density-based aggregation model;
[0220] Determine the log type of each line of the log data to be processed according to the output result of the log merging model; wherein the log type includes a single-line log type and a multi-line log type;
[0221] Multiple lines of log data are merged to obtain a processing result of the log data to be processed.
[0222] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device.
[0223] Computer-readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0224] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0225] Computer program code for performing the operations of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0226] Embodiment 5
[0227] Embodiment 5 of the present application provides an electronic device. Figure 4 Schematic diagram of the structure of an electronic device provided in Embodiment 5 of the present application. Figure 4 As shown, this embodiment provides an electronic device 400, which includes: one or more processors 420; a storage device 410, which is used to store one or more programs. When the one or more programs are executed by the one or more processors 420, the one or more processors 420 implement the log data processing method provided in the embodiment of the present application, and the method includes:
[0228] Get the log data to be processed on the target platform;
[0229] Extract characteristic characters from each line of the log data to be processed, and construct a characteristic code of the log data to be processed;
[0230] Inputting the feature code into a log merging model pre-merged by a cluster center-based aggregation model and a density-based aggregation model;
[0231] Determine the log type of each line of the log data to be processed according to the output result of the log merging model; wherein the log type includes a single-line log type and a multi-line log type;
[0232] Multiple lines of log data are merged to obtain a processing result of the log data to be processed.
[0233] Of course, those skilled in the art can understand that the processor 420 also implements the technical solution of the log data processing method provided in any embodiment of the present application.
[0234] Figure 4 The electronic device 400 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0235] like Figure 4 As shown, the electronic device 400 includes a processor 420, a storage device 410, an input device 430, and an output device 440; the number of the processor 420 in the electronic device can be one or more. Figure 4 A processor 420 is taken as an example; the processor 420, the storage device 410, the input device 430 and the output device 440 in the electronic device can be connected via a bus or other means. Figure 4 The connection via bus 450 is taken as an example.
[0236] The storage device 410 is a computer-readable storage medium that can be used to store software programs, computer executable programs, and module units, such as program instructions corresponding to the log data processing method in the embodiment of the present application.
[0237] The storage device 410 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal, etc. In addition, the storage device 410 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the storage device 410 may further include a memory remotely arranged relative to the processor 420, and these remote memories may be connected via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0238] The input device 430 may be used to receive input numbers, character information or voice information, and generate key signal input related to user settings and function control of the electronic device. The output device 440 may include electronic devices such as a display screen and a speaker.
[0239] The electronic device provided by the embodiment of the present application can enhance the applicability of log data processing without the need for complicated management and maintenance, while also being able to improve the accuracy of log data processing.
[0240] The log data processing device, medium, and electronic device provided in the above embodiments can execute the log data processing method provided in any embodiment of the present application, and have the corresponding functional modules and beneficial effects of executing the method. For technical details not described in detail in the above embodiments, please refer to the log data processing method provided in any embodiment of the present application.
[0241] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for processing log data, characterized in that: The method comprises: Obtaining log data to be processed of the target platform; the log data to be processed includes single-line log type data and multi-line log type data; Extract characteristic characters from each line of the log data to be processed, and construct a characteristic code of the log data to be processed; Inputting the feature code into a log merging model pre-merged by a cluster center-based aggregation model and a density-based aggregation model; Determine the log type of each line of data of the to-be-processed log data according to the output result of the log merging model; wherein the log type of each line of data includes a single-line log type and a multi-line log type; The data of the multiple log types are merged to obtain the processing result of the log data to be processed, including: If it is a single-line log type, determine that the log data of the current line is complete data; If it is a multi-line log type, the log data of the current line is merged with the log data of the multi-line log type of the adjacent line to obtain complete data; The training process of the log merging model includes: Collect log data of at least one platform, and divide a training set from the log data; extract feature characters from each line of log data in the training set; construct feature codes for each line of log data according to the feature characters; cluster the feature codes using a log merging model formed by merging a cluster center-based aggregation model and a density-based aggregation model to obtain a training feature set; train the log merging model according to recognition results of a single-line log feature set and a multi-line log feature set of the training feature set, and pre-labeling results of the log data in the training set, to obtain a training result; The construction process of the log merging model includes: Determining an initial merging weight of the cluster center-based aggregation model and the density-based aggregation model; According to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set, respectively determine a first accuracy of the cluster center-based aggregation model and a second accuracy of the density-based aggregation model; According to the magnitude relationship between the first accuracy and the second accuracy, the initial merging weight of the log merging model is adjusted to obtain a final merging weight.
2. The method according to claim 1, characterized in that After obtaining the training result, the method further includes: Dividing a test set from the log data; Extracting characteristic characters from the log data of the test set; Constructing a feature code of the log data according to the feature character; Inputting the feature code of the log data of the test set into the log merging model to obtain a test feature set; Determining whether the log merging model meets the accuracy requirement according to the recognition results of the single-line log feature set and the multi-line log feature set of the test feature set and the pre-labeling results of the log data in the training set; If so, it is determined that the training result is available.
3. The method according to claim 1, characterized in that Extracting characteristic characters from each line of log data in the training set includes: According to the predetermined number of at least two characteristic characters, starting from the first characteristic character of each line of log data, a target character is determined; Determine the target character as a characteristic character, and extract the characteristic characters of each characteristic character quantity; Correspondingly, after obtaining the training results of the number of each characteristic character, the method further includes: The target number of characteristic characters is determined according to the accuracy of the training results of the number of each characteristic character.
4. The method according to claim 3, characterized in that: Before determining the target number of characteristic characters according to the accuracy of the training results of the number of characteristic characters, the method further includes: Obtain the training time factor and training computing power factor of each characteristic character quantity; According to the accuracy of the training results of the number of each characteristic character, the target number of characteristic characters is determined, including: The target number of characteristic characters is determined according to the accuracy of the training results of the number of characteristic characters, the training time factor and the training computing power factor.
5. A log data processing device, characterized in that: The device comprises: The to-be-processed log data acquisition module is used to acquire the to-be-processed log data of the target platform; the to-be-processed log data includes single-line log type data and multi-line log type data; A feature code construction module, used to extract feature characters from each line of the log data to be processed, and construct a feature code for the log data to be processed; A feature code input module, used for inputting the feature code into a log merging model pre-merged by a cluster center-based aggregation model and a density-based aggregation model; A log type output module, used to determine the log type of each line of data of the to-be-processed log data according to the output result of the log merging model; wherein the log type of each line of data includes a single-line log type and a multi-line log type; A log data processing module, used for merging the data of the multiple log types to obtain a processing result of the log data to be processed; The log data processing module is specifically used to determine that the log data of the current row is complete data if it is a single-row log type; if it is a multi-row log type, merge the log data of the current row with the log data of the multi-row log type of the adjacent row to obtain complete data; A log merging model training module is used to collect log data of at least one platform and divide a training set from the log data; extract feature characters from each line of log data in the training set; construct feature codes for each line of log data according to the feature characters; cluster the feature codes using a log merging model formed by merging a cluster center-based aggregation model and a density-based aggregation model to obtain a training feature set; train the log merging model according to the recognition results of a single-line log feature set and a multi-line log feature set of the training feature set and the pre-labeling results of the log data in the training set to obtain a training result; A log merging model building module, including an initial merging weight determination unit and a final merging weight calculation unit; The initial merging weight determining unit is used to determine the initial merging weights of the cluster center-based aggregation model and the density-based aggregation model; The final merge weight calculation unit is used to determine the first accuracy of the cluster center-based aggregation model and the second accuracy of the density-based aggregation model according to the recognition result of each line of log data in the training set and the pre-labeling result of the log data in the training set; according to the relationship between the first accuracy and the second accuracy, the initial merge weight of the log merge model is adjusted to obtain the final merge weight.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the log data processing method according to any one of claims 1 to 4 is implemented.
7. An electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for processing log data according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Distributed log system
CN109542750A