A method for analyzing large amounts of log data and a storage medium

By building a decision tree on the first terminal and distributing it to multiple log analysis terminals for parallel analysis, the problem of large-data log data analysis is solved, efficient and reliable analysis results are achieved, and resource occupation and cost are reduced.

CN119576888BActive Publication Date: 2025-05-16TIANJIN HAOYANG HUANYU TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510135191.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-16
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

In the large amount of log data, it is difficult for the existing technology to effectively analyze and ensure the reliability of the analysis, especially in terms of resource occupation and event balance, resulting in increased costs.

Method used

By obtaining the log sample data set based on the first rule on the first terminal, a decision tree is constructed, and distributed to multiple log analysis terminals. These terminals perform log analysis in parallel, update the decision tree, and merge the update content to obtain the final log analysis results.

Benefits of technology

It realizes efficient analysis and reliability of large-data log data, reduces resource occupation and cost, and improves analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576888B_ABST
    Figure CN119576888B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of information processing, and specifically relates to a method for analyzing large amounts of log data and a storage medium. The method provided by the present invention comprises: a first terminal obtains a first log sample data set based on a first rule, and constructs a first decision tree based on the first log sample data set; the first decision tree is distributed to multiple log analysis terminals; multiple log analysis terminals analyze logs in parallel, and update the decision trees used by them; and the decision trees of multiple log analysis terminals are merged to obtain log analysis results. The method of the present invention can improve the data processing capability of large amounts of data and ensure repeatability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information processing, and in particular relates to an analysis method and storage medium for large-volume log data. Background Art

[0002] With the widespread application of cloud computing and the increase in data processing needs, it is necessary to monitor, analyze and audit various components of an enterprise. Currently, enterprises often involve multiple cloud environments. Although logs can be recorded in a variety of ways and information can be mined through log analysis, it is difficult to analyze logs when the log data is huge and scattered. For example, some logs are stored in databases such as SQLServer and Mysql, some are stored in Elasticsearch or listed databases, and some logs are stored in files. It is difficult to analyze them and ensure the reliability of the analysis.

[0003] In addition, it is also necessary to consider the occupancy of resources and the balance of events to avoid the increase of costs. Therefore, the present invention proposes an analysis method and storage medium for large amounts of log data. Summary of the invention

[0004] At least one aspect or advantage of the invention will be set forth in part in the following description, or may be obvious from the description, or may be learned by practice of the disclosed subject matter.

[0005] According to an embodiment of the present invention, a method for analyzing large amounts of log data includes:

[0006] The first terminal obtains a first log sample data set based on the first rule, and constructs a first decision tree based on the first log sample data set;

[0007] The first decision tree is distributed to a plurality of log analysis terminals;

[0008] The multiple log analysis terminals analyze the logs in parallel, update the decision trees used by them, and send the updated content of the decision trees to the first terminal;

[0009] Merge the decision trees of multiple log analysis terminals to obtain log analysis results;

[0010] The first rule includes log source information, sampling ratio information and sampling rules.

[0011] According to an embodiment of the present invention, obtaining a first log sample data set based on a first rule includes:

[0012] Determine the log source of the log according to the first rule, and obtain a first log set to be analyzed based on the log source;

[0013] Determine the size of the first log set and the size of the space occupied by the logs according to the sampling ratio of the first rule to determine the sampling amount;

[0014] Obtaining first log sample data of the log using the sampling method included in the first rule for the first log file;

[0015] The ratio of the first log sample data to the total number of records in the log file is not less than the sampling ratio.

[0016] According to one embodiment of the present invention, the first decision tree is constructed based on a first log sample data set; the row data of the first log sample data set includes at least timestamp, application context information, load information, connection information, exception type and exception information; and the root node of the first decision tree is constructed according to the exception type.

[0017] According to an embodiment of the present invention, the multiple log analysis terminals concurrently performing log analysis and decision tree updating include:

[0018] The first log analysis terminal receives the log source sent by the first terminal;

[0019] Acquire a first log set from a log source according to the first window, and after acquiring the first log set, synchronize identification information of the first log set to the first terminal;

[0020] Using a second decision tree to analyze the logs in the first log set;

[0021] After completing the analysis of the first log set;

[0022] The identification information of the first log set is sent to the first terminal.

[0023] According to one embodiment of the present invention, the first terminal further includes a first task unit and a first buffer;

[0024] The first task unit is used to configure the analysis task of the log analysis terminal, the analysis task includes a log source, a sliding window size and a task index, the first terminal determines the analysis tasks for all log analysis terminals according to batches, and the analysis tasks created in the same batch have a consistent task index;

[0025] The first buffer is used to store changes made by the multiple log analysis terminals to the second decision tree used by the multiple log analysis terminals.

[0026] According to an embodiment of the present invention, the multiple log analysis terminals concurrently performing log analysis and decision tree updating include:

[0027] The second log analysis terminal receives the analysis task sent by the first terminal;

[0028] Acquire a second log set from the log source according to the second window configured for the analysis task, and send identification information of the second log set to the first terminal;

[0029] creating a copy of the second decision tree as a third decision tree;

[0030] Using a second decision tree to analyze the logs in the second log set;

[0031] After completing the analysis of the second log set, determining tree difference information based on the difference between the second decision tree and the third decision tree;

[0032] The identification information of the second log set is sent to the first terminal, and the tree difference information is sent to the first buffer of the first terminal.

[0033] According to one embodiment of the present invention, after the tasks with the same task index are completed, in response to the first buffer being empty, the first decision tree is not modified;

[0034] or in response to the first buffer being not empty and the tree difference information in the first buffer being consistent, updating the tree difference information in the first buffer to the first decision tree;

[0035] Or in response to the first buffer being not empty and the tree difference information in the first buffer being inconsistent, determine the analysis task sequence corresponding to the tree difference information, select any log analysis terminal to execute the analysis task in the analysis task sequence, obtain the first tree difference information, and update the first tree difference information to the first decision tree.

[0036] According to an embodiment of the present invention, when the first decision tree is updated, the updated node information of the first decision tree is synchronized to all log analysis terminals, and the log analysis terminals update the decision trees used based on the updated node information.

[0037] According to one embodiment of the present invention, when the tree difference information of the second log analysis terminal is empty, the first terminal assigns a new analysis task to the second log analysis terminal;

[0038] When the tree difference information of the second log analysis terminal is not empty, the tree difference information is sent to the first terminal. After the task with the same task index is completed, in response to the tree difference information in the first buffer being consistent or the number being 1, the tree difference information in the first buffer is updated to the first decision tree; or in response to the tree difference information in the first buffer being inconsistent, the analysis task sequence corresponding to the tree difference information is determined, any log analysis terminal is selected to execute the analysis task in the analysis task sequence, the first tree difference information is obtained, and the first tree difference information is updated to the first decision tree.

[0039] According to an embodiment of the present invention, a computer-readable storage medium stores a program, and the program is executed by a processor to analyze the large amount of log data in the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The present invention is a flowchart of an analysis method for large amounts of log data in an example of the present invention. DETAILED DESCRIPTION

[0041] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0042] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. According to one embodiment of the present invention, a method for analyzing large amounts of log data includes:

[0043] The first terminal obtains a first log sample data set based on the first rule, and constructs a first decision tree based on the first log sample data set;

[0044] The first decision tree is distributed to a plurality of log analysis terminals;

[0045] The multiple log analysis terminals analyze the logs in parallel, update the decision trees used by them, and send the updated content of the decision trees to the first terminal;

[0046] Merge the decision trees of multiple log analysis terminals to obtain log analysis results;

[0047] The first rule includes log source information, sampling ratio information and sampling rules.

[0048] The first terminal of the present invention can be a server deployed in an enterprise, which provides services through WebSocket communication or by providing an application program interface API. The first terminal can access log data or at least basic information of log data, such as storage form, log size, log record format, etc. The log information accessed by the present invention has a data set that exceeds the size of ordinary capabilities, or is processed within a tolerable time.

[0049] The first terminal is not limited to the operating system or device type used. For example, it can work by being deployed in a container, so that it has portability and flexibility.

[0050] In the present invention, the first terminal is mainly used for the scheduling of the analysis process and the generation of the final decision tree, wherein the scheduling of the analysis process depends on the generation of a decision tree using a small amount of sampled data. Since the data generation process may be smooth or unbalanced, but the data are correlated, the corresponding data can be representative by setting a suitable sampling method, for example, the obtained sampling points are evenly distributed according to time, or cover the main time period, or correspond to a certain category of sampling, and the obtained logs are recorded with a rough time interval, then a rough decision tree can be constructed based on this, and analysis can be performed based on this decision tree, so that specific analysis can be completed on this basis and basic accuracy is guaranteed.

[0051] Among them, the decision tree is a machine learning algorithm. When the program configuration exception information interface and log recording function of the application layer, or in most cases, the information recorded by the application server is used for analysis to obtain a tree structure, which can be used for classification tasks or regression tasks. The decision tree mainly involves nodes, edges and paths, where a node represents a feature or attribute, the root node represents the top-level node of the decision tree, and represents the initial segmentation of the entire data set; the internal node or child node is a node other than the root node and the leaf node, which is used to further segment the data; the leaf node is the end node of the tree, which represents the classification or regression result; the edge is the connection between the nodes, which represents the possible value of the feature or attribute; the path is a path from the root node to the leaf node, which represents a series of decisions. The construction of the decision tree goes through the selection of the best features, data segmentation, recursive construction of subtrees and generation of leaf nodes; the present invention performs feature selection based on the Gini index, specifically the CART algorithm. The Gini index and the CART algorithm are not repeated here.

[0052] For the convenience of comparison, the present invention selects the exception type as the root node by default, wherein the exception type is a user-defined integer or enumerated data, for example, when set to an integer 0-7, 8 different types of risk information can be indicated respectively.

[0053] After the first decision tree is constructed, the analysis results can be obtained by performing analysis on different analysis terminals, and since it is mainly based on the first decision tree, the corresponding result merging or updating has a consistent basis, so that even if anomalies are detected later, their impact on the first decision tree is minimized.

[0054] After the analysis is completed, the actual decision trees of each analysis terminal can be merged, and the obtained decision tree is used as the analysis result.

[0055] Obviously, parallel analysis can be performed in this way, and the calculation process is reduced because the previous sampling allows the subsequent analysis to be performed on the same tree. According to an embodiment of the present invention, obtaining the first log sample data set based on the first rule includes:

[0056] Determine the log source of the log according to the first rule, and obtain a first log set to be analyzed based on the log source;

[0057] Determine the size of the first log set and the size of the space occupied by the logs according to the sampling ratio of the first rule to determine the sampling amount;

[0058] Obtain first log sample data using the sampling method included in the first rule for the first log file;

[0059] The ratio of the first log sample data to the total number of records in the log file is not less than the sampling ratio.

[0060] The first rule is mainly used to define the data source actually used, such as how the data source is provided or obtained. The sampling ratio is used to determine the size of the sampling amount when constructing the first decision tree. It should be understood that when the data source is a database, the corresponding connection method, driver name and other information should be provided, or obtained through the API port; when the data source is a file, the corresponding URI should be provided for acquisition; when the data source is a proprietary file system, the file should be obtained in a way that the log analysis terminal can support, such as obtaining the file through the fs command, etc. Similarly, the file source information involved in the first rule should also apply to the log analysis server.

[0061] The sampling ratio here can generally be set to a value that is completed in a shorter time and has representative sampling. For example, when the analyzed data is mGB, the actual sampling ratio can be set so that at least 50MB of data is analyzed, so that the sampling ratio is 1 / (20*m). When sampling, each file can be sampled, such as by skipping a specified number of bytes according to the file size when reading the input stream, and then reading the next available log record line.

[0062] The sampling ratio should be set reasonably according to the hardware and the amount of data. It can be adjusted by setting a sampling ratio of 0.001, for example, and determining the actual processing time.

[0063] If the generation process is uniform, the data obtained according to the file size or offset should also be roughly uniform in time; if the data generation process is uneven, such as more during the day and less at night, the sampling process may be uniform in time, but it is still referenceable for the system status. According to one embodiment of the present invention, the first decision tree is constructed based on the first log sample data set; the row data of the first log sample data set includes at least timestamp, application context information, load information, connection information, exception type and exception information; and the root node of the first decision tree is constructed according to the exception type.

[0064] The above format is only a recommendation. The actual log output format may have different column names or arrangement orders. When constructing the first decision tree, you can select columns related to performance or risk as root nodes to maintain the relative balance of the first decision tree.

[0065] In addition, the log format should remain basically consistent, and when some data is missing, it defaults to a null value or a specific value (such as -1) to enable the analysis to be completed. According to one embodiment of the present invention, the multiple log analysis terminals perform log analysis and decision tree update in parallel, including:

[0066] The first log analysis terminal receives the log source sent by the first terminal;

[0067] Acquire a first log set from a log source according to the first window, and after acquiring the first log set, synchronize identification information of the first log set to the first terminal;

[0068] Using a second decision tree to analyze the logs in the first log set;

[0069] After completing the analysis of the first log set;

[0070] The identification information of the first log set is sent to the first terminal.

[0071] When the structure of the decision tree does not change significantly, which is typically the case when data is generated relatively evenly and the data in the sampling process is relatively representative, the structure of the first decision tree constructed will not change significantly. At this time, the first terminal will construct the analysis task according to the load of each log analysis terminal.

[0072] If the first terminal provides tasks in a synchronous manner, for the first terminal, only one log analysis terminal's request is processed at the same time. After completing the request of one log analysis terminal, other log processing tasks are assigned to the next log analysis terminal. The log analysis task here mainly includes the data source, the size of each log or the size of the sliding window, and the identification of the current task. It should be understood that the decision tree obtained is reliable only when a task is completed, so identification information can be set here to track the processing progress.

[0073] To explain in the most simplified way, the first terminal allocates tasks to the log analysis terminal 1 and the log analysis terminal 2; the log analysis terminal 1 and the log analysis terminal 2 send task requests at the same time, the first terminal blocks the request of the log analysis terminal 2, and sends the log processing task 1 to the log analysis terminal 1 at the same time. After completing the sending of the log processing task 1, the first terminal sends the log processing task 2 to the log analysis terminal 2. Obviously, the log processing task 2 and the log processing task 1 have no intersection; when the log analysis terminal 1 obtains the log data, it obtains the log line data according to the data source and the sliding window, and sends the corresponding identification (such as the offset data and the actual number of data obtained) to the first terminal; after completing the analysis, the corresponding identification is sent to the first terminal again. At this time, the first terminal compares the two received identifications to determine that the corresponding analysis task is completed, so that the subsequent consistent analysis tasks are no longer assigned, or when some analysis tasks are not completed, the analysis tasks are completed through other log analysis nodes. Afterwards, the log analysis terminal 1 or the log analysis terminal 2 continues to request tasks after completing the analysis task until all logs are analyzed.

[0074] Afterwards, all decision trees are merged, and the weights of the decision trees are adjusted. The new decision tree obtained is the log analysis result, that is, the impact of various factors on the normal operation of the system. At this time, due to its structural approximation, the weights of the edges can be directly added to obtain the final weight. According to one embodiment of the present invention, the first terminal also includes a first task unit and a first buffer;

[0075] The first task unit is used to configure the analysis task of the log analysis terminal, the analysis task includes a log source, a sliding window size and a task index, the first terminal determines the analysis tasks for all log analysis terminals according to batches, and the analysis tasks created in the same batch have a consistent task index;

[0076] The first buffer is used to store changes made by the multiple log analysis terminals to the second decision tree used by the multiple log analysis terminals.

[0077] According to an embodiment of the present invention, the multiple log analysis terminals concurrently performing log analysis and decision tree updating include:

[0078] The second log analysis terminal receives the analysis task sent by the first terminal;

[0079] Acquire a second log set from the log source according to the second window configured for the analysis task, and send identification information of the second log set to the first terminal;

[0080] creating a copy of the second decision tree as a third decision tree;

[0081] Using a second decision tree to analyze the logs in the second log set;

[0082] After completing the analysis of the second log set, determining tree difference information based on the difference between the second decision tree and the third decision tree;

[0083] The identification information of the second log set is sent to the first terminal, and the tree difference information is sent to the first buffer of the first terminal.

[0084] This method can avoid the defect of unreliable analysis results caused by uneven sampling.

[0085] In this step, the second log analysis terminal receives the analysis task sent by the first terminal, obtains the second log set from the log source according to the second window configured for the analysis task, and sends the identification information of the second log set to the first terminal to synchronize the tracking information of the task;

[0086] Afterwards, a copy of the second decision tree is created as a third decision tree, and the third decision tree is used to compare the differences of the decision trees after the analysis is completed;

[0087] Afterwards, the logs in the second log set are analyzed using the second decision tree, during which the weight or structure of the decision tree may change, thereby generating a tree;

[0088] After completing the analysis of the second log set, the tree difference information is determined based on the difference between the second decision tree and the third decision tree. If the structure has not changed, it can be considered that the decision tree has not changed; if pruning occurs, the structure of the decision tree will change. At this time, the tree change information is obtained by comparing the differences between the trees. When obtaining the tree change information, since the root node is specified in the previous text and sampling and training are provided, the change of the tree structure is generally limited. For example, if the node A with a height of 7 is removed and replaced by the node B, the difference can be expressed as [{"row":7,"add":"B","remove":"A"}]; approximately, the change of the node at each tree height is determined by comparing the nodes within the row height; in extreme cases, only a tree difference information needs to be registered in the first buffer without providing content. At this time, the tree difference information only needs to provide a mark of the change;

[0089] The identification information of the second log set is sent to the first terminal, and the tree difference information is sent to the first buffer of the first terminal.

[0090] At this time, the allocation of tasks can be blocked in accordance with the previous embodiment, and the allocation of analysis tasks can continue after the first buffer is processed; other methods can also be used, such as continuing task analysis for nodes whose tree structure has not changed.

[0091] The acquisition of the tree difference information in the first buffer is mainly used to form a new decision tree, which contains the association information between the newly obtained features. The decision tree can be distributed to each log analysis terminal for further analysis. According to one embodiment of the present invention, after the tasks with the same task index are completed, in response to the first buffer being empty, the first decision tree is not changed;

[0092] or in response to the first buffer being not empty and the tree difference information in the first buffer being consistent, updating the tree difference information in the first buffer to the first decision tree;

[0093] Or in response to the first buffer being not empty and the tree difference information in the first buffer being inconsistent, determine the analysis task sequence corresponding to the tree difference information, select any log analysis terminal to execute the analysis task in the analysis task sequence, obtain the first tree difference information, and update the first tree difference information to the first decision tree.

[0094] In this way, the reliability of the analysis results can be guaranteed and the number of data sets that need to be analyzed can be reduced.

[0095] After the task with the same task index is completed, in response to the first buffer being empty, it is obvious that the first decision tree does not need to be changed. In some cases, a decision tree can be re-obtained as the first decision tree based on the third decision tree of each log analysis terminal and synchronized to each log analysis terminal;

[0096] When the first buffer is not empty and the tree difference information in the first buffer is consistent, the tree difference information in the first buffer is updated to the first decision tree. At this time, the selected tree difference information should include information such as weights. When synchronizing, any tree difference information can be selected, or multiple trees that are consistent with the first decision tree can be constructed respectively, and changes based on the tree difference information are made respectively. After that, the decision tree obtained by merging multiple decision trees is used as the first difference tree, and synchronized to each log analysis terminal;

[0097] Or in response to the first buffer being not empty and the tree difference information in the first buffer being inconsistent, determine the analysis task sequence corresponding to the tree difference information, select any log analysis terminal to execute the analysis task in the analysis task sequence, obtain the first tree difference information, and update the first tree difference information to the first decision tree; because the weight information of the decision tree recalculation process is less reflected in the difference information, and because the tree pruning process may be smaller in this process, it is more convenient to directly generate a new tree by recalculation; when calculating, any node can be selected for calculation, and only the analysis task that generates the tree difference information needs to be considered during calculation, without considering the log information involved in other analysis tasks. This process can also be selected to be performed at the node that generates the tree difference information, because the third decision tree it uses has already reflected the correlation of various factors in some log information, and only other analysis tasks need to be processed. According to one embodiment of the present invention, when the first decision tree is updated, the node information updated by the first decision tree is synchronized to all log analysis terminals, and the log analysis terminal updates the decision tree used based on the updated node information.

[0098] In this way, each analysis node can analyze the log on a consistent basis, thereby improving the efficiency of the analysis. According to one embodiment of the present invention, when the tree difference information of the second log analysis terminal is empty, the first terminal assigns a new analysis task to the second log analysis terminal;

[0099] When the tree difference information of the second log analysis terminal is not empty, the tree difference information is sent to the first terminal. After the task with the same task index is completed, in response to the tree difference information in the first buffer being consistent or the number being 1, the tree difference information in the first buffer is updated to the first decision tree; or in response to the tree difference information in the first buffer being inconsistent, the analysis task sequence corresponding to the tree difference information is determined, any log analysis terminal is selected to execute the analysis task in the analysis task sequence, the first tree difference information is obtained, and the first tree difference information is updated to the first decision tree.

[0100] In this way, the blocking time can be reduced and the reliability of the analysis results can be guaranteed.

[0101] That is, when the tree difference information of the second log analysis terminal is empty, the first terminal assigns a new analysis task to the second log analysis terminal, which may cause the tree to be pruned or the structure to change, and then rollback is performed at this time; if there is no change, the analysis task is continued to be executed. At this stage, even if the tree structure of the first terminal changes, it does not affect the tree structure of the newly assigned task node;

[0102] When the tree difference information of the second log analysis terminal is not empty, the tree difference information is sent to the first terminal. After the task with the same task index is completed, when the tree difference information in the first buffer is consistent or the number is 1, it means that an approximate association relationship is found at this stage. At this time, the tree difference information in the first buffer is updated to the first decision tree; when updating or synchronizing, any tree difference information can be selected, or multiple trees that are consistent with the first decision tree can be constructed respectively, and changes based on the tree difference information are made respectively, and then the decision tree obtained by merging multiple decision trees is used as the first difference tree, and synchronized to each log analysis terminal;

[0103] When the tree difference information in the first buffer is inconsistent, determine the analysis task sequence corresponding to the tree difference information, select any log analysis terminal to execute the analysis task in the analysis task sequence, obtain the first tree difference information, update the first tree difference information to the first decision tree, and use the newly obtained first decision tree for the subsequent analysis process.

[0104] According to one embodiment of the present invention, a computer-readable storage medium is further provided, wherein the storage medium stores a program, and the program is executed by a processor to perform the analysis method of log data with a large amount of data in the above embodiments.

[0105] Those of ordinary skill in the art will appreciate that the modules and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0106] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and equipment can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0107] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0108] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of the present invention.

[0109] In addition, each functional module in the embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0110] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the energy-saving signal sending / receiving method of each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.

[0111] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other.

[0112] It should be understood that the size of the sequence number of each step in the content of the invention and the embodiments of the present invention does not absolutely mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. For the purpose of example and description, the foregoing description of the implementation of the present disclosure has been given. The foregoing description is not exhaustive and is not intended to limit the present disclosure to the exact form disclosed. Various deformations and modifications may exist according to the above teachings, or various deformations and modifications may be obtained from the practice of the present disclosure. These embodiments are selected and described to illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can use the present disclosure in various embodiments and various modifications suitable for the specific purpose conceived.

Claims

1. A method for analyzing large amounts of log data, characterized in that: include: The first terminal obtains a first log sample data set based on the first rule, and constructs a first decision tree based on the first log sample data set; The first decision tree is distributed to a plurality of log analysis terminals; The multiple log analysis terminals analyze the logs in parallel, update the second decision tree used by them, and send tree difference information to the first terminal, where the tree difference information is the update content of the second decision tree; Merge the second decision trees of multiple log analysis terminals to obtain log analysis results; The first rule includes log source information, sampling ratio information and sampling rules; The first terminal also includes a first task unit and a first buffer; The first task unit is used to configure the analysis task of the log analysis terminal, the analysis task includes a log source, a sliding window size and a task index, the first terminal determines the analysis tasks for all the log analysis terminals according to batches, and the analysis tasks created in the same batch have a consistent task index; The first buffer is used to store tree difference information of the multiple log analysis terminals; When the first decision tree is updated, the updated node information of the first decision tree is synchronized to all log analysis terminals, and the log analysis terminals update the used second decision tree based on the updated node information; After the tasks with the same task index are completed, in response to the first buffer being empty, not changing the first decision tree, or in response to the first buffer being not empty, updating the first decision tree based on the tree difference information in the first buffer; The obtaining of a first log sample data set based on a first rule includes: Determine the log source of the log according to the first rule, and obtain a first log set to be analyzed based on the log source; Determine the size of the first log set and the size of the space occupied by the logs to determine the sampling amount according to the sampling ratio of the first rule; Obtain first log sample data using the sampling method included in the first rule for the first log file; The ratio of the first log sample data to the total number of records in the log file is not less than the sampling ratio.

2. The method for analyzing large amounts of log data according to claim 1, characterized in that: The first decision tree is constructed based on a first log sample data set; the row data of the first log sample data set includes at least timestamp, application context information, load information, connection information, exception type and exception information; and the root node of the first decision tree is constructed according to the exception type.

3. The method for analyzing large amounts of log data according to claim 1, characterized in that: The multiple log analysis terminals concurrently perform log analysis and update the second decision tree, including: The first log analysis terminal receives the log source sent by the first terminal; Acquire a first log set from a log source according to the first window, and after acquiring the first log set, synchronize identification information of the first log set to the first terminal; Using a second decision tree to analyze the logs in the first log set; After completing the analysis of the first log set; The identification information of the first log set is sent to the first terminal.

4. The method for analyzing large amounts of log data according to claim 1, characterized in that: The multiple log analysis terminals concurrently perform log analysis and update the second decision tree, including: The second log analysis terminal receives the analysis task sent by the first terminal; Acquire a second log set from the log source according to the second window configured for the analysis task, and send identification information of the second log set to the first terminal; creating a copy of the second decision tree as a third decision tree; Using a second decision tree to analyze the logs in the second log set; After completing the analysis of the second log set, determining tree difference information based on the difference between the second decision tree and the third decision tree; The identification information of the second log set is sent to the first terminal, and the tree difference information is sent to the first buffer of the first terminal.

5. The method for analyzing large amounts of log data as claimed in claim 4, characterized in that: In response to the first buffer being not empty, updating the first decision tree based on the tree difference information in the first buffer comprises: or in response to the first buffer being not empty and the tree difference information in the first buffer being consistent, updating the tree difference information in the first buffer to the first decision tree; Or in response to the first buffer being not empty and the tree difference information in the first buffer being inconsistent, determine the analysis task sequence corresponding to the tree difference information, select any log analysis terminal to execute the analysis task in the analysis task sequence, obtain the first tree difference information, and update the first tree difference information to the first decision tree.

6. The method for analyzing large amounts of log data as claimed in claim 4, characterized in that: When the tree difference information of the second log analysis terminal is empty, the first terminal assigns a new analysis task to the second log analysis terminal; When the tree difference information of the second log analysis terminal is not empty, the tree difference information is sent to the first terminal. After the task with the same task index is completed, in response to the tree difference information in the first buffer being consistent or the number being 1, the tree difference information in the first buffer is updated to the first decision tree; or in response to the tree difference information in the first buffer being inconsistent, the analysis task sequence corresponding to the tree difference information is determined, any log analysis terminal is selected to execute the analysis task in the analysis task sequence, the first tree difference information is obtained, and the first tree difference information is updated to the first decision tree.

7. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method for analyzing log data with a large amount of data as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • A non-uniform sampling-based stream anomaly detection method and device, equipment and a medium

    CN113868866A

  • Threat detection method, device and system

    CN115168841A