Fault prediction method and device for distributed system
By establishing a log knowledge graph in a distributed system and using neural network models for fault prediction, the problem of inaccurate prediction in the existing technology is solved, and efficient fault prediction is achieved.
Patent Information
- Application Number
- CN202111265142.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-10-28
AI Technical Summary
In the fault prediction of distributed system, the prior art has inaccurate prediction information and incomplete prediction considerations, resulting in inaccurate prediction results and inefficient efficiency, which cannot meet the diverse fault needs.
By establishing a log knowledge graph of a distributed system, the characteristic information of the log data is extracted, and the failure prediction is performed using a semi-supervised learning neural network model, including data preprocessing, log number extraction, call relationship determination, triple data generation, graph representation analysis and neural network model training.
It improves the accuracy and efficiency of fault prediction of distributed system and meets the actual needs of enterprises.
Smart Images

Figure CN113961424B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis, and particularly to a fault prediction method for a distributed system and a fault prediction device for a distributed system. Background Art
[0002] With the continuous development of technology and business, technicians have proposed distributed systems to meet actual business and functional requirements. However, errors or anomalies may occur during the application of distributed systems, which will cause troubles to users.
[0003] To solve the above technical problems, technicians have proposed a fault prediction method for distributed systems. By extracting log data in the distributed system, specifically, extracting log data corresponding to multiple time intervals with the same step length, preprocessing it and then inputting it into a pre-trained fault prediction model, and generating a fault prediction result for the next time interval.
[0004] However, in the actual application process, although a distributed system is composed of multiple sub-modules, there is a propagation mechanism between each sub-module. Therefore, in the process of fault prediction for distributed systems in the prior art, there are technical problems such as inaccurate fault prediction information and incomplete prediction consideration factors, resulting in inaccurate prediction results and low prediction efficiency when predicting faults in distributed systems, and unable to meet diverse fault requirements. Summary of the Invention
[0005] In order to overcome the above technical problems existing in the prior art, embodiments of the present invention provide a fault prediction method and a prediction device for a distributed system. By establishing a log knowledge graph in the process of log generation of the distributed system, the log generation behavior characteristics of the distributed system are deeply analyzed, so as to effectively identify the fault situation of the distributed system, and effectively improve the accuracy of fault prediction for the distributed system.
[0006] To achieve the above object, an embodiment of the present invention provides a fault prediction method for a distributed system, the method includes: obtaining original log data within a variety of sliding time windows of the distributed system; establishing a log knowledge graph based on the original log data; extracting feature information based on the log knowledge graph; generating a fault prediction model based on the feature information; and performing corresponding fault prediction operations based on the fault prediction model.
[0007] Preferably, the method further includes: before establishing the log knowledge graph, preprocessing the original log data to obtain preprocessed log data; and establishing the log knowledge graph based on the preprocessed log data.
[0008] Preferably, establishing the log knowledge graph based on the preprocessed log data includes: extracting log numbers from each of the preprocessed log data; performing a concatenation operation on the log numbers to determine the call relationship between each of the log numbers; establishing triple data based on the preprocessed log data, the log numbers, and the call relationship; and generating the log knowledge graph based on the triple data.
[0009] Preferably, extracting feature information based on the log knowledge graph includes: performing a transformation process on the preprocessed log data based on a preset vector transformation rule to obtain a first log feature; determining a second log feature according to the multiple sliding time windows; performing a graph representation analysis operation on the log knowledge graph to obtain a third log feature; and generating the feature information based on the first log feature, the second log feature, and the third log feature.
[0010] Preferably, performing a graph representation analysis operation on the log knowledge graph to obtain a third log feature includes: extracting nodes from the log knowledge graph based on a preset node extraction rule to obtain a corresponding node sequence; training a preset graph embedding learning algorithm based on the node sequence to obtain a trained algorithm; and determining the third log feature based on the trained algorithm, where the trained algorithm is represented as: where u represents a node in the log knowledge graph, and N s (u) represents the neighborhood nodes of the u node obtained by sampling method N s and f(u) is the third log feature.
[0011] Preferably, generating a fault prediction model based on the feature information includes: obtaining a preset neural network model, where the preset neural network model is based on a semi-supervised learning algorithm; determining a first anomaly score based on a first calculation rule, where the first anomaly score IS(x) is represented as: where h(x) represents the average depth of the random forest dividing the sample x, and c(n) represents a normalization parameter; determining a second anomaly score based on a second calculation rule, where the second anomaly score SS(x) is represented as: SS(x) = max e -(x-u)2 ; where u represents an anomaly center obtained according to a preset clustering algorithm; obtaining an overall anomaly score based on the first anomaly score and the second anomaly score, where the overall anomaly score TS(x) is represented as: TS(x) = θIS(x) + (1 - θ)ss(x); where θ represents a weight coefficient; processing the feature information based on the overall anomaly score TS(x) to obtain corresponding prediction samples; and training the preset neural network model based on the prediction samples to generate the fault prediction model.
[0012] Correspondingly, an embodiment of the present invention further provides a fault prediction device for a distributed system, and the device includes: a data acquisition unit, configured to acquire original log data within multiple sliding time windows of the distributed system; a knowledge graph construction unit, configured to construct a log knowledge graph based on the original log data; a feature extraction unit, configured to extract feature information based on the log knowledge graph; a model generation unit, configured to generate a fault prediction model based on the feature information; and a fault prediction unit, configured to perform a corresponding fault prediction operation based on the fault prediction model.
[0013] Preferably, the device further includes a data preprocessing unit, and the data preprocessing unit is configured to: preprocess the original log data before constructing the log knowledge graph to obtain preprocessed log data; and construct the log knowledge graph based on the preprocessed log data.
[0014] Preferably, the knowledge graph construction unit includes: a number extraction module, configured to extract a log number from each of the preprocessed log data; a call determination module, configured to perform a concatenation operation on the log numbers to determine a call relationship between each of the log numbers; a triple data generation module, configured to construct triple data based on the preprocessed log data, the log number, and the call relationship; and a knowledge graph generation module, configured to generate the log knowledge graph based on the triple data.
[0015] Preferably, the feature extraction unit includes: a first feature acquisition module, configured to perform a transformation process on the preprocessed log data based on a preset vector transformation rule to obtain a first log feature; a second feature acquisition module, configured to determine a second log feature according to the multiple sliding time windows; a third feature acquisition module, configured to perform a graph representation analysis operation on the log knowledge graph to obtain a third log feature; and a feature information generation module, configured to generate the feature information based on the first log feature, the second log feature, and the third log feature.
[0016] Preferably, performing the graph representation analysis operation on the log knowledge graph to obtain a third log feature includes: extracting nodes from the log knowledge graph based on a preset node extraction rule to obtain a corresponding node sequence; training a preset graph embedding learning algorithm based on the node sequence to obtain a trained algorithm; and determining the third log feature based on the trained algorithm, where the trained algorithm is represented as: where u represents a node in the log knowledge graph, and N s (u) represents the neighborhood nodes of the u node obtained by sampling, and f(u) is the third log feature. s
[0017] Preferably, the model generation unit includes: a preset model acquisition module configured to acquire a preset neural network model based on a semi-supervised learning algorithm; a first score determination module configured to determine a first anomaly score based on a first calculation rule, where the first anomaly score IS(x) is characterized as: where h(x) represents the average depth of the random forest for partitioning the sample x, and c(n) represents a normalization parameter; a second score determination module configured to determine a second anomaly score based on a second calculation rule, where the second anomaly score SS(x) is characterized as: SS(x) = max e -(x-u)2 ; where u represents an anomaly center obtained according to a preset clustering algorithm; a total score determination module configured to obtain an overall anomaly score based on the first anomaly score and the second anomaly score, where the overall anomaly score TS(x) is characterized as: TS(x) = θIS(x) + (1 - θ)ss(x); where θ represents a weight coefficient; a sample determination module configured to process the feature information based on the overall anomaly score TS(x) to obtain a corresponding predicted sample; a model generation module configured to train the preset neural network model based on the predicted sample to generate the fault prediction model.
[0018] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the fault prediction method for a distributed system provided by the embodiment of the present invention.
[0019] Through the technical solution provided by the present invention, the present invention has at least the following technical effects:
[0020] By extracting the original log data of the distributed system during operation based on multiple sliding time windows, the log data in various operating states can be effectively extracted, which can comprehensively represent various operating conditions of the distributed system. On this basis, by establishing a log knowledge graph, extracting the features of the log data, and performing fault prediction through a semi-supervised learning neural network model, the accuracy of fault prediction for the distributed system is greatly improved, meeting the actual needs of enterprises.
[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. Description of the Drawings
[0022] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0023] Figure 1It is a specific implementation flowchart of the fault prediction method for the distributed system provided by the embodiments of the present invention;
[0024] Figure 2 It is a specific implementation flowchart of establishing a log knowledge graph in the fault prediction method for the distributed system provided by the embodiments of the present invention;
[0025] Figure 3 It is a specific implementation flowchart of extracting feature information in the fault prediction method for the distributed system provided by the embodiments of the present invention;
[0026] Figure 4 It is a schematic structural diagram of the fault prediction device for the distributed system provided by the embodiments of the present invention. Detailed implementation manners
[0027] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0028] The terms "system" and "network" in the embodiments of the present invention can be used interchangeably. "Multiple" means two or more. In view of this, in the embodiments of the present invention, "multiple" can also be understood as "at least two". "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the front and rear associated objects unless otherwise specified. In addition, it should be understood that in the description of the embodiments of the present invention, words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0029] Please refer to Figure 1 , the embodiments of the present invention provide a fault prediction method for a distributed system, and the method includes:
[0030] S10) Obtain the original log data within multiple sliding time windows of the distributed system;
[0031] S20) Establish a log knowledge graph based on the original log data;
[0032] S30) Extract feature information based on the log knowledge graph;
[0033] S40) Generate a fault prediction model based on the feature information;
[0034] S50) Perform corresponding fault prediction operations based on the fault prediction model.
[0035] In order to solve the above-mentioned technical problems existing in the background technology, in a possible implementation mode, during the operation of the distributed system, the original log data in multiple sliding time windows is first obtained. For example, multiple sliding time panes are first determined according to a randomly generated method, and the above-mentioned multiple sliding time panes are used as observation periods. At this time, the original log data of the distributed system in the above-mentioned observation period is further obtained, and then the log knowledge graph of the distributed system is established based on the original log data.
[0036] However, in actual applications, since the original log data of the distributed system is in the format of computer communication or data operation, when analyzing the above original log data, the computer may not be able to recognize the accurate log data, thereby reducing the accuracy of subsequent fault prediction.
[0037] In order to solve the above technical problems, in an embodiment of the present invention, the method further includes: before establishing the log knowledge graph, preprocessing the original log data to obtain preprocessed log data; and establishing the log knowledge graph based on the preprocessed log data.
[0038] In a possible implementation, after obtaining the above-mentioned original log data, the original log data is further preprocessed, for example, the original log data is first parsed to obtain a log template and log content, specifically, by deleting invalid characters in the original log data, converting multiple terms in the original log data into standard terms, and uniformly replacing the variables in the original log data, for example, they can be uniformly replaced with the same token, and a log template column is generated. In the process of parsing and generating log content, since the characteristics of log variables have the following rules: the string contains numbers, the string is a meaningless sequence, and the string is located near specific punctuation marks, such as being located in various brackets, quotation marks, after a colon, before and after an operator, etc., in the process of parsing, the text in the original log data can be identified based on regular expressions, word segmentation and other text analysis technologies. On the basis of the above-mentioned generated log template, the log sentence where the template characters are located is removed, and the remaining whole sentence is retained as the log content column, that is, the log content is obtained, thereby converting the original log data into standardized content that can be machine-recognized and processed. At this time, a log knowledge graph is established based on the preprocessed log data obtained after the above preprocessing.
[0039] In an embodiment of the present invention, by converting and processing the format and content of the original log data, the textual and instructional original log data is converted into standardized log data that can be recognized and processed by the machine, thereby facilitating the machine to establish an accurate log knowledge graph of the distributed system on this basis, thereby improving the accuracy of subsequent fault prediction.
[0040] See alsoFigure 2 In the embodiments of the present invention, establishing the log knowledge graph based on the preprocessed log data includes:
[0041] S221) Extracting log numbers from each of the preprocessed log data;
[0042] S222) Performing a concatenation operation on the log numbers to determine the call relationship between each of the log numbers;
[0043] S223) Establishing triple data based on the preprocessed log data, the log numbers, and the call relationship;
[0044] S224) Generating the log knowledge graph based on the triple data.
[0045] In a possible implementation manner, in order to accurately obtain the call relationship between different log sequences in a distributed system, the call relationship is obtained by means of logging. Specifically, a corresponding request ID is generated at each request entry and printed into the corresponding original log data. During the process of establishing the log knowledge graph, first, log numbers are extracted from each preprocessed log data. For example, the log number is the above-mentioned request ID. Then, the above log numbers are concatenated to extract the call relationship of log sequences within the above multiple sliding time windows. The above call relationship includes, but is not limited to, call relationships such as HTTP request service / client, RPC request service / client, database access, middleware call, and local method call. Then, according to the callers and callees of the above existing call relationships in the log sequences as entities, with the log sequence ID as the entity ID, and combined with the above call relationship, triple data of "entity - relationship - entity ID" is generated, and all the triple data is imported into a graph database to generate a log knowledge graph.
[0046] In the embodiments of the present invention, by collecting and processing the call relationship between original log data, and creating a log knowledge graph of log data based on the above call relationship, during the subsequent fault prediction process, the call relationship between each log data can be effectively combined to accurately predict the faults of the distributed system, rather than simply performing fault analysis on the original log data, thereby greatly improving the prediction accuracy of fault prediction for the distributed system.
[0047] Please refer to Figure 3 In the embodiments of the present invention, extracting feature information based on the log knowledge graph includes:
[0048] S31) Performing transformation processing on the preprocessed log data based on a preset vector transformation rule to obtain first log features;
[0049] S32) Determine the second log feature according to the multiple sliding time windows;
[0050] S33) Perform a graph representation analysis operation on the log knowledge graph to obtain a third log feature;
[0051] S34) Generate the feature information based on the first log feature, the second log feature, and the third log feature.
[0052] After creating the log knowledge graph, in order to accurately predict the faults of the distributed system, it is necessary to extract and analyze various features in the log data of the distributed system. In a possible implementation manner, first, the preprocessed log data is transformed through a preset vector transformation rule. For example, through the word2vec technology, the log template column can be transformed into a vector to obtain the log text feature, that is, the first log feature is obtained. Then, according to the above multiple sliding time windows, the time series features within the time corresponding to each sliding time window are calculated, and the time series features are used as the log statistical features, that is, the second log feature is obtained. Then, further obtain the third log feature according to the log knowledge graph. For example, using the graph embedding learning technology, the graph representation features of the nodes in the log graph are extracted.
[0053] For example, in the embodiment of the present invention, the performing a graph representation analysis operation on the log knowledge graph to obtain a third log feature includes: extracting nodes from the log knowledge graph based on a preset node extraction rule to obtain a corresponding node sequence; training a preset graph embedding learning algorithm based on the node sequence to obtain a trained algorithm; determining the third log feature based on the trained algorithm, and the trained algorithm is characterized as: where u represents a node in the log knowledge graph, and N s (u) represents the neighborhood nodes of node u obtained by sampling method N s and f(u) is the third log feature.
[0054] In a possible implementation manner, first, nodes are extracted from the log knowledge graph based on a preset node extraction rule. For example, an improved random walk strategy can be used to generate a node sequence of the log knowledge graph. Then, the preset graph embedding learning algorithm is trained through the above node sequence. For example, in the skip-gram manner, sample pairs are generated according to the above node sequence, and the sample pairs are input into the Node2Vec algorithm, and the node vector representation is obtained from the hidden layer of the Node2Vec algorithm. Specifically, the algorithm can be characterized as: where the feature vector of node u is characterized as f(u), that is, the third log feature is obtained. At this time, the corresponding feature information is generated according to the above first log feature, second log feature, and third log feature.
[0055] In an embodiment of the present invention, by analyzing and determining the characteristics of log data in various dimensions, and generating corresponding characteristic information according to the characteristics of the above-mentioned multiple dimensions, the above-mentioned characteristic information can effectively represent each characteristic in the log behavior during the operation of the distributed system. Based on the above-mentioned characteristics, possible faults of the distributed system can be effectively analyzed, thereby improving the accuracy of fault prediction.
[0056] Further, in an embodiment of the present invention, the generating a fault prediction model based on the characteristic information includes: obtaining a preset neural network model, where the preset neural network model is based on a semi-supervised learning algorithm; determining a first anomaly score based on a first calculation rule, and the first anomaly score IS(x) is characterized as: where h(x) represents the average depth of the random forest dividing the sample x, and c(n) represents a normalization parameter; determining a second anomaly score based on a second calculation rule, and the second anomaly score SS(x) is characterized as: SS(x) = max e -(x-u)2 ; where u represents an anomaly center obtained according to a preset clustering algorithm; obtaining an overall anomaly score based on the first anomaly score and the second anomaly score, and the overall anomaly score TS(x) is characterized as: TS(x) = θIS(x) + (1 - θ)ss(x); where θ represents a weight coefficient; processing the characteristic information based on the overall anomaly score TS(x) to obtain corresponding prediction samples; training the preset neural network model based on the prediction samples to generate the fault prediction model.
[0057] In a possible implementation manner, in order to further improve the prediction efficiency and prediction accuracy of the faults of the distributed system, automatic analysis is performed by creating a neural network model based on a semi-supervised learning algorithm. In an embodiment of the present invention, first, a preset neural network model is obtained. For example, the preset neural network model is generated based on the improved ADOA (Anomaly Detection with Partially Observed Anomalies). By inputting the above-mentioned characteristic information to train the preset neural network model, a fault prediction model is generated.
[0058] Specifically, first, a first anomaly score is determined according to the first calculation rule. For example, the first anomaly score IS(x) is characterized as: where h(x) represents the average depth of the random forest dividing the sample x, and c(n) represents a normalization parameter; then, a second anomaly score is further determined according to the second calculation rule, and the second anomaly score SS(x) is characterized as: SS(x) = max e -(x-u)2; where u represents the anomaly center obtained according to a preset clustering algorithm; then add the above first anomaly score IS(x) and the second anomaly score SS(x) to obtain the total anomaly score, and the total anomaly score TS(x) is characterized as:
[0059] TS(x) = θIS(x) + (1 - θ)ss(x); where θ represents the weight coefficient, for example, the weight coefficient of the first anomaly score, and the value of θ is [0, 1], which is used to balance the importance of the first anomaly score and the second anomaly score.
[0060] At this time, further divide the feature information by a threshold to obtain reliable positive and negative samples. For example, mark the feature information with an anomaly total score higher than the threshold as a fault sample. At this time, use the existing positive samples and the reliable negative samples to form a corresponding sample set, that is, obtain the prediction samples, and then input the prediction samples into the above preset neural network model for training to obtain the final fault prediction model. At this time, through the above fault prediction model, the faults of the distributed system can be accurately and reliably predicted.
[0061] In the embodiment of the present invention, by adopting the fault prediction method for the distributed system based on the log knowledge graph, fault analysis is performed according to the characteristics of the logs generated during the operation of the distributed system, and the faults in the distributed system can be accurately and effectively predicted, greatly improving the prediction accuracy in the fault prediction of the distributed system and meeting the actual needs of enterprises.
[0062] The following will describe the fault prediction device for the distributed system provided by the embodiment of the present invention with reference to the accompanying drawings.
[0063] Please refer to Figure 4 , based on the same inventive concept, the embodiment of the present invention provides a fault prediction device for a distributed system, and the device includes: a data acquisition unit for acquiring the original log data within a plurality of sliding time windows of the distributed system; a knowledge graph building unit for building a log knowledge graph based on the original log data; a feature extraction unit for extracting feature information based on the log knowledge graph; a model generation unit for generating a fault prediction model based on the feature information; and a fault prediction unit for performing corresponding fault prediction operations based on the fault prediction model.
[0064] In the embodiment of the present invention, the device further includes a data preprocessing unit, and the data preprocessing unit is used for: before building the log knowledge graph, preprocessing the original log data to obtain the preprocessed log data; and building the log knowledge graph based on the preprocessed log data.
[0065] In an embodiment of the present invention, the knowledge graph building unit includes: a number extraction module, configured to extract a log number from each of the preprocessed log data; a call determination module, configured to perform a concatenation operation on the log numbers to determine the call relationship between each of the log numbers; a triple data generation module, configured to build triple data based on the preprocessed log data, the log numbers, and the call relationship; and a knowledge graph generation module, configured to generate the log knowledge graph based on the triple data.
[0066] In an embodiment of the present invention, the feature extraction unit includes: a first feature acquisition module, configured to perform a transformation process on the preprocessed log data based on a preset vector transformation rule to obtain a first log feature; a second feature acquisition module, configured to determine a second log feature according to the multiple sliding time windows; a third feature acquisition module, configured to perform a graph representation analysis operation on the log knowledge graph to obtain a third log feature; and a feature information generation module, configured to generate the feature information based on the first log feature, the second log feature, and the third log feature.
[0067] In an embodiment of the present invention, the performing a graph representation analysis operation on the log knowledge graph to obtain a third log feature includes: performing node extraction on the log knowledge graph based on a preset node extraction rule to obtain a corresponding node sequence; training the preset graph embedding learning algorithm based on the node sequence to obtain a trained algorithm; and determining the third log feature based on the trained algorithm, where the trained algorithm is represented as: where u represents a node in the log knowledge graph, and N s (u) represents the neighborhood nodes of the u node obtained by sampling method N s and f(u) is the third log feature.
[0068] In an embodiment of the present invention, the model generation unit includes: a preset model acquisition module, configured to acquire a preset neural network model, where the preset neural network model is based on a semi-supervised learning algorithm; a first score determination module, configured to determine a first anomaly score based on a first calculation rule, where the first anomaly score IS(x) is represented as: where h(x) represents the average depth of the random forest dividing the sample x, and c(n) represents a normalization parameter; a second score determination module, configured to determine a second anomaly score based on a second calculation rule, where the second anomaly score SS(x) is represented as: SS(x) = max e -(x-u)2; where u represents the anomaly center obtained according to a preset clustering algorithm; a total score determination module, configured to obtain an anomaly total score based on the first anomaly score and the second anomaly score, and the anomaly total score TS(x) is represented as: TS(x) = θIS(x) + (1 - θ)ss(x); where θ represents a weight coefficient; a sample determination module, configured to process the feature information based on the anomaly total score TS(x) to obtain a corresponding predicted sample; a model generation module, configured to train the preset neural network model based on the predicted sample to generate the fault prediction model.
[0069] Further, an embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the fault prediction method of the distributed system described in the embodiment of the present invention.
[0070] The above has described in detail the optional embodiments of the embodiments of the present invention with reference to the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention.
[0071] In addition, it should be noted that, among the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without conflict. To avoid unnecessary repetition, the embodiments of the present invention do not separately describe various possible combination methods.
[0072] Those skilled in the art can understand that all or part of the steps of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a program, and the program is stored in a storage medium, including several instructions for causing a single-chip microcomputer, a chip, or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.
[0073] In addition, any combination can be made between various different embodiments of the embodiments of the present invention as long as it does not violate the idea of the embodiments of the present invention, and it should also be regarded as the content disclosed in the embodiments of the present invention.
Claims
1. A fault prediction method for a distributed system, characterized in that The method includes: Obtaining the original log data within multiple sliding time windows of a distributed system; Preprocessing the original log data to obtain preprocessed log data; Building a log knowledge graph based on the preprocessed log data; Extracting feature information based on the log knowledge graph; Generating a fault prediction model based on the feature information; Performing corresponding fault prediction operations based on the fault prediction model; The extracting feature information based on the log knowledge graph includes: Performing transformation processing on the preprocessed log data according to a preset vector transformation rule to obtain first log features; Determining second log features according to the multiple sliding time windows; Performing graph representation analysis operations on the log knowledge graph to obtain third log features; Generating the feature information based on the first log features, the second log features, and the third log features; The generating a fault prediction model based on the feature information includes: Obtaining a preset neural network model, where the preset neural network model is based on a semi-supervised learning algorithm; Determining a first anomaly score based on a first calculation rule, and the first anomaly score IS(x) is characterized as: Among them, h(x) represents the average depth of the random forest for partitioning the sample x, and c(n) represents the normalization parameter; Determine a second anomaly score based on a second calculation rule, and the second anomaly score SS(x) is characterized as: SS(x) = maxe -(x-u)2 ; where u represents an anomaly center obtained according to a preset clustering algorithm; Obtaining an overall anomaly score based on the first anomaly score and a second anomaly score, and the overall anomaly score TS(x) is characterized as: TS(x) = θIS(x) + (1 - θ)ss(x); where θ represents a weight coefficient; Processing the feature information based on the overall anomaly score TS(x) to obtain corresponding prediction samples; Training the preset neural network model based on the prediction samples to generate the fault prediction model.
2. The method according to claim 1, wherein The building the log knowledge graph based on the preprocessed log data includes: Extracting log numbers from each of the preprocessed log data; Performing a concatenation operation on the log numbers to determine the call relationship between each of the log numbers; Building triple data based on the preprocessed log data, the log numbers, and the call relationship; Generating the log knowledge graph based on the triple data.
3. The method according to claim 1, wherein The performing graph representation analysis operations on the log knowledge graph to obtain third log features includes: Performing node extraction on the log knowledge graph according to a preset node extraction rule to obtain a corresponding node sequence; Training a preset graph embedding learning algorithm based on the node sequence to obtain a trained algorithm; Determining the third log features based on the trained algorithm, and the trained algorithm is characterized as: Among them, u represents a node in the log knowledge graph, N S (u) represents the domain nodes of the u node obtained by sampling method N S The resulting u node's domain nodes, and f(u) is the third log feature.
4. A fault prediction device for a distributed system, characterized in that, The apparatus includes: A data acquisition unit for obtaining the original log data within multiple sliding time windows of a distributed system; A data preprocessing unit for preprocessing the original log data to obtain preprocessed log data; A knowledge graph building unit for building a log knowledge graph based on the preprocessed log data; A feature extraction unit for extracting feature information based on the log knowledge graph; A model generation unit for generating a fault prediction model based on the feature information; A fault prediction unit for performing corresponding fault prediction operations based on the fault prediction model; The feature extraction unit includes: The first feature acquisition module is used to perform transformation processing on the preprocessed log data based on a preset vector transformation rule to obtain the first log feature; The second feature acquisition module is used to determine the second log feature according to the multiple sliding time windows; The third feature acquisition module is used to perform graph representation analysis operations on the log knowledge graph to obtain the third log feature; The feature information generation module is used to generate the feature information based on the first log feature, the second log feature, and the third log feature; The model generation unit includes: The preset model acquisition module is used to acquire a preset neural network model, and the preset neural network model is based on a semi-supervised learning algorithm; The first scoring determination module determines a first anomaly score based on a first calculation rule, and the first anomaly score IS(x) is characterized as: where h(x) represents the average depth of the random forest for partitioning the sample x, and c(n) represents a normalization parameter; The second scoring determination module is used to determine a second anomaly score based on a second calculation rule, and the second anomaly score SS(x) is characterized as: SS(x) = maxe -(x-u)2 ; where u represents the anomaly center obtained according to a preset clustering algorithm; The total score determination module is used to obtain an abnormal total score based on the first abnormal score and the second abnormal score. The abnormal total score TS(x) is characterized as: TS(x) = θIS(x) + (1 - θ)ss(x); where θ represents the weight coefficient; The sample determination module is used to process the feature information based on the abnormal total score TS(x) to obtain corresponding prediction samples; The model generation module is used to train the preset neural network model based on the prediction samples to generate the fault prediction model.
5. The device according to claim 4, characterized in that The knowledge graph establishment unit includes: The number extraction module is used to extract log numbers from each of the preprocessed log data; The call determination module is used to perform a concatenation operation on the log numbers to determine the call relationship between each of the log numbers; The triple data generation module is used to establish triple data based on the preprocessed log data, the log numbers, and the call relationship; The knowledge graph generation module is used to generate the log knowledge graph based on the triple data.
6. The device according to claim 4, characterized in that, The performing graph representation analysis operations on the log knowledge graph to obtain the third log feature includes: Performing node extraction on the log knowledge graph based on a preset node extraction rule to obtain a corresponding node sequence; Training a preset graph embedding learning algorithm based on the node sequence to obtain a trained algorithm; Determining the third log feature based on the trained algorithm, and the trained algorithm is characterized as: Among them, u represents a node in the log knowledge graph, and N S (u) represents the domain nodes of the u node obtained by sampling method N S The f(u) is the third log feature.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the fault prediction method of the distributed system described in any one of claims 1-3.
Citation Information
Patent Citations
Log analysis early warning method based on deep learning
CN110958136A
Power equipment fault knowledge graph construction method
CN111737496A