Cloud platform alarm root cause positioning method, device and equipment and storage medium
By cleaning and feature extraction of the alarm data and logs of the cloud platform, building a Bayesian network and performing causal reasoning, the accuracy and efficiency of alarm root cause positioning on the cloud platform are solved, and efficient alarm root cause positioning and operation and maintenance support are achieved.
Patent Information
- Application Number
- CN202510482109.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology is difficult to quickly and accurately locate the alarm cause on cloud platforms, resulting in uneven positioning results and low efficiency, and it is impossible to cope with complex operating environment changes.
By cleaning the alarm data and the operation log, extracting feature data and building a Bayesian network, establishing a causal inference model, using F1 value to evaluate the model performance, and locate the root cause of the alarm in real time.
It improves the accuracy of alarm root cause positioning on cloud platforms, supports efficient operation and maintenance, and improves the stability and reliability of cloud platforms.
Smart Images

Figure CN120276903A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a method, device, equipment and storage medium for locating the root cause of alarms in a cloud platform. Background Art
[0002] With the rapid development of cloud computing technology, cloud platforms have become an important part of enterprise IT infrastructure. However, during the operation of cloud platforms, a large amount of alarm information is generated, and how to quickly and accurately locate the root cause of alarms has become a major challenge for operation and maintenance personnel.
[0003] Existing methods mainly rely on manual experience or simple rules. The method dominated by manual experience is greatly affected by individual differences among operation and maintenance personnel. Different operation and maintenance personnel have different familiarity with each component of the cloud platform and different fault handling experiences, which makes the results of locating the root cause of alarms uneven. For the method of simple rules, it usually judges the root cause of alarms based on fixed condition matching. However, the operating environment of cloud platforms is dynamically variable, and new fault modes will continuously emerge. Therefore, existing technologies are difficult to handle complex scenarios and have low location efficiency.
[0004] Therefore, how to improve the accuracy of locating the root cause of alarms on cloud platforms is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for locating the root cause of alarms in a cloud platform, which can improve the accuracy of locating the root cause of alarms on the cloud platform. The specific scheme is as follows:
[0006] In a first aspect, the present application provides a method for locating the root cause of alarms in a cloud platform, including:
[0007] Clean the alarm data and operation logs on the target cloud platform, extract target feature data from the obtained cleaned data, and divide the cleaned data and the target feature data into a training set and a test set;
[0008] Extract the running state of the cloud platform and alarm events from the training set, use the running state of the cloud platform as the alarm root cause node and the alarm events as the alarm result nodes, and establish connections between the alarm root cause nodes and the alarm result nodes with causal relationships to obtain a Bayesian network;
[0009] Construct an initial causal inference model based on the Bayesian network, and use the initial causal inference model and the test set to determine the F1 value. When the F1 value meets the preset performance conditions, determine the initial causal inference model as the target causal inference model;
[0010] During the operation of the target cloud platform, the current alarm data is determined in real time, and the root cause of the alarm is located based on the current alarm data and the target causal inference model.
[0011] Optionally, the data cleaning of the alarm data and operation logs on the target cloud platform includes:
[0012] Obtain the alarm data and operation logs from the target cloud platform;
[0013] Use the Z-score method to determine the first outliers from the alarm data and the operation logs, and determine the second outliers from the alarm data and the operation logs based on the Isolation Forest algorithm;
[0014] Remove the first outliers and the second outliers from the alarm data and the operation logs to complete the data cleaning operation.
[0015] Optionally, the extraction of the target feature data from the obtained cleaned data includes:
[0016] Use the principal component analysis algorithm to reduce the dimensionality of the data in the obtained cleaned data whose data dimension is greater than the preset dimension threshold, so as to extract the target feature data based on the obtained data after dimensionality reduction;
[0017] Extract the alarm time feature, alarm type feature, alarm level feature, and alarm location feature from the cleaned data based on the random forest algorithm and the linear regression algorithm with the L1 regularization term; wherein, the alarm time feature includes the moment feature when the alarm occurs and the duration feature of the alarm;
[0018] Determine the alarm time feature, the alarm type feature, the alarm level feature, and the alarm location feature as the target feature data.
[0019] Optionally, the division of the cleaned data and the target feature data into a training set and a test set includes:
[0020] Divide the cleaned data and the target feature data into a training set and a test set based on the k-fold cross-validation technique.
[0021] Optionally, taking the operating state of the cloud platform as the alarm root cause node, taking the alarm event as the alarm result node, and establishing a connection line between the alarm root cause node and the alarm result node with a causal relationship to obtain a Bayesian network, includes:
[0022] Taking the operating status of the cloud platform as the alarm root cause node and the alarm event as the alarm result node, and assigning corresponding first weight values to the alarm root cause node and the alarm result node by calculating the information gain of the alarm root cause node and the alarm result node to complete node configuration;
[0023] Establish a connection between the alarm root cause node and the alarm result node with a causal relationship, determine the conditional probability of the causal relationship, and assign a corresponding second weight value to the connection using the conditional probability to complete the connection configuration;
[0024] Determine the Bayesian network based on the node configuration and the connection configuration.
[0025] Optionally, the step of determining the F1 value using the initial causal inference model and the test set, and when the F1 value meets the preset performance condition, determining the initial causal inference model as the target causal inference model includes:
[0026] Input the test set into the initial causal inference model so that the initial causal inference model outputs the alarm root cause;
[0027] Statistical quantity information of the alarm root cause, and determine the F1 value based on the quantity information and the actual alarm root cause quantity information;
[0028] Judge whether the F1 value is not less than the preset performance threshold and obtain a judgment result;
[0029] If the judgment result indicates not less than, then determine the initial causal inference model as the target causal inference model;
[0030] If the judgment result indicates less than, then adjust the model parameters of the initial causal inference model and optimize the target feature data based on the preset feature optimization rules.
[0031] Optionally, the step of locating the alarm root cause based on the current alarm data and the target causal inference model includes:
[0032] Input the current alarm data into the target causal inference model so that the target causal inference model generates each alarm root cause and the probability value corresponding to each alarm root cause based on the current alarm data.
[0033] In a second aspect, the present application provides a cloud platform alarm root cause location device, including:
[0034] A data partitioning module for cleaning the alarm data and operation logs on the target cloud platform, extracting target feature data from the obtained cleaned data, and partitioning the cleaned data and the target feature data into a training set and a test set;
[0035] A network establishment module, configured to extract the operation status of the cloud platform and alarm events from the training set, use the operation status of the cloud platform as the alarm root cause node, use the alarm events as the alarm result nodes, and establish connections between the alarm root cause nodes and the alarm result nodes with a causal relationship to obtain a Bayesian network;
[0036] A model determination module, configured to construct an initial causal inference model based on the Bayesian network, determine the F1 value by using the initial causal inference model and the test set, and when the F1 value meets a preset performance condition, determine the initial causal inference model as the target causal inference model;
[0037] An alarm location module, configured to, during the operation of the target cloud platform, determine current alarm data in real time, and perform alarm root cause location based on the current alarm data and the target causal inference model.
[0038] In a third aspect, the present application provides an electronic device, including:
[0039] A memory, configured to store a computer program;
[0040] A processor, configured to execute the computer program to implement the foregoing cloud platform alarm root cause location method.
[0041] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the foregoing cloud platform alarm root cause location method is implemented.
[0042] In this application, data cleaning is performed on the alarm data and operation logs on the target cloud platform, target feature data is extracted from the obtained cleaned data, and the cleaned data and the target feature data are divided into a training set and a test set; the cloud platform operation status and alarm events are extracted from the training set, the cloud platform operation status is used as the alarm root cause node, the alarm event is used as the alarm result node, and a connection is established between the alarm root cause node and the alarm result node with a causal relationship to obtain a Bayesian network; an initial causal inference model is constructed based on the Bayesian network, and the F1 value is determined using the initial causal inference model and the test set. When the F1 value meets the preset performance condition, the initial causal inference model is determined as the target causal inference model; during the operation of the target cloud platform, the current alarm data is determined in real time, and the alarm root cause is located based on the current alarm data and the target causal inference model. As can be seen from the above, in this application, first, data cleaning work is carried out on the alarm data and operation logs generated by the target cloud platform. After completing the data cleaning, the target feature data is extracted from the obtained cleaned data. Then, this part of the cleaned data and the extracted target feature data are divided into a training set and a test set. After that, the operation status information of the cloud platform and the content related to alarm events are refined from the divided training set. The operation status of the cloud platform is set as the alarm root cause node, and the alarm event is set as the alarm result node. For those alarm root cause nodes and alarm result nodes with a causal relationship, a connection is established between these nodes to construct a Bayesian network to represent the causal relationship between the nodes. An initial causal inference model is further built based on the constructed Bayesian network. The performance of this initial causal inference model is evaluated using the previously divided test set, and the performance of the model is measured by calculating the F1 value. When the calculated F1 value meets the preset performance condition, this initial causal inference model can be determined as the target causal inference model. Finally, during the actual operation of the target cloud platform, the currently newly generated alarm data is continuously and real-time determined. These real-time alarm data are input into the already determined target causal inference model, and the target causal inference model is used to analyze and reason about the current alarm data to achieve accurate positioning of the alarm root cause. In this way, this application can improve the accuracy of alarm root cause positioning on the cloud platform and provide certain support for the efficient operation and maintenance of the cloud platform. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0044] Figure 1 Flowchart of a method for locating the root cause of alarms in a cloud platform disclosed in this application;
[0045] Figure 2 Schematic structural diagram of a device for locating the root cause of alarms in a cloud platform disclosed in this application;
[0046] Figure 3 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] Existing cloud platform alarm root cause location solutions mainly rely on manual experience or simple rules. The method dominated by manual experience is greatly affected by the individual differences of operation and maintenance personnel. Different operation and maintenance personnel have different familiarity with each component of the cloud platform and different fault handling experiences, which makes the alarm root cause location results uneven. And the method of simple rules usually judges the alarm root cause based on fixed condition matching. However, the operating environment of the cloud platform is dynamically changing, and new fault modes will emerge continuously. Therefore, the existing technologies are difficult to cope with complex scenarios and have low location efficiency. For this reason, this application provides a method, device, equipment and storage medium for locating the root cause of alarms in a cloud platform, which can improve the accuracy of locating the root cause of alarms on the cloud platform.
[0049] Refer to Figure 1 As shown, an embodiment of the present invention discloses a method for locating the root cause of alarms in a cloud platform, including:
[0050] Step S11: Clean the alarm data and operation logs on the target cloud platform, extract target feature data from the obtained cleaned data, and divide the cleaned data and the target feature data into a training set and a test set.
[0051] In this embodiment, first, alarm data and operation logs are obtained from the target cloud platform. Among them, during the daily operation of the target cloud platform, a large amount of alarm data will be generated, and this alarm data contains information such as the timestamp when the alarm occurs and the alarm type. At the same time, the operation logs record various status information during the operation of the target cloud platform, such as resource utilization rate, network traffic, etc.
[0052] After obtaining the alarm data and operation logs, one of the steps in data cleaning is to remove outliers. Specifically, the Z-score method is used to determine the first outliers from the alarm data and operation logs. The Z-score method is an outlier detection method based on statistical metrics, and its principle is to judge whether a data point is an outlier by calculating the multiple of the standard deviation of each data point from the mean. For each numerical data in the alarm data and operation logs, calculate its Z-score value. When the Z-score value of a certain data point exceeds a certain multiple, it is determined as the first outlier. Among them, the multiple can be 3 times. In this way, the outlier data that is significantly different from the overall data distribution can be effectively identified.
[0053] Meanwhile, based on the Isolation Forest (i.e., iForest) algorithm, the second outliers are determined from the alarm data and operation logs. It should be noted that the Isolation Forest algorithm is an outlier detection algorithm based on the idea of clustering or classification, and it divides data points by constructing a tree structure. Therefore, in this embodiment, based on the Isolation Forest algorithm, the path length of each data point in the tree reflects its difference degree from other data points, and the data points with shorter paths are more likely to be outliers. By applying the Isolation Forest algorithm to the alarm data and operation logs, the outlier data that is relatively isolated in the data distribution, that is, the second outliers, can be further discovered.
[0054] Next, after determining the first outliers and the second outliers, the first outliers and the second outliers are removed from the alarm data and operation logs to complete the outlier removal step in the data cleaning operation. It can be understood that removing these outliers can avoid their interference with subsequent model training and analysis and improve the quality and reliability of the data.
[0055] In addition, in this embodiment, missing values are a common problem in data cleaning, and the methods for handling missing values depend on the characteristics of the data and the reasons for the missing values. For numerical data, if its distribution is relatively uniform and the number of missing values is small, the mean filling method is adopted, that is, calculate the average value of all non-missing values in this data column and fill the missing values with this average value; if the data distribution is skewed and the median is less affected by extreme values, the median filling method is adopted, and the median of this data column is used to fill the missing values. When there is a certain correlation between data features and the amount of data is relatively sufficient, use the model-based prediction filling method, such as constructing a linear regression model or a decision tree model, etc., to predict the missing values according to the values of other relevant features. For time series data, if there are missing values, interpolation methods such as linear interpolation can be used to estimate the missing values according to the values of adjacent data points to ensure the continuity and integrity of the data in the time dimension.
[0056] In addition, the data may have inconsistent formats and need to be processed uniformly. For example, the format of the timestamp may be different in different systems or tools. This embodiment converts all timestamps into a unified format, such as the ISO8601 (International Organization for Standardization, data storage and exchange format, information exchange, and date and time representation method) standard. There may also be differences in the representation of log levels. The log levels of different systems or tools can be mapped to a unified standard, such as DEBUG (debugging), INFO (system status), WARN (repairable problems), ERROR (system errors), FATAL (serious errors), etc., so that when analyzing log data, the information conveyed by the log and its importance can be more clearly understood.
[0057] Furthermore, after completing data cleaning operations such as outlier removal, missing value processing, and format unification, in order to further process the data, the principal component analysis (PCA) algorithm is used to reduce the dimensionality of the data in the cleaned data whose data dimension is greater than the preset dimensionality threshold. It should be pointed out that in the actual cloud platform data, there may be high-dimensional data, which not only increases the complexity of calculation, but also may introduce noise and redundant information. Therefore, the principal component analysis algorithm projects the original data into a low-dimensional space by performing a linear transformation on the data, while retaining the main features of the data as much as possible. By setting an appropriate preset dimensionality threshold, the data with higher dimensions in the cleaned data is reduced in dimensionality, so as to extract the target feature data based on the obtained dimensionality-reduced data.
[0058] On the basis of dimensionality reduction, the alarm time feature, alarm type feature, alarm level feature and alarm location feature are extracted from the cleaned data based on the random forest algorithm and the linear regression algorithm with the L1 regularization term (i.e. Lasso regression, Least Absolute Shrinkage and Selection Operator). Among them, the L1 regularization term refers to the sum of the absolute values of each element in the weight vector. The random forest algorithm is an integrated learning algorithm that improves the accuracy and stability of the model by constructing multiple decision trees and synthesizing their results. Therefore, the random forest algorithm can effectively mine various feature information related to alarms from the cleaned data. At the same time, the linear regression algorithm can screen and reduce the dimension of features by introducing the L1 regularization term in the linear regression model.
[0059] Specifically, the alarm time feature includes the moment feature when the alarm occurs and the duration feature of the alarm. The moment when the alarm occurs is of great significance for analyzing the cause of the alarm and the time series pattern, while the duration of the alarm reflects the severity and impact scope of the alarm. Through the random forest algorithm and the linear regression algorithm, these time features can be accurately extracted from the cleaned data. The alarm type feature includes but is not limited to CPU overload, memory shortage, network bandwidth congestion, etc. Different alarm types correspond to different failure causes of the cloud platform. Accurately extracting the alarm type feature is a key step in locating the root cause of the alarm. The alarm level feature can reflect the urgency and importance of the alarm, and the alarm location feature indicates the specific location where the alarm occurs or the cloud service components involved, which helps to locate the specific source of the failure. Then, the alarm time feature, the alarm type feature, the alarm level feature, and the alarm location feature are determined as the target feature data.
[0060] Finally, based on the k-fold cross-validation technique, the cleaned data and the target feature data are divided into a training set and a test set. K-fold cross-validation is a commonly used data partitioning and model evaluation technique. It divides the data set into k non-overlapping subsets. Each time, k - 1 of these subsets are selected as the training set, and the remaining one subset is used as the test set. This process is repeated k times, and finally, the results of the k validations are averaged. In this way, the data can be utilized more fully, avoiding the evaluation bias caused by the randomness of data set partitioning and improving the accuracy of model evaluation. The divided training set is used for subsequent training of the causal inference model to enable it to learn the causal relationship patterns in the data, while the test set is used to evaluate the trained model to examine the generalization ability and performance of the model.
[0061] Step S12: Extract the cloud platform running status and alarm events from the training set. Use the cloud platform running status as the alarm root cause node and the alarm events as the alarm result nodes, and establish connections between the alarm root cause nodes and the alarm result nodes with causal relationships to obtain a Bayesian network.
[0062] In this embodiment, the training set contains high-quality data that has been cleaned and feature-extracted. Extracting the cloud platform running status and alarm events from it is the basis for constructing the Bayesian network. The cloud platform running status covers various metrics such as CPU usage rate, memory utilization rate, network bandwidth, disk I / O, etc. These status information reflect the real-time running situation of the cloud platform. The alarm events are various alarms triggered during the operation of the cloud platform, such as CPU overload alarm, memory shortage alarm, network connection interruption alarm, etc. Through in-depth analysis and mining of the training set data, these cloud platform running status and alarm events are accurately identified, providing a basis for subsequent node construction.
[0063] Furthermore, a node system of the Bayesian network is constructed with the operating status of the cloud platform as the warning root cause node and the warning event as the warning result node. After determining the nodes, corresponding weight values need to be assigned to each node to reflect its importance in the entire causal relationship system. This goal is achieved by calculating the information gain between the warning root cause node and the warning result node.
[0064] Specifically, for the warning root cause node and the warning result node, by calculating their information gain, the importance of the information represented by this node for causal relationship judgment can be evaluated. The greater the information gain, the more significant the impact of this node on the causal relationship, and the higher its importance in the network. For example, for the warning root cause node of CPU usage rate, if its information gain is high, it indicates that the change in CPU usage rate has a greater impact on the occurrence of the warning event. Then, when configuring the node, a higher first weight value will be assigned to it. In this way, corresponding first weight values are assigned to each warning root cause node and warning result node, enabling the Bayesian network to more accurately reflect the importance of each node in the causal relationship.
[0065] After completing the node configuration, connections need to be established between the warning root cause nodes and warning result nodes with causal relationships to represent the causal connections between them. Determining the conditional probability of the causal relationship is a key step in connection configuration. The conditional probability can reflect the likelihood of the warning result node occurring under the condition that a certain warning root cause node occurs.
[0066] For example, when the CPU usage rate is too high, the probability of the server response timeout warning occurring is a conditional probability. Through statistical analysis of the training set data, these conditional probabilities can be calculated. After calculating the conditional probabilities, the corresponding second weight values are assigned to the connections using this conditional probability. The greater the conditional probability, the stronger the causal relationship, and the higher the second weight value of the connection. For example, if the probability of the server response timeout warning occurring is very high when the CPU usage rate is too high, then the connection from the "CPU usage rate too high" node to the "server response timeout warning" node will be assigned a higher second weight value. In this way, the connection configuration is completed, enabling the Bayesian network to more accurately describe the strength of the causal relationship between each node.
[0067] Based on the above node configuration and connection configuration, the Bayesian network is finally determined. It can be understood that the Bayesian network integrates the causal relationship between the operating status of the cloud platform and the warning event, and through the weight values of the nodes and the weight values of the connections, can accurately reflect the importance and causal relationship strength of each factor in warning root cause location.
[0068] Step S13: Construct an initial causal inference model based on the Bayesian network, and use the initial causal inference model and the test set to determine the F1 value. When the F1 value meets the preset performance condition, determine the initial causal inference model as the target causal inference model.
[0069] In this embodiment, the test set is input into the initial causal inference model so that the initial causal inference model outputs the alarm root cause. The test set contains various types of alarm data and corresponding running status information. After the test set is input into the initial causal inference model, the initial causal inference model analyzes and reasons the input data according to the causal relationship learned in the Bayesian network and outputs possible alarm root causes.
[0070] Next, count the quantity information of the alarm root causes, and determine the F1 value based on the quantity information and the actual alarm root cause quantity information. Among them, the F1 value is an index that comprehensively considers the accuracy rate and the recall rate, and it can more comprehensively evaluate the performance of the model. The accuracy rate reflects the proportion of the number of alarm root causes correctly predicted by the model in the total predicted number, that is, the accuracy of the model prediction, while the recall rate reflects the proportion of the number of actual alarm root causes that the model can correctly identify in the total number of actual alarm root causes, reflecting the coverage of the model for the true root cause.
[0071] When specifically calculating the F1 value, it is first necessary to determine the number of alarm root causes correctly predicted by the model, the number of alarm root causes wrongly predicted by the model, and the number of actual alarm root causes missed by the model, and then calculate according to these three quantities to obtain the accuracy rate and the recall rate.
[0072] Furthermore, determine whether the F1 value is not less than the preset performance threshold and obtain the judgment result. The preset performance threshold can be a standard value set according to actual requirements and experience, which can represent the minimum performance level that the initial causal inference model needs to achieve in actual applications. If the F1 value is not less than the preset performance threshold, it means that the initial causal inference model has achieved a good balance between the accuracy rate and the recall rate and can accurately locate the alarm root cause. That is to say, if the judgment result shows not less than, determine the initial causal inference model as the target causal inference model. At this time, the initial causal inference model has been verified by the test set and has high accuracy and reliability, and can be used in the actual cloud platform operation and maintenance.
[0073] If the judgment result indicates less than, it means that the performance of the initial causal inference model has not yet met the requirements, and the initial causal inference model needs to be optimized. The model parameters of the initial causal inference model can be adjusted, such as adjusting the weights of nodes and edges in the Bayesian network, or changing parameters such as the learning rate of the initial causal inference model, so that the initial causal inference model can better fit the data. At the same time, optimize the target feature data based on the preset feature optimization rules, which may include removing some features with low correlation, or further combining and transforming the features to improve the performance of the initial causal inference model.
[0074] Step S14, during the operation of the target cloud platform, continuously determine the current alarm data, and perform alarm root cause localization based on the current alarm data and the target causal inference model.
[0075] During the operation of the target cloud platform, new alarm data will be continuously generated. The monitoring system and log recording mechanism of the target cloud platform can be used to collect and organize these alarm data in real time. At the same time, perform preliminary preprocessing on the collected current alarm data, similar to the data cleaning operation in step S11, to remove possible duplicate alarms and data with abnormal formats, so as to ensure the data quality input to the target causal inference model.
[0076] Input the current alarm data into the target causal inference model, so that the target causal inference model can generate each alarm root cause and the probability value corresponding to each alarm root cause based on the current alarm data. After receiving the current alarm data, the target causal inference model will deeply analyze and reason about the input data according to the causal relationship pattern learned in the Bayesian network, and finally output the alarm root cause and the corresponding probability value.
[0077] It can be understood that the generated alarm root causes and corresponding probability values are presented to the operation and maintenance personnel through a visual interface. Each alarm root cause and its probability value are presented on the visual interface, and relevant causal relationship explanations can be provided to help the operation and maintenance personnel understand why this factor is considered a possible root cause. The operation and maintenance personnel can quickly take corresponding measures for fault troubleshooting and resolution based on this information, combined with the actual operation of the cloud platform, so as to improve the stability and reliability of the cloud platform.
[0078] As can be seen from the above, in this application, first, data cleaning is carried out on the alarm data and operation logs generated by the target cloud platform. After completing the data cleaning, target feature data is extracted from the obtained cleaned data. Then, this part of the cleaned data and the extracted target feature data are divided into a training set and a test set. After that, the operation status information of the cloud platform and the content related to alarm events are refined from the divided training set. The operation status of the cloud platform is set as the alarm root cause node, and the alarm event is set as the alarm result node. For those alarm root cause nodes and alarm result nodes with causal associations, connections are established between these nodes to construct a Bayesian network to represent the causal relationship between the nodes. Based on the constructed Bayesian network, an initial causal reasoning model is further built. The performance of this initial causal reasoning model is evaluated using the previously divided test set, and the performance of the model is measured by calculating the F1 value. When the calculated F1 value meets the preset performance conditions, this initial causal reasoning model can be determined as the target causal reasoning model. Finally, during the actual operation of the target cloud platform, newly generated alarm data is continuously and real-time determined. These real-time alarm data are input into the determined target causal reasoning model, and the target causal reasoning model is used to analyze and reason about the current alarm data to achieve accurate positioning of the alarm root cause. In this way, this application can improve the accuracy of alarm root cause positioning on the cloud platform and provide certain support for the efficient operation and maintenance of the cloud platform.
[0079] The technical solution of the embodiment of this application will be specifically described below in combination with the application scenario of a certain institution's cloud platform.
[0080] The cloud platform of a certain institution undertakes the operation of the core business system, including key tasks such as online transaction processing and customer data storage. In daily operations, the cloud platform needs to maintain a high degree of stability at all times. If a large number of alarms suddenly occur in the cloud platform monitoring system, involving problems such as slow database response and inability to access some business functions normally, it will seriously affect the normal development of the business.
[0081] According to the process of the embodiment of this application, first, data cleaning is carried out on the historical alarm data and operation logs of the cloud platform. In the original data, it is found that there are data noises, such as some abnormal values in the memory usage records of some servers that far exceed the reasonable range, and there are also problems such as inconsistent timestamp formats and chaotic log level representations. The Z - score method and the isolation forest algorithm are used to identify and remove the abnormal values, and the mean / median filling method is used to process the missing values. All timestamps are converted to the ISO 8601 standard format, and the log level representation is unified to ensure the high quality of the data.
[0082] Subsequently, key features are extracted from the cleaned data. The dimensionality of the high-dimensional data is reduced through the principal component analysis algorithm. Combining the random forest algorithm and the linear regression algorithm with the L1 regularization term, the alarm time features (such as the occurrence time and duration of the alarm), alarm type features, alarm level features, and alarm location features are extracted.
[0083] Based on the data after the above cleaning and feature extraction, a Bayesian network is constructed. Taking the operating state of the cloud platform as the alarm root cause node and the alarm event as the alarm result node, the information gain of the node is calculated to assign weights to the nodes, and the conditional probability is used to assign weights to the connections between the nodes, thereby constructing an initial causal inference model. The model is verified through the test set, the F1 value is calculated, and after adjusting the model parameters and optimizing the feature data, the target causal inference model with excellent performance is finally determined.
[0084] When the cloud platform is running in real time, once new alarm data is generated, it is immediately input into the target causal inference model. The model quickly analyzes and outputs the possible alarm root causes and corresponding probability values, indicating that the high disk I / O of the database server and the large network latency are the main causes of this alarm. Based on the model output, the operation and maintenance team quickly optimizes the disk of the database server and improves the network speed and stability, successfully solving the cloud platform failure and ensuring the normal operation of the business. This process demonstrates the efficiency and practicality of the embodiments of this application in actual scenarios, greatly improving the operation and maintenance efficiency and stability of the cloud platform.
[0085] Correspondingly, as shown in Figure 2 the embodiments of this application provide a cloud platform alarm root cause positioning device, including:
[0086] A data division module 11, configured to clean the alarm data and operation logs on the target cloud platform, extract target feature data from the obtained cleaned data, and divide the cleaned data and the target feature data into a training set and a test set;
[0087] A network establishment module 12, configured to extract the operating state of the cloud platform and alarm events from the training set, use the operating state of the cloud platform as the alarm root cause node and the alarm event as the alarm result node, and establish a connection between the alarm root cause node and the alarm result node with a causal relationship to obtain a Bayesian network;
[0088] A model determination module 13, configured to construct an initial causal inference model based on the Bayesian network, and use the initial causal inference model and the test set to determine the F1 value. When the F1 value meets the preset performance conditions, the initial causal inference model is determined as the target causal inference model;
[0089] The alarm locating module 14 is used to determine the current alarm data in real time during the operation of the target cloud platform, and locate the root cause of the alarm based on the current alarm data and the target causal reasoning model.
[0090] As can be seen from the above, in this application, first, data cleaning work is carried out for the alarm data and operation logs generated by the target cloud platform. After completing the data cleaning, the target feature data is extracted from the cleaned data. Then, the cleaned part of the data and the extracted target feature data are divided into a training set and a test set. After that, the operation status information of the cloud platform and the alarm event related content are extracted from the divided training set. The operation status of the cloud platform is set as the alarm root cause node, and the alarm event is set as the alarm result node. For those alarm root cause nodes and alarm result nodes with causal associations, a connection is established between these nodes to construct a Bayesian network to represent the causal relationship between the nodes. The initial causal reasoning model is further constructed based on the constructed Bayesian network. The performance of this initial causal reasoning model is evaluated using the test set previously divided, and the performance of the model is measured by calculating the F1 value. When the calculated F1 value meets the preset performance conditions, the initial causal reasoning model can be determined as the target causal reasoning model. Finally, during the actual operation of the target cloud platform, the current newly generated alarm data is continuously and in real time determined. These real-time alarm data are input into the determined target causal reasoning model, and the target causal reasoning model is used to analyze and reason the current alarm data, so as to accurately locate the root cause of the alarm. In this way, this application can improve the accuracy of locating the root cause of alarms on the cloud platform and provide certain support for the efficient operation and maintenance of the cloud platform.
[0091] In some specific implementations, the data partitioning module 11 specifically includes:
[0092] A data extraction unit, used to obtain alarm data and operation logs from the target cloud platform;
[0093] a first outlier determination unit, configured to determine a first outlier from the alarm data and the operation log using a Z-score method, and to determine a second outlier from the alarm data and the operation log based on an isolation forest algorithm;
[0094] A data cleaning unit is used to remove the first abnormal value and the second abnormal value from the alarm data and the operation log to complete the data cleaning operation.
[0095] In some specific implementations, the data partitioning module 11 specifically includes:
[0096] A data dimensionality reduction unit, which is used to perform dimensionality reduction on the data in the cleaned data whose data dimension is greater than a preset dimension threshold by using the principal component analysis algorithm, so as to extract target feature data based on the obtained data after dimensionality reduction;
[0097] A feature extraction unit, which is used to extract alarm time features, alarm type features, alarm level features, and alarm location features from the cleaned data based on the random forest algorithm and the linear regression algorithm with an L1 regularization term; wherein, the alarm time features include the moment feature when the alarm occurs and the duration feature of the alarm;
[0098] A feature determination unit, which is used to determine the alarm time features, the alarm type features, the alarm level features, and the alarm location features as target feature data.
[0099] In some specific embodiments, the data partitioning module 11 specifically includes:
[0100] A data partitioning unit, which is used to partition the cleaned data and the target feature data into a training set and a test set based on the k-fold cross-validation technique.
[0101] In some specific embodiments, the network establishment module 12 specifically includes:
[0102] A node configuration unit, which is used to use the operating state of the cloud platform as the alarm root cause node and the alarm event as the alarm result node, and assign corresponding first weight values to the alarm root cause node and the alarm result node by calculating the information gain of the alarm root cause node and the alarm result node to complete node configuration;
[0103] A node connection unit, which is used to establish a connection between the alarm root cause node and the alarm result node with a causal relationship, determine the conditional probability of the causal relationship, and assign a corresponding second weight value to the connection by using the conditional probability to complete connection configuration;
[0104] A network determination unit, which is used to determine a Bayesian network based on the node configuration and the connection configuration.
[0105] In some specific embodiments, the model determination module 13 specifically includes:
[0106] A model testing unit, which is used to input the test set into the initial causal inference model so that the initial causal inference model outputs the alarm root cause;
[0107] An F1 value determination unit, which is used to count the quantity information of the alarm root cause and determine the F1 value based on the quantity information and the actual alarm root cause quantity information;
[0108] The F1 value judgment unit is configured to judge whether the F1 value is not less than a preset performance threshold and obtain a judgment result;
[0109] The model determination unit is configured to, if the judgment result indicates not less than, determine the initial causal inference model as the target causal inference model;
[0110] The data optimization unit is configured to, if the judgment result indicates less than, adjust the model parameters of the initial causal inference model and optimize the target feature data based on a preset feature optimization rule.
[0111] In some specific embodiments, the alarm location module 14 specifically includes:
[0112] The alarm location unit is configured to input the current alarm data into the target causal inference model, so that the target causal inference model generates each alarm root cause and a probability value corresponding to each alarm root cause based on the current alarm data.
[0113] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 3 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure cannot be considered as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the cloud platform alarm root cause location method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0114] In this embodiment, the power supply 23 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0115] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.
[0116] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the cloud platform alarm root cause location method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0117] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the cloud platform alarm root cause location method disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0118] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for related parts.
[0119] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0120] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0121] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0122] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for locating the root cause of alarms in a cloud platform, characterized in that, Including: Clean the alarm data and operation logs on the target cloud platform, extract target feature data from the obtained cleaned data, and divide the cleaned data and the target feature data into a training set and a test set; Extract the cloud platform operation status and alarm events from the training set, use the cloud platform operation status as the alarm root cause node and the alarm event as the alarm result node, and establish a connection between the alarm root cause node and the alarm result node with a causal relationship to obtain a Bayesian network; Construct an initial causal inference model based on the Bayesian network, and use the initial causal inference model and the test set to determine the F1 value. When the F1 value meets the preset performance conditions, determine the initial causal inference model as the target causal inference model; During the operation of the target cloud platform, determine the current alarm data in real time, and perform alarm root cause positioning based on the current alarm data and the target causal inference model.
2. The method for locating the root cause of cloud platform alarms according to claim 1, wherein The cleaning of the alarm data and operation logs on the target cloud platform includes: Obtain the alarm data and operation logs from the target cloud platform; Use the Z-score method to determine the first outliers from the alarm data and the operation logs, and determine the second outliers from the alarm data and the operation logs based on the isolation forest algorithm; Remove the first outliers and the second outliers from the alarm data and the operation logs to complete the data cleaning operation.
3. The method for locating the root cause of cloud platform alarms according to claim 1, wherein The extraction of the target feature data from the obtained cleaned data includes: Use the principal component analysis algorithm to perform dimensionality reduction on the data in the obtained cleaned data with a data dimension greater than the preset dimension threshold, so as to extract the target feature data based on the obtained dimension-reduced data; Extract the alarm time feature, alarm type feature, alarm level feature, and alarm location feature from the cleaned data based on the random forest algorithm and the linear regression algorithm with the L1 regularization term introduced; wherein, the alarm time feature includes the moment feature when the alarm occurs and the duration feature of the alarm; Determine the alarm time feature, the alarm type feature, the alarm level feature, and the alarm location feature as the target feature data.
4. The method for locating the root cause of cloud platform alarms according to claim 1, characterized in that, The division of the cleaned data and the target feature data into a training set and a test set includes: Divide the cleaned data and the target feature data into a training set and a test set based on the k-fold cross-validation technique.
5. The method for locating the root cause of cloud platform alarms according to any one of claims 1 to 4, characterized in that The use of the cloud platform operation status as the alarm root cause node, the alarm event as the alarm result node, and the establishment of a connection between the alarm root cause node and the alarm result node with a causal relationship to obtain a Bayesian network includes: Use the cloud platform operation status as the alarm root cause node and the alarm event as the alarm result node, and assign corresponding first weight values to the alarm root cause node and the alarm result node by calculating the information gain of the alarm root cause node and the alarm result node to complete the node configuration; Establish a connection line between the alarm root cause node and the alarm result node with a causal relationship, determine the conditional probability of the causal relationship, and assign a corresponding second weight value to the connection line using the conditional probability to complete the connection line configuration; Determine a Bayesian network based on the node configuration and the connection line configuration.
6. The method for locating the root cause of cloud platform alarms according to claim 1, characterized in that The determining the F1 value using the initial causal inference model and the test set, and when the F1 value meets the preset performance condition, determining the initial causal inference model as the target causal inference model includes: Input the test set into the initial causal inference model so that the initial causal inference model outputs the alarm root cause; Count the quantity information of the alarm root cause, and determine the F1 value based on the quantity information and the actual alarm root cause quantity information; Judge whether the F1 value is not less than the preset performance threshold, and obtain a judgment result; If the judgment result indicates not less than, determine the initial causal inference model as the target causal inference model; If the judgment result indicates less than, adjust the model parameters of the initial causal inference model, and optimize the target feature data based on the preset feature optimization rules.
7. The method for locating the root cause of cloud platform alarms according to claim 1, wherein The performing alarm root cause location based on the current alarm data and the target causal inference model includes: Input the current alarm data into the target causal inference model so that the target causal inference model generates each alarm root cause and the probability value corresponding to each alarm root cause based on the current alarm data.
8. An alarm root cause location device for a cloud platform, characterized in that, Including: A data partitioning module, configured to perform data cleaning on the alarm data and operation logs on the target cloud platform, extract target feature data from the obtained cleaned data, and partition the cleaned data and the target feature data into a training set and a test set; A network establishment module, configured to extract the cloud platform operation status and alarm events from the training set, use the cloud platform operation status as the alarm root cause node and the alarm event as the alarm result node, and establish a connection line between the alarm root cause node and the alarm result node with a causal relationship to obtain a Bayesian network; A model determination module, configured to construct an initial causal inference model based on the Bayesian network, and determine the F1 value using the initial causal inference model and the test set. When the F1 value meets the preset performance condition, determine the initial causal inference model as the target causal inference model; An alarm location module, configured to determine the current alarm data in real time during the operation of the target cloud platform, and perform alarm root cause location based on the current alarm data and the target causal inference model.
9. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the cloud platform alarm root cause location method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, it implements the cloud platform alarm root cause location method according to any one of claims 1 to 7.
Citation Information
Cited By
Root cause reasoning method and system based on event driving and large model fine tuning
CN121094102A