Fault detection rule determination method, computing device and storage medium
By acquiring and analyzing the running data characteristics of the data processing object and generating timing and module fault association rules, the accuracy of the data processing object's fault detection is solved, ensuring the continuity and reliability of data processing.
Patent Information
- Application Number
- CN202410021898.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, data processing objects are prone to failure during processing, resulting in interruption of data processing and unable to meet the needs of various application scenarios. How to determine effective fault detection rules has become an urgent problem to be solved.
By determining the operating data of the failed data processing object, fault characteristics, abnormal timing characteristics and abnormal module characteristics are obtained, timing fault association rules and module fault association rules are determined based on these characteristics, and fault detection rules are generated to predict and avoid faults of data processing objects.
It realizes accurate detection of faults of data processing objects, avoids interruptions in data processing, and meets the actual needs of various application scenarios.
Smart Images

Figure CN120276929A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for determining a fault detection rule, a computing device, and a storage medium. Background Art
[0002] With the continuous development of computer technology, data processing objects are used to process data in different actual application scenarios to meet different actual application requirements.
[0003] However, during the process of using a data processing object to process data, the data processing object may malfunction for various reasons, resulting in the interruption of the data processing process and the inability to meet the actual application requirements of various application scenarios. Therefore, how to determine a fault detection rule for a data processing object has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a method for determining a fault detection rule, a computing device, and a storage medium. One or more embodiments of this specification also relate to a device for determining a fault detection rule, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a method for determining a fault detection rule is provided, including:
[0006] Determine the operation data of a malfunctioning data processing object;
[0007] Process the operation data through an operation data processing rule to obtain a fault feature, an abnormal timing feature, and an abnormal module feature, where the abnormal module feature is a feature corresponding to a data processing module in the data processing object;
[0008] Determine a timing fault association rule based on the fault feature and the abnormal timing feature, and determine a module fault association rule based on the fault feature and the abnormal module feature;
[0009] Determine a fault detection rule based on the timing fault association rule and the module fault association rule.
[0010] According to the second aspect of the embodiments of this specification, a device for determining a fault detection rule is provided, including:
[0011] A data determination module, configured to determine the operation data of a malfunctioning data processing object;
[0012] A feature determination module, configured to process the operation data by running data processing rules to obtain fault features, abnormal timing features, and abnormal module features, where the abnormal module features are features corresponding to the data processing modules in the data processing object;
[0013] A first rule determination module, configured to determine a timing fault association rule based on the fault features and the abnormal timing features, and determine a module fault association rule based on the fault features and the abnormal module features;
[0014] A second rule determination module, configured to determine a fault detection rule based on the timing fault association rule and the module fault association rule.
[0015] According to the third aspect of the embodiments of the present specification, a fault detection rule determination system is provided. The system includes a cloud server and a monitoring platform, where
[0016] The cloud server is configured to provide operation data generated during operation to the monitoring platform;
[0017] The monitoring platform is configured to determine the operation data of the faulty cloud server; process the operation data by running data processing rules to obtain fault features, abnormal timing features, and abnormal module features, where the abnormal module features are features corresponding to the data processing modules in the data processing object; determine a timing fault association rule based on the fault features and the abnormal timing features, and determine a module fault association rule based on the fault features and the abnormal module features; determine a fault detection rule based on the timing fault association rule and the module fault association rule.
[0018] According to the fourth aspect of the embodiments of the present specification, a fault detection method is provided, which is applied to a detection end and includes:
[0019] Configure the obtained fault detection rules, where the fault detection rules include a timing fault association rule and a module fault association rule;
[0020] Determine the operation data of the cloud server in the host and determine the abnormal timing features and abnormal module features based on the operation data;
[0021] When it is determined that the abnormal timing features satisfy the timing fault association rule and the abnormal module features satisfy the module fault association rule, it is determined that the cloud server has a fault risk, and hot migration is performed on the cloud server in the host.
[0022] According to the fifth aspect of the embodiments of the present specification, a fault detection device is provided, which is applied to a detection end and includes:
[0023] A rule configuration module, configured to configure the obtained fault detection rules, where the fault detection rules include a timing fault association rule and a module fault association rule;
[0024] A feature determination module, configured to determine the operation data of a cloud server in a host, and determine an abnormal timing feature and an abnormal module feature based on the operation data;
[0025] A fault determination module, configured to determine that there is a fault risk for the cloud server in the case that the determined abnormal timing feature satisfies the timing fault association rule and the abnormal module feature satisfies the module fault association rule, and perform a hot migration on the cloud server in the host.
[0026] According to a sixth aspect of the embodiments of the present specification, a computing device is provided, including:
[0027] A memory and a processor;
[0028] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the above-mentioned fault detection rule determination method or the steps of the fault detection method are implemented.
[0029] According to a seventh aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the above-mentioned fault detection rule determination method or the steps of the fault detection method are implemented.
[0030] According to an eighth aspect of the embodiments of the present specification, a computer program is provided, where when the computer program is executed in a computer, the computer is made to execute the above-mentioned fault detection rule determination method or the steps of the fault detection method.
[0031] In the embodiments of the present specification, a method for determining a fault detection rule is provided, including: determining the operation data of a processed object with a fault; processing the operation data through an operation data processing rule to obtain a fault feature, an abnormal timing feature, and an abnormal module feature, where the abnormal module feature is a feature corresponding to a data processing module in the data processing object; determining a timing fault association rule based on the fault feature and the abnormal timing feature, and determining a module fault association rule based on the fault feature and the abnormal module feature; determining a fault detection rule based on the timing fault association rule and the module fault association rule.
[0032] Specifically, the method for determining a fault detection rule can process the operation data of a fault data processing object through an operation data processing rule, obtain fault features, abnormal timing features, and abnormal module features, determine a timing fault association rule based on the fault features and abnormal timing features, and determine a module fault association rule based on the fault features and abnormal module features. Finally, based on the timing fault association rule and the module fault association rule, a fault detection rule that can accurately detect the fault problems of the data processing object is determined, thereby avoiding the problem of interruption of the data processing process caused by the fault of the data processing object, and enabling the data processing object to meet the actual application requirements of various application scenarios. Description of the Drawings
[0033] Figure 1 is a schematic application diagram of a method for determining a fault detection rule provided by an embodiment of this specification;
[0034] Figure 2 is a flowchart of a method for determining a fault detection rule provided by an embodiment of this specification;
[0035] Figure 3 is a flowchart of the processing process of a method for determining a fault detection rule provided by an embodiment of this specification;
[0036] Figure 4 is a flowchart of the processing process of frequent sequence mining and generating an association relationship in a method for determining a fault detection rule provided by an embodiment of this specification;
[0037] Figure 5 is a flowchart of the processing process of frequent itemset mining and generating an association relationship in a method for determining a fault detection rule provided by an embodiment of this specification;
[0038] Figure 6 is a schematic diagram of a fault detection rule determination system provided by an embodiment of this specification;
[0039] Figure 7 is a flowchart of a fault detection method provided by an embodiment of this specification;
[0040] Figure 8 is a schematic structural diagram of a fault detection rule determination device provided by an embodiment of this specification;
[0041] Figure 9 is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed Embodiments
[0042] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0043] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any and all possible combinations of one or more of the associated listed items.
[0044] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0045] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0046] First, the noun terms involved in one or more embodiments of this specification are explained.
[0047] ECS: Elastic Compute Service (ECS for short) is an IaaS (Infrastructure as a Service) level cloud computing service that provides good performance, is stable and reliable, and has elastic scalability.
[0048] Frequent itemset: An itemset is a collection of one or more items. The elements in the itemset are non-repeating and unordered. A frequent itemset refers to an itemset whose support degree is greater than or equal to the minimum support threshold.
[0049] Frequent sequence: A sequence is an ordered (chronological) set of items, and item repetition is allowed. A frequent sequence refers to a sequence whose support is greater than the minimum support threshold.
[0050] Association rule: Also known as an association relationship, it is used to infer the occurrence of another item set based on a certain item set.
[0051] Downtime prediction: Downtime refers to the phenomenon that an ECS becomes unavailable due to software or hardware problems such as shutdown or crash. Downtime prediction is to predict the downtime behavior of an ECS through a series of analysis methods.
[0052] Live migration: Live migration means migrating a running ECS from one physical machine to another physical machine without stopping its data processing, that is, the ECS is running continuously during the migration process.
[0053] PrefixSpan algorithm: It is an association rule mining algorithm used to mine frequently occurring data. For example, this PrefixSpan algorithm can be used to discover sequence patterns that frequently appear in a sequence dataset.
[0054] FP-growth: It is a frequent item set mining algorithm. Its core idea is to transform the transaction dataset into a data structure called an FP-tree and use this FP-tree for frequent pattern mining.
[0055] A / B test: Also called AB testing, it means providing two different alternative solutions for the downtime prediction rules to be investigated, and then letting a part of users use solution A and another part of users use solution B. Finally, the better solution is determined by observing and comparing the data.
[0056] Dryrun: Dry run; it refers to a simulation test rather than a real test.
[0057] With the continuous development of computer technology, data processing objects are used to process data in different actual application scenarios to meet different actual application requirements. For example, the data processing object is an Elastic Compute Service (ECS) cloud server. When the ECS cloud server provides services, there is a probability of downtime, which will cause the ECS to be unavailable and affect the customer's data processing tasks. For some customers with single-machine application scenarios and without high-availability capabilities, once a downtime occurs, it will cause the unavailability of the customer's related services. For example, for some single-machine game customers, it will cause phenomena such as players being disconnected. For these customers, they are very sensitive to downtime because it will cause the unavailability of their services. In order to reduce the downtime probability of these downtime-sensitive customers. Based on the above background, these downtime-sensitive customers have put forward the need to reduce their downtime probability within the range of performance tolerance. However, a downtime prediction rule provided in this specification is based on some hardware for downtime prediction, such as predicting downtime based on the performance data of memory, but the prediction efficiency is poor and cannot meet the actual needs.
[0058] Based on this, in this specification, a method for determining a fault detection rule is provided. This specification also relates to a device for determining a fault detection rule, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0059] See Figure 1 , Figure 1 shows an application schematic diagram of a method for determining a fault detection rule provided according to an embodiment of this specification. Based on Figure 1 it can be known that when the method for determining a fault detection rule provided in this specification is applied to the ECS cloud server scenario, the server to which the method for determining a fault detection rule is applied can obtain the log data of the downtime ECS cloud server; determine abnormal data from the log data, and perform feature extraction on the abnormal data to extract various types of abnormal features and downtime features. Perform frequent sequence mining based on the abnormal features to obtain frequent sequences, and then generate association relationships based on the frequent sequences and the downtime features. And perform frequent itemset mining based on the abnormal features to obtain frequent itemsets, and then generate association relationships based on the frequent itemsets and the downtime features. Combine and configure the two obtained association relationships to obtain a downtime prediction rule for predicting downtime of the ECS cloud server. Based on this, it can be known that the method for determining a fault detection rule provided in one or more embodiments of this specification is a solution for generating a downtime prediction rule based on frequent itemset and frequent sequence mining. This downtime prediction rule is used for predicting downtime phenomena. Once the method predicts that the ECS will have a downtime behavior, it will perform a live migration on the ECS, migrating the ECS from the current host to another healthier host, thereby avoiding the occurrence of downtime.
[0060] See Figure 2 , Figure 2 which shows a flowchart of a method for determining a fault detection rule provided according to an embodiment of this specification, specifically including the following steps.
[0061] Step 202: Determine the operation data of the failed data processing object.
[0062] Among them, the data processing object can be understood as an object capable of processing data; when the fault detection rule determination method is applied to different scenarios, the data processing object is also different. When the fault detection rule determination method is applied to the virtual machine fault detection scenario, the data processing object can be a virtual machine; when the fault detection rule determination method is applied to the server fault detection scenario, the data processing object can be a server; when the fault detection rule determination method is applied to the cloud server fault detection scenario, the data processing object can be a cloud server; therefore, the data processing object can be a cloud server, a virtual machine, a container, a server (i.e., a physical machine), or a terminal, and this specification does not make specific limitations on this.
[0063] The failed data processing object can be understood as a data processing object that has failed; for example, the failed data processing object can be a failed cloud server that has failed; the failed data processing object can also be a failed server that has failed, a failed virtual machine that has failed, a failed terminal that has failed, or a failed container that has failed, etc. It should be noted that the fault that occurs to the data processing object can be set according to the actual application scenario. For example, the fault can be a downtime, a crash, a system error, a system hardware fault problem, etc. In the fault detection rule determination method provided in this specification, it is illustrated by taking the fault as a downtime as an example. Based on this, the failed data processing object can be a downed data processing object that has experienced a downtime. For example, the failed data processing object can be a downed cloud server that has experienced a downtime; or, the failed data processing object can be a downed server that has experienced a downtime, a downed virtual machine, a downed terminal, a downed container, etc.
[0064] The operation data can be understood as the data generated by the failed data processing object during operation; for example, the operation data can be log data.
[0065] In one or more embodiments provided in this specification, the data processing object is a cloud server;
[0066] Correspondingly, the determining the operation data of the failed data processing object includes:
[0067] Obtain the log data of the cloud server from the log data storage unit, where the log data is obtained through the data collection module corresponding to the cloud server, and the data collection module stores the obtained log file in the log data storage unit;
[0068] Determine the faulty cloud server that has failed, and determine the log data of the faulty cloud server from the log data.
[0069] Among them, the log data storage unit can be understood as the storage unit for storing this log data; this log data storage unit can be a log database for storing log data. The data collection module can be understood as a data collection component (agent), and this data collection component can be a component deployed in the cloud server for collecting the log data of the cloud server.
[0070] Specifically, the method for determining the fault detection rule provided in this specification can obtain the log data of this cloud server obtained by the data collection module from the log data storage unit. Then, determine the faulty cloud server that has failed, and quickly determine the log data of the faulty cloud server from the log data of the cloud server, which is convenient for subsequently determining the fault detection rule based on this log data and improving the rule generation efficiency.
[0071] Taking the method for determining the fault detection rule provided in this specification applied to the scenario of determining the cloud server fault detection rule as an example, the method for determining the fault detection rule provided in this specification is described. Among them, the data processing object is the cloud server (ECS), and the running data is the log data. Based on this, when the ECS is providing services, the data collection component will collect various log data of the ECS in real time and store the log data in the log database through the log backflow method. And the monitoring platform will collect relevant log data from the log database and determine the log data of the downed cloud server that has experienced a downtime event from this log data.
[0072] Step 204: Process the running data through the running data processing rule to obtain the fault feature, the abnormal timing feature, and the abnormal module feature.
[0073] Among them, the abnormal module feature is the feature corresponding to the data processing module in the data processing object.
[0074] Among them, the operation data processing rule can be understood as a rule for processing operation data, and this operation data processing rule can be set according to the actual application scenario. For example, taking the application of the fault detection rule determination method provided in the embodiments of this specification in the scenario of downtime detection as an example, the operation data processing rule may refer to performing frequent sequence mining operations on the operation data to obtain abnormal temporal features and fault features. And / or the operation data processing rule may refer to performing frequent item set mining operations on the operation data to obtain abnormal module features and fault features. For the explanations of the frequent sequence mining operation and the frequent item set mining operation, reference can be made to the corresponding content in this specification, and details will not be elaborated here.
[0075] The fault feature can be understood as a feature representing the fault event that occurs to the data processing object. For example, when the data processing object is a cloud server, the fault feature can be understood as a feature representing the downtime event that occurs to the cloud server. When the data processing object is a virtual machine, the fault feature can be understood as a feature representing the downtime event that occurs to the virtual machine.
[0076] The abnormal temporal feature can be understood as a feature composed of abnormal features that can be sorted by time; when the data processing object is a cloud server, the abnormal temporal feature can be understood as a sequence of abnormal event features that occur before the cloud server downtime. For example, the abnormal temporal feature can be a sequence composed of a series of abnormal event features such as full CPU utilization and downtime, and this sequence is sorted by time.
[0077] The data processing module can be understood as the module in the data processing object for performing data processing; when the data processing object is a cloud server, the data processing module can be hardware modules such as the motherboard, CPU, and memory; or, the data processing module can be software modules such as the system kernel. The abnormal module feature can be understood as a feature representing the abnormal event that occurs to the data processing module in the data processing object. For example, the abnormal module feature can be a feature representing abnormal events such as serious motherboard failure, CPU failure, memory failure, or kernel exception.
[0078] In one or more embodiments provided in this specification, in order to obtain a fault detection rule capable of accurately performing fault detection, the fault detection rule determination method provided in this specification can mine abnormal timing features and abnormal module features from operation data, so as to more comprehensively mine various abnormal events that have occurred to the faulty data processing object, facilitating the subsequent generation of fault detection rules to comprehensively consider various abnormal events, and thus generating a fault detection rule capable of accurately performing fault detection. Specifically, the processing of the operation data through the operation data processing rule to obtain fault features, abnormal timing features, and abnormal module features includes steps one to three:
[0079] Step 1: Determine the abnormal data features corresponding to the faulty data processing object based on the operation data.
[0080] Among them, the abnormal data features can be understood as the features corresponding to the abnormal operation data in the operation data.
[0081] Specifically, the determining of the abnormal data features corresponding to the faulty data processing object based on the operation data includes:
[0082] Determine the abnormal operation data of the faulty data processing object from the operation data through the abnormal data screening rule;
[0083] Extract features from the abnormal operation data to determine the abnormal data features corresponding to the faulty data processing object.
[0084] Among them, the abnormal data screening rule can be understood as a rule for screening out abnormal operation data from the operation data; the abnormal data screening rule can be set according to the actual application scenario. For example, the abnormal data screening rule can be: According to the Io performance data, screen out the event data with a high delay (delay rate greater than 80%) from the log data as abnormal data. Abnormal operation data can be understood as operation data with abnormalities; for example, the abnormal data can be: event data such as CPU failure, motherboard failure, network device failure, GPU card dropout, CPU utilization, network delay, etc. that may cause the ECS to crash.
[0085] Continuing with the above example, after the monitoring platform collects the log data of the crashed cloud server, it configures some anomaly detection rules (i.e., anomaly data screening rules) for the log data, and determines the anomaly data (i.e., abnormal operation data) from the log data based on the anomaly detection rules; feature extraction is performed based on the anomaly data, and various types of anomaly data features are extracted. For example, the anomaly data feature can be a hardware feature indicating the hardware status, a network feature related to the network, a crash feature describing the availability of the instance, and so on. This facilitates the subsequent rapid mining of abnormal time series features and abnormal module features from the operation data.
[0086] In the embodiments provided in this specification, the anomaly data feature can be an abstraction of the anomaly data. For example, CPU failure anomaly events and motherboard failure anomaly events are uniformly abstracted as hardware failure anomaly features.
[0087] Step two: By performing time series feature screening processing and time series feature mining processing on the anomaly data features, the fault feature and the abnormal time series feature are obtained.
[0088] The fault feature can be understood as a feature characterizing the fault event that occurs to the data processing object. For example, when the data processing object is a cloud server, the fault feature can be understood as a feature characterizing the crash event that occurs to the cloud server. When the data processing object is a virtual machine, the fault feature can be understood as a feature characterizing the crash event that occurs to the virtual machine. It should be noted that the method for determining the fault detection rules provided in this specification determines the abnormal time series feature and the abnormal module feature through two steps respectively. And the fault feature needs to be determined in both steps, so as to be used for the abnormal time series feature or the abnormal module feature. Therefore, the fault feature needs to be determined in both step two and step three.
[0089] In one or more embodiments provided in this specification, in order to improve the mining efficiency of abnormal time series features, the anomaly data features are initially selected to obtain candidate time series features; then the abnormal time series features are determined from the candidate time series features, so as to improve the acquisition efficiency of the abnormal time series features and further improve the efficiency of determining the fault detection rules. Specifically, the obtaining of the fault feature and the abnormal time series feature by performing time series feature screening processing and time series feature mining on the anomaly data features includes:
[0090] Based on the anomaly data features, the fault feature is determined, and time series feature screening processing is performed on the anomaly data features to obtain candidate time series features;
[0091] The candidate time series features are input into a time series feature mining module for processing to determine the abnormal time series features in the candidate time series features.
[0092] Among them, the timing feature screening process can be understood as a process of screening features for the abnormal data features to determine candidate timing features from the abnormal data features. The timing feature mining process can be understood as a process of mining abnormal timing features from the candidate timing features to determine abnormal timing features from the candidate timing features.
[0093] The candidate timing features can be understood as features used to determine abnormal timing features and composed of abnormal features that can be sorted by time. For example, the abnormal timing features can be a sequence composed of a series of abnormal event features such as full CPU utilization and system crashes, and this sequence is used to determine abnormal timing features.
[0094] The timing feature mining module can be understood as a module capable of mining abnormal timing features from candidate timing features. For example, the sequence processing module can be a timing feature mining model. By inputting the candidate timing features into the timing feature mining model, the abnormal timing features in the candidate timing features can be mined.
[0095] Specifically, in one or more embodiments provided in this specification, the timing feature screening process for the abnormal data features to obtain candidate timing features includes:
[0096] Based on the abnormal data features, determine the failure time corresponding to the failed data processing object;
[0097] Based on the failure time and a preset timing feature threshold, determine a feature time interval;
[0098] From the abnormal data features, obtain the candidate timing features corresponding to the feature time interval.
[0099] Among them, the failure time can be understood as the moment when the failed data processing object fails. The preset timing feature threshold can be understood as a pre-set time range threshold required to obtain candidate timing features. For example, the preset timing feature threshold can be 1 day, 2 days, etc. That is to say, in the process of obtaining candidate timing features, a time range threshold needs to be determined, and based on this time range threshold, the time range for obtaining candidate timing features is controlled.
[0100] The feature time interval can be understood as the time interval for obtaining the candidate timing features. It should be noted that the feature time interval needs to be determined by the failure time (such as the system crash moment) and the preset timing range threshold (such as 24 hours). Based on this, it is possible to determine which period of time in the historical time is used as the time interval where the candidate timing features are located.
[0101] Continuing with the above example, after identifying the downed ECS and determining the abnormal characteristics (i.e., abnormal data characteristics) of the downed ECS; based on the abnormal data, determine the downtime of the cloud server, and determine that the preset characteristic acquisition time interval is 2 days. Based on this, starting from the moment when the cloud server went down, a time window of 2 days (i.e., the characteristic acquisition time interval) is traced back for the abnormal characteristics forward, obtaining the time series within this time window (i.e., the candidate time series characteristics). Thus, the acquisition efficiency of the abnormal time series characteristics is improved based on the candidate time series characteristics.
[0102] Specifically, in one or more embodiments provided in this specification, the step of inputting the candidate time series characteristics into the time series characteristic mining module for processing to determine the abnormal time series characteristics in the candidate time series characteristics includes:
[0103] Input the candidate time series characteristics into the frequent feature sequence mining model, and perform frequent feature sequence mining on the candidate time series characteristics through the frequent feature sequence mining model to obtain the frequent feature sequences in the candidate time series characteristics, where the frequent feature sequences include frequent features, and the frequent features are features that appear more than a preset number threshold in the frequent feature sequences;
[0104] Use the frequent feature sequences as the abnormal time series characteristics.
[0105] Among them, the frequent feature sequence mining model can be understood as a model capable of mining frequent feature sequences. For example, the frequent feature sequence mining model can be the PrefixSpan algorithm.
[0106] The frequent feature sequences can be understood as being composed of frequent features, and the frequent feature sequences can be understood as the above-mentioned frequent sequences; the frequent features can be understood as features that appear more than a preset number threshold in the sequence, and the preset number threshold can be understood as the support threshold. When the support of a candidate time series is greater than the minimum support threshold, it can be determined as a frequent sequence. That is to say, when the number of occurrences of features in a candidate time series is greater than the minimum support threshold, it can be determined as a frequent sequence. In one or more embodiments provided in this specification, the number of occurrences of the features can be represented by the number of features. Based on this, when the number of features in a candidate time series is greater than the minimum support threshold, it can also be determined as a frequent sequence. For example, the frequent feature sequence can be a sequence composed of abnormal events that occur more than a preset number threshold (support) within 2 days before the cloud server ECS goes down. For example, the frequent sequence can be a series of abnormal events such as the CPU being fully loaded and the server going down that occur relatively frequently.
[0107] Continuing with the above example, the frequent feature sequence mining model is the PrefixSpan algorithm. Based on this, after determining the time series, the PrefixSpan algorithm is used to mine the time series to obtain frequent sequences. Specifically, during the frequent sequence mining process, first, a support threshold (i.e., a preset number threshold) is set. Second, the time series is input into the PrefixSpan algorithm model, and all frequent sequences are discovered through the PrefixSpan sequence pattern mining algorithm, obtaining frequent sequences that retain the downtime features (abnormal time series features), and this frequent sequence can be called a sequence pattern. Thus, accurate fault detection rules are generated based on this abnormal time series feature.
[0108] It should be noted that frequent sequence mining aims to find all time series containing downtime features with the downtime feature as the target, and find all frequent sequences by setting a support threshold, thereby discovering the correlation between frequent sequences and downtime events. If the support threshold is set too large, some useful features may be missed; if the support threshold is set too small, the discovered frequent sequences are too large and it is not easy to calculate relevant metrics and analyze. Therefore, the method for determining the fault detection rules provided in this specification sets a reasonable support threshold in combination with the actual application situation, so that useful features are not missed, and the discovered frequent sequences are not too large and not easy to calculate relevant metrics and analyze.
[0109] Step 3: By performing module feature screening processing and module feature mining processing on the abnormal data features, the fault features and the abnormal module features are obtained.
[0110] Among them, the fault feature can be understood as the feature representing the fault event that occurs to the data processing object. For example, when the data processing object is a cloud server, the fault feature can be understood as the feature representing the downtime event that occurs to the cloud server. When the data processing object is a virtual machine, the fault feature can be understood as the feature representing the downtime event that occurs to the virtual machine. It should be noted that in the embodiments provided in this specification, the fault feature and the fault feature can be the same.
[0111] In one or more embodiments provided in this specification, during the ECS operation and maintenance process, it is often found that some hardware features are highly correlated with the downtime of the cloud server. Therefore, in order to improve the accuracy of the fault detection rules, it is necessary to determine the hardware features. The specific content is that the obtaining of the fault features and the abnormal module features by performing module feature screening processing and module feature mining processing on the abnormal data features includes:
[0112] Determine the fault features based on the abnormal data features, and perform module feature screening processing on the abnormal data features to obtain candidate module features;
[0113] Input the candidate module features into a module feature mining module for processing to determine the abnormal module features among the candidate module features.
[0114] Among them, module feature screening processing can be understood as a processing step of screening features for the abnormal data features to determine candidate module features from the abnormal data features. Module feature mining processing can be understood as a processing step of mining the abnormal module features among the candidate module features to determine abnormal module features from the candidate time series features.
[0115] The candidate module features can refer to all features associated with data processing modules in abnormal operation data. For example, the candidate module features can be features corresponding to event data such as severe motherboard failures, CPU failures, memory failures, or kernel exceptions.
[0116] The module feature mining module can be understood as a module capable of mining abnormal module features from candidate module features. For example, the module feature mining module can be a module feature mining model. By inputting the candidate module features into the module feature mining model, the abnormal module features among the candidate module features can be mined.
[0117] In one or more embodiments provided in this specification, the step of inputting the candidate module features into a module feature mining module for processing to determine the abnormal module features among the candidate module features includes:
[0118] Perform deduplication processing on the candidate module features to obtain deduplicated candidate module features;
[0119] Input the deduplicated candidate module features into a frequent feature item set mining model for processing, and perform frequent feature item set mining on the deduplicated candidate module features through the frequent feature item set mining model to obtain a frequent feature item set, where the initial frequent feature item set includes frequent module features, and the frequent module features are module features whose occurrence times exceed a preset number threshold;
[0120] Use the frequent feature item set as the abnormal module features.
[0121] Among them, the frequent feature item set mining model can be understood as a model capable of performing frequent feature item set mining. For example, the frequent feature item set mining model can be the FP-Growth algorithm.
[0122] The frequent feature item set can be understood as a set composed of frequent module features. For example, the frequent feature item set can be the above-mentioned frequent item set; the frequent module feature is an abnormal module feature whose occurrence times exceed a preset number threshold; the preset number threshold can be understood as the support threshold. When the support of an item set (i.e., candidate module feature) is greater than the minimum support threshold, it can be determined as a frequent item set. That is to say, when the number of features in an item set is greater than the minimum support threshold, it can be determined as a frequent item set. The frequent item set can include features of abnormal events such as serious motherboard failures, CPU failures, memory failures, or kernel anomalies that occur frequently.
[0123] Continuing with the above example, the frequent feature item set mining model is the FP-Growth algorithm. Based on this, considering that there is a strong correlation between hardware features and cloud server downtime. Therefore, frequent item set mining is carried out within the selected feature range; First, the selected feature set (i.e., candidate module feature) is the union of hardware features and downtime features. Since the frequent item set targets hardware features, only data containing hardware features is extracted. Second, an item set database is generated based on the selected feature set. This item set database can be understood as a feature item set that contains the features in the feature set; Finally, since duplicate elements are not allowed in the frequent item set mining algorithm, that is, the frequent item set mining algorithm requires non-repetitive items, data deduplication operations need to be performed; After deduplicating the duplicate elements (abnormal hardware features) in the item set database and extracting the transaction data of the features, the FP-Growth algorithm is used to mine the frequent item set in the item set database to obtain the frequent item set (i.e., frequent feature item set). This facilitates subsequent improvement of the accuracy of the fault detection rules based on the mined frequent feature item set.
[0124] Step 206: Determine the timing fault association rule based on the fault feature and the abnormal timing feature, and determine the module fault association rule based on the fault feature and the abnormal module feature.
[0125] Among them, the timing fault association rule can be understood as the association rule between the fault feature and the abnormal timing feature, that is, the association relationship between the fault feature and the abnormal timing feature. For example, in the embodiments provided in this specification, after determining the fault feature and the abnormal timing feature, a timing fault association rule needs to be constructed. This timing fault association rule can be used to indicate that the occurrence of the abnormal timing feature can be inferred through this abnormal feature, and the occurrence or discovery of the abnormal feature can be inferred through the abnormal timing feature. That is to say, when the occurrence of the abnormal timing feature can be inferred through the abnormal feature, and the occurrence of the abnormal feature can be inferred through the abnormal timing feature, then the timing fault association rule between the two can be determined, which can also be called the timing fault association relationship. Therefore, this timing fault association rule can be understood as the association rule in which the fault feature and the abnormal timing feature can infer each other's occurrence. Among them, the occurrence of this abnormal timing feature can be understood as the appearance of this abnormal timing feature. The occurrence of this fault feature can be understood as the appearance of this fault feature.
[0126] The module fault association rule can be understood as the association rule between the fault feature and the abnormal module feature, that is, the association relationship between the fault feature and the abnormal module feature. For example, in the embodiments provided in this specification, after determining the fault feature and the abnormal module feature, a module fault association rule needs to be constructed. This module fault association rule can be used to indicate that the occurrence of the abnormal module feature can be inferred through this abnormal feature, or the occurrence of the abnormal feature can be inferred through the abnormal module feature. That is to say, when the occurrence of the abnormal module feature can be inferred through the abnormal feature, and the occurrence of the abnormal feature can be inferred through the abnormal module feature, then the timing fault association rule between the two can be determined, which can also be called the timing fault association relationship. Therefore, this module fault association rule can be understood as the association rule in which the fault feature and the abnormal module feature can infer each other's occurrence. Among them, the occurrence of this abnormal module feature can be understood as the appearance of this abnormal module feature. The occurrence of this fault feature can be understood as the appearance of this fault feature.
[0127] In one or more embodiments provided in this specification, taking the failed data processing object as the downed cloud server as an example, this fault detection can be understood as downtime prediction. This downtime prediction is to discover the abnormal features that are more associated with the downtime feature, that is, to discover the association relationship between these abnormal features and the downtime feature, so as to determine the abnormal events that may cause the cloud server to go down through this association relationship. Therefore, after determining the fault feature and the abnormal timing feature, it is necessary to determine the association relationship between the two, so as to facilitate the subsequent generation of fault detection rules for detecting the downtime of the cloud server. The specific content is that determining the timing fault association rule based on the fault feature and the abnormal timing feature includes:
[0128] Determine the timing correlation data between the fault feature and the abnormal timing feature;
[0129] When it is determined that the timing correlation data satisfies the preset correlation condition, determine the timing fault correlation rule between the fault feature and the abnormal timing feature.
[0130] Among them, the timing correlation data can be understood as a correlation parameter characterizing the degree of correlation between the fault feature and the abnormal timing feature; the preset correlation condition can be set according to the actual application scenario. When the correlation data is a correlation parameter, the preset correlation condition can be to determine whether the correlation parameter is greater than or equal to the preset correlation parameter threshold.
[0131] Specifically, after the fault feature and the abnormal timing feature, it is necessary to calculate the timing correlation data between the fault feature and the abnormal timing feature, such as timing confidence, timing lift, and timing imbalance rate; when it is determined that the timing confidence, timing lift, and timing imbalance rate satisfy the preset correlation parameter threshold, determine the module fault correlation rule between the fault feature and the abnormal timing feature.
[0132] Continuing with the above example, after determining the downtime feature and the frequent sequence, it is necessary to calculate the correlation parameter between the downtime feature and the frequent sequence; when it is determined that the correlation parameter satisfies (i.e., is greater than or equal to) the preset correlation parameter threshold, establish the correlation relationship between the downtime feature and the frequent sequence.
[0133] In one or more embodiments provided in this specification, the determining the timing correlation data between the fault feature and the abnormal timing feature includes:
[0134] Based on the fault feature and the abnormal timing feature, construct a candidate timing fault correlation rule;
[0135] Based on the fault feature and the abnormal timing feature, perform confidence calculation to obtain timing confidence;
[0136] Based on the abnormal data feature, the abnormal timing feature, and the fault feature, perform lift calculation to obtain timing lift;
[0137] Determine the fault feature support of the fault feature and the timing feature support of the abnormal timing feature, and perform imbalance rate calculation based on the fault feature support and the timing feature support to obtain timing imbalance rate;
[0138] Take the timing confidence, the timing lift, and the timing imbalance rate as the timing correlation data of the candidate timing fault correlation rule.
[0139] Among them, the timing confidence can be understood as the confidence of this fault feature and the abnormal timing feature. In one or more embodiments provided in this specification, the timing confidence can be used to measure the accuracy rate of the candidate timing fault association rule. Based on this, the timing confidence can be used to screen the candidate timing fault association rule. By comparing the timing confidence with a preset association parameter threshold, a timing fault association rule with better performance can be obtained. That is to say, the timing confidence is used to determine the confidence of the association rule during the data mining process.
[0140] The timing lift can be understood as the lift of this fault feature and the abnormal timing feature. In one or more embodiments provided in this specification, the timing lift can be used to reflect the correlation between the fault feature and the abnormal timing feature in the candidate timing fault association rule. Based on this, the timing lift can be used to screen the candidate timing fault association rule. By comparing the timing lift with a preset association parameter threshold, a timing fault association rule with better performance can be obtained. That is to say, the timing lift is used to determine the lift of the association rule during the data mining process.
[0141] The fault feature support can be understood as the support of the fault feature. In one or more embodiments provided in this specification, the fault feature support can be understood as the proportion of this fault feature in the abnormal data features. The timing feature support can be understood as the support of the abnormal timing feature. In one or more embodiments provided in this specification, the timing feature support can be understood as the proportion of the abnormal timing feature in the abnormal data features.
[0142] The timing imbalance rate can be understood as the confidence imbalance rate of this fault feature and the abnormal timing feature. In one or more embodiments provided in this specification, the timing imbalance rate is used to screen the candidate timing fault association rule. By comparing the timing imbalance rate with a preset association parameter threshold, a timing fault association rule with better performance can be obtained. That is to say, the timing imbalance rate is used to determine the imbalance rate of the association rule during the data mining process.
[0143] Among them, the candidate timing fault association rule can be understood as the candidate timing fault association rule between the fault feature and the abnormal timing feature, that is, the candidate timing fault association relationship between the fault feature and the abnormal timing feature; subsequently, it is necessary to determine whether the timing confidence, timing lift, and timing imbalance rate of the candidate timing fault association rule meet the preset association parameter threshold, so as to screen the candidate timing fault association rule. When it is determined that the timing confidence, timing lift, and timing imbalance rate meet the preset association parameter threshold, the candidate timing fault association rule is used as the timing fault association rule, so as to obtain the timing fault association rule.
[0144] Continuing with the above example, frequent sequence pattern mining is to find sequence patterns containing downtime features. Therefore, among the time series feature data of ECS, select the sequence feature data containing downtime features, and after mining the frequent sequences (also known as sequence patterns) containing downtime features from it, at this time, generate an association relationship based on the obtained sequence patterns and downtime features. The generated association relationship contains A and B, and A => B. Among them, A is the feature set containing downtime features, and B is the downtime feature.
[0145] After determining the above feature set containing downtime features and downtime features, it is necessary to calculate indicators such as the confidence, lift, and imbalance rate of the association relationship through the feature set containing downtime features and downtime features. That is to say, after finding the frequent sequences, it is necessary to calculate the indicators of the association relationship between the frequent sequences and the downtime features, specifically including the calculation of indicators such as confidence, lift, and imbalance rate. Among them, the calculation formula for confidence P(B|A) can be seen in the following formula (1).
[0146]
[0147] Among them, the count function in the formula is to count the number of ECS containing the target feature set. count(A∪B) represents the number of ECS in which both the A feature set and the B feature set appear.
[0148] The calculation formula for lift Lift(B|A) can be seen in the following formula (2).
[0149]
[0150] Among them, the count(total) function is to count the number of all ECS. That is to say, to calculate the lift, it is necessary to traverse the abnormal transaction data of all ECS for calculation.
[0151] The calculation formula for the imbalance rate IR(A,B) can be seen in the following formula (3).
[0152]
[0153] Among them, the sup function is the calculation function of support. That is to say, in the process of calculating the imbalance rate, it is necessary to calculate the support of the feature set containing downtime features and the support of the downtime feature.
[0154] Based on the above, it can be seen that the method for determining a fault detection rule provided in this specification can calculate the confidence level of the association relationship through a feature set containing downtime features and downtime features; through the abnormal data features, the feature set containing downtime features, and the downtime features, the lift of the association relationship can be calculated; by calculating the support of the feature set containing downtime features and the support of the downtime features, and through this support, the imbalance rate of the association relationship can be calculated. Based on this, in the embodiments of this specification, by calculating the temporal confidence level, temporal lift, and temporal imbalance rate between the fault features and the abnormal temporal features, and using them as the association data of the candidate temporal fault association rules, it is convenient to accurately determine the temporal fault association rules based on this association data subsequently.
[0155] In one or more embodiments provided in this specification, when it is determined that the temporal association data satisfies the preset association condition, determining the temporal fault association rule between the fault feature and the abnormal temporal feature includes:
[0156] When it is determined that the temporal confidence level is greater than or equal to the preset association parameter threshold, the temporal lift is greater than or equal to the preset association parameter threshold, and the temporal imbalance rate is less than or equal to the preset association parameter threshold, the candidate temporal fault association rule is used as the temporal fault association rule between the fault feature and the abnormal temporal feature.
[0157] Among them, the preset association parameter threshold can be set according to the actual application scenario, and this specification does not make specific restrictions on this. For example, when the temporal confidence level, temporal lift, and temporal imbalance rate are all any values in the interval [0, 1], the preset association parameter threshold can be 0.1.
[0158] Continuing with the above example, after calculating indicators such as the confidence level, lift, and imbalance rate of the association relationship, the association relationships that meet the indicator threshold (i.e., the above-mentioned preset association parameter threshold) are filtered out through the set indicator threshold. That is to say, after determining the confidence level, lift, and imbalance rate, the indicator threshold defined for indicators such as the confidence level, lift, and imbalance rate is determined, and then the confidence level, lift, and imbalance rate are compared with this indicator threshold. When the confidence level and lift are greater than or equal to this indicator threshold, and the imbalance rate is less than or equal to this indicator threshold, all association relationships that meet the indicator threshold are extracted. It is convenient to accurately determine the fault detection rule based on this association relationship subsequently.
[0159] In one or more embodiments provided in this specification, taking a failed data processing object as a downed cloud server as an example, this fault detection can be understood as downtime prediction. This downtime prediction is to discover abnormal features associated with the downtime features, that is, to discover the association relationship between these abnormal features and the downtime features, so as to determine abnormal events that may cause the cloud server to go down through this association relationship. Therefore, after determining the fault features and abnormal module features, it is necessary to determine the association relationship between the two, so as to facilitate the subsequent generation of fault detection rules for detecting the downtime of the cloud server. Specifically, the determining the module fault association rule based on the fault feature and the abnormal module feature includes:
[0160] Determine the module association data between the fault feature and the abnormal module feature;
[0161] When it is determined that the module association data meets the preset association condition, determine the module fault association rule between the fault feature and the abnormal module feature.
[0162] Among them, this module association data can be understood as an association parameter representing the degree of association between the fault feature and the abnormal module feature; this preset association condition can be set according to the actual application scenario. When the module association data can be module confidence, module lift, and module imbalance rate, this preset association condition can be to judge whether the module confidence, module lift, and module imbalance rate meet the preset association parameter threshold.
[0163] Specifically, after the fault feature and the abnormal module feature, it is necessary to calculate the module association data between the fault feature and the abnormal module feature, such as module confidence, module lift, and module imbalance rate; when it is determined that the module confidence, module lift, and module imbalance rate meet the preset association parameter threshold, determine the module fault association rule between the fault feature and the abnormal module feature.
[0164] Continuing with the above example, after determining the downtime feature and the frequent item set, it is necessary to calculate the module confidence, module lift, and module imbalance rate between the downtime feature and the frequent item set; when it is determined that the module confidence, module lift, and module imbalance rate meet the preset association parameter threshold, establish the association relationship between the downtime feature and the frequent item set.
[0165] In one or more embodiments provided in this specification, the determining the association data between the fault feature and the abnormal module feature includes:
[0166] Based on the fault feature and the abnormal module feature, construct a candidate module fault association rule;
[0167] Calculate the confidence based on the fault feature and the abnormal module feature to obtain the module confidence;
[0168] Calculate the lift based on the abnormal data feature, the abnormal module feature, and the fault feature to obtain the module lift;
[0169] Determine the fault feature support degree of the fault feature and the module feature support degree of the abnormal module feature, and calculate the imbalance rate based on the fault feature support degree and the module feature support degree to obtain the module imbalance rate;
[0170] Use the module confidence, the module lift, and the module imbalance rate as the module association data of the candidate module fault association rule.
[0171] Among them, the module confidence can be understood as the confidence between the fault feature and the abnormal module feature. In one or more embodiments provided in this specification, the module confidence can be used to measure the accuracy of the candidate module fault association rule. Based on this, the module confidence can be used to screen the candidate module fault association rule by comparing the module confidence with a preset association parameter threshold, so as to obtain a module fault association rule with better performance. That is to say, the module confidence is used to determine the confidence of the association rule during the data mining process.
[0172] The module lift can be understood as the lift between the fault feature and the abnormal module feature. In one or more embodiments provided in this specification, the module lift can be used to reflect the correlation between the fault feature and the abnormal module feature in the candidate module fault association rule. Based on this, the module lift can be used to screen the candidate module fault association rule by comparing the module lift with a preset association parameter threshold, so as to obtain a module fault association rule with better performance. That is to say, the module lift is used to determine the lift of the association rule during the data mining process.
[0173] The fault feature support degree can be understood as the support degree of the fault feature. In one or more embodiments provided in this specification, the fault feature support degree can be understood as the proportion of the fault feature in the abnormal data feature. The module feature support degree can be understood as the support degree of the abnormal module feature. In one or more embodiments provided in this specification, the module feature support degree can be understood as the proportion of the abnormal module feature in the abnormal data feature.
[0174] The module imbalance rate can be understood as the confidence imbalance rate between this fault feature and the abnormal module feature. In one or more embodiments provided in this specification, this module imbalance rate can be used to screen the candidate module fault association rules. By comparing this module imbalance rate with a preset association parameter threshold, a module fault association rule with better performance can be obtained. That is to say, this module imbalance rate is the imbalance rate used to determine the association rule during the data mining process.
[0175] Among them, the candidate module fault association rule can be understood as the candidate module fault association rule between the fault feature and the abnormal module feature. That is, the candidate module fault association relationship between the fault feature and the abnormal module feature; subsequently, it is necessary to determine whether the module confidence, module lift, and module imbalance rate of this candidate module fault association rule meet the preset association parameter threshold, so as to screen this candidate module fault association rule. When it is determined that the module confidence, module lift, and module imbalance rate meet the preset association parameter threshold, this candidate module fault association rule is used as the module fault association rule, thereby obtaining the module fault association rule.
[0176] Continuing with the above example, after obtaining the frequent item set (i.e., the hardware feature item set) through frequent item set mining, based on this frequent item set, it can be determined that A is in the association relationships to be generated; the association relationships to be generated include A and B, and A => B. Among them, A is the hardware feature item set, and B is the downtime feature (fault feature).
[0177] After determining the hardware feature item set and the downtime feature, it is necessary to calculate indicators such as the confidence, lift, and imbalance rate of the association relationship through the hardware feature item set and the downtime feature. That is, after finding the hardware feature item set, it is necessary to calculate the association relationship between the hardware feature item set and the downtime feature, specifically including the calculation of indicators such as confidence, lift, and imbalance rate. The calculation methods for this confidence, lift, and imbalance rate are the same as those for the confidence, lift, and imbalance rate between the above-mentioned frequent sequence and the downtime feature; therefore, for the calculation methods of indicators such as the confidence, lift, and imbalance rate between the hardware feature item set and the downtime feature, reference can be made to the corresponding content of formulas (1), (2), and (3) above, and this specification will not elaborate on it.
[0178] Based on this, it can be known that the fault detection rule determination method provided in this specification can calculate the confidence of the association relationship through the hardware feature item set and the downtime feature; the lift of the association relationship can be calculated through the abnormal data feature, the hardware feature item set, and the downtime feature; by calculating the support of the hardware feature item set and the support of the downtime feature, and through this support, the imbalance rate of the association relationship can be calculated. Based on this, in the embodiments of this specification, by calculating the module confidence, module lift, and module imbalance rate between the fault feature and the abnormal module feature, and using them as the association data between the fault feature and the abnormal module feature, it is convenient to accurately establish the module fault association rule based on this association data subsequently.
[0179] In one or more embodiments provided in this specification, when it is determined that the module association data meets the preset association condition, determining the module fault association rule between the fault feature and the abnormal module feature includes:
[0180] When it is determined that the module confidence is greater than or equal to the preset association parameter threshold, the module lift is greater than or equal to the preset association parameter threshold, and the module imbalance rate is less than or equal to the preset association parameter threshold, the candidate module fault association rule is used as the module fault association rule between the fault feature and the abnormal module feature.
[0181] Following the above example, after calculating indicators such as the confidence, lift, and imbalance rate of the association relationship between the hardware feature item set and the downtime feature, the association relationships that meet the indicator thresholds are filtered out through the set indicator thresholds. That is to say, first, after determining the association relationship based on the mined frequent item set and the downtime feature, calculate indicators such as the confidence and lift of this feature association relationship; secondly, determine the indicator thresholds (i.e., the preset association parameter thresholds) defined for indicators such as confidence, lift, and imbalance rate, and then compare the confidence, lift, and imbalance rate with this indicator threshold. When the confidence and lift are greater than or equal to this indicator threshold, and the imbalance rate is less than or equal to this indicator threshold, it can be determined that the association relationship between the hardware feature item set and the downtime feature meets the requirements, so as to retain all the association relationships with the downtime feature as the right-side feature, achieving the purpose of extracting all the association relationships that meet the preset association parameter thresholds. It is convenient to accurately determine the fault detection rule based on this association relationship subsequently.
[0182] In one or more embodiments provided in this specification, when generating the association relationship, a minimum confidence threshold (i.e., the preset association parameter threshold) needs to be satisfied. Two points need to be considered in the setting of this threshold. First, the method for determining the fault detection rule provided in this specification generates a rule for predicting downtime. When the rule is hit, hot migration of the ECS needs to be performed to ensure that the availability of the ECS is not affected. However, there is a certain probability that the performance of the ECS will be damaged during hot migration, such as a probability of 0.2%. Therefore, to ensure that the confidence level of downtime prediction is at least greater than the probability of performance degradation during hot migration, the downtime prediction rule can be considered valuable. Secondly, if the confidence level is only slightly greater than the performance degradation rate during hot migration, it is not enough. If the number of hot migrations is too large, it will also cause a greater degree of interference to the customer's ECS and lead to performance degradation. Therefore, the method for determining the fault detection rule provided in this specification sets a confidence level of more than 10% to generate relevant operation and maintenance rules. On the one hand, it reduces the probability of downtime of the customer as much as possible, and on the other hand, it reduces the number of hot migrations of the customer's ECS as much as possible. That is to say, in terms of threshold setting, the method for determining the fault detection rule provided in this specification requires that the threshold of the confidence level should be greater than the probability of performance degradation during hot migration and should not cause an increase in the degree of interference to the customer due to frequent operation and maintenance.
[0183] Step 208: Determine a fault detection rule based on the time-series fault association rule and the module fault association rule.
[0184] Among them, this fault detection rule can be understood as a rule that can pre-detect possible faults of the data processing object. For example, this fault detection rule can be understood as a downtime detection rule that can pre-detect possible downtime of the data processing object. That is to say, this fault detection rule can be understood as a fault prediction rule.
[0185] In one embodiment provided in this specification, the determining a fault detection rule based on the time-series fault association rule and the module fault association rule includes:
[0186] Merge and configure the time-series fault association rule and the module fault association rule to obtain the fault detection rule.
[0187] It should be noted that in one or more embodiments provided in this specification, merging and configuring the time-series fault association rule and the module fault association rule to obtain the fault detection rule can be understood as merging and processing the time-series fault association rule and the module fault association rule to obtain a fault detection rule that includes the time-series fault association rule and the module fault association rule.
[0188] Alternatively, the step of combining and configuring the timing fault association rule and the module fault association rule to obtain the fault detection rule can be understood as follows: obtaining some timing fault association rules from the timing fault association rule, obtaining some module fault association rules from the module fault association rule, and combining and processing the some timing fault association rules and the some module fault association rules to obtain a fault detection rule including some timing fault association rules and some module fault association rules. Among them, the some timing fault association rules and the some module fault association rules can be selected according to the actual application scenario, and no specific limitation is made here.
[0189] Continuing with the above example, after obtaining the feature association relationships mined based on frequent item sets and frequent sequences, the feature association relationships are combined and configured to obtain the final operation and maintenance rules for downtime prediction. The specific method is as follows: for the association relationship A => B obtained from frequent sequence mining and the association relationship C => D obtained from frequent item set mining, these association relationships need to be combined and configured to obtain the downtime prediction rules for cloud servers.
[0190] It should be noted that after obtaining the downtime prediction rules, the downtime prediction rules can be applied to the maintenance process of cloud servers, and the downtime prediction rules are used as operation and maintenance rules. The operation and maintenance rules include the above-mentioned feature set A with downtime features and / or the hardware feature item set C. By monitoring the events in the log data of the cloud server, when it is determined that the abnormal event features generated during the operation of the cloud server match the operation and maintenance rules, that is, when all the features in the feature set A with downtime features appear, or all the features in the hardware feature item set C appear during the operation of the cloud server, it is determined that the rule is hit at this time, and there is a risk of downtime for this cloud server.
[0191] In the case where it is determined that the cloud server has a risk of downtime, the ECS of the physical machine with a risk of downtime and sensitive customers can be hot migrated to achieve the effects of downtime prediction and downtime avoidance.
[0192] In addition, in one or more embodiments provided in this specification, before applying the fault detection rule, the fault detection rule can be tested. The specific method can be: performing a dry run on the generated fault detection rule through an A / B test, and using an operation and maintenance evaluation algorithm (methods such as hypothesis testing) to evaluate whether the fault detection rule is significantly effective. Thus, it can be confirmed that the fault detection rule can effectively reduce the downtime probability of customers, and at the same time, it has little impact on the performance of customer ECS.
[0193] The fault detection rule determination method provided by the embodiments of this specification can process the operation data of a faulty data processing object through an operation data processing rule, obtain fault features, abnormal timing features, and abnormal module features, determine a timing fault association rule based on the fault features and abnormal timing features, and determine a module fault association rule based on the fault features and abnormal module features. Finally, based on the timing fault association rule and the module fault association rule, a fault detection rule that can accurately detect the fault problems of the data processing object is determined, thereby avoiding the problem of interruption of the data processing process caused by the fault of the data processing object, and enabling the data processing object to meet the actual application requirements of various application scenarios.
[0194] The following combines the attached Figure 3 , taking the application of the fault detection rule determination method provided by this specification in the scenario of cloud server downtime as an example, further illustrates the fault detection rule determination method. Among them, Figure 3 shows the processing process flowchart of a fault detection rule determination method provided by an embodiment of this specification. Based on Figure 3 it can be known that the processing process of the fault detection rule determination method provided by this specification includes four stages: 1. Data collection; 2. The monitoring platform extracts downtime features; 3. Association relationship generation; 4. The operation and maintenance platform applies the downtime prediction rule.
[0195] Among them, the data collection stage means that in the data collection stage, when the ECS is providing services, the data collection component ( Figure 3 the agent in it) will collect various log data of the ECS in real time, perform log backflow, and store the collected log data in the log database.
[0196] Among them, the monitoring platform extracts downtime features stage means that the monitoring platform will obtain the log data related to the downtime cloud server from the log data, configure some anomaly detection rules for the log data, generate anomaly data based on the anomaly detection rule, and perform feature extraction on the anomaly data to extract downtime features and various types of anomaly features.
[0197] Among them, the association relationship generation stage includes two execution modules: the frequent sequence mining module and the frequent itemset mining module;
[0198] The frequent sequence mining module includes four parts: 1. Perform time series extraction; 2. Perform frequent sequence mining; 3. Perform index statistics on the sequence pattern; 4. Generate an association relationship. For the steps of performing frequent sequence mining and generating an association relationship for this frequent sequence mining module, reference can be made to Figure 4 , Figure 4It is a process flow chart for frequent sequence mining and generating association relationships in a fault detection rule determination method provided by an embodiment of this specification, specifically including the following steps:
[0199] Step 402: Determine the ECS feature time series data.
[0200] Specifically, extract features from the abnormal data of the downed ECS to obtain downed features and abnormal features existing in the form of a time series.
[0201] Step 404: Extract the time series containing the downed features.
[0202] Specifically, use a time window that traces back 2 days from the moment where the downed feature is located to obtain the time series containing the downed feature within this time window.
[0203] Step 406: Conduct frequent sequence mining.
[0204] Input the time series into the PrefixSpan algorithm model to conduct frequent sequence mining on the time series and obtain frequent sequences, that is, sequence patterns.
[0205] Step 408: Calculate indicators such as the confidence level and lift of the sequence pattern.
[0206] Specifically, based on the sequence pattern and the downed feature, form an association relationship, and calculate indicators such as the confidence level, lift, and imbalance rate of this association relationship. The calculation methods of these indicators can refer to the content of the above formulas (1), (2), and (3).
[0207] Step 410: Determine whether the indicator threshold is met. If so, execute Step 412; if not, end.
[0208] Specifically, define an indicator threshold for these indicators, and compare the confidence level, lift, and imbalance rate with this indicator threshold. When both the confidence level and lift are greater than or equal to the indicator threshold, and the imbalance rate is less than or equal to the indicator threshold, execute Step 412; otherwise, end.
[0209] Step 412: Generate the association relationship of the downed feature.
[0210] Specifically, extract all association relationships where the confidence level, lift, and imbalance rate are all greater than or equal to the indicator threshold.
[0211] This frequent sequence mining module includes four parts: 1. Extract feature item sets; 2. Conduct frequent item set mining; 3. Statistic item set indicators. For the steps of conducting frequent item set mining and generating association relationships for this frequent item set mining module, reference can be made to Figure 5 , Figure 5It is a process flow chart for frequent item set mining and generating association relationships in a fault detection rule determination method provided by an embodiment of this specification, specifically including the following steps:
[0212] Step 502: Determine the ECS feature time series data.
[0213] Specifically, extract features from the abnormal data of the downed ECS to obtain downed features and abnormal features existing in the form of time series.
[0214] Step 504: Extract the time series containing hardware features.
[0215] Specifically, select a hardware feature set from the abnormal features. This hardware feature set is the union of hardware features and downed features. Generate an item set database based on the selected feature set.
[0216] Step 506: Perform frequent item set mining.
[0217] Specifically, remove duplicates from the duplicate elements in the item set database (duplicate elements are not allowed in the frequent item set mining algorithm). Then input the item set database into the FP-Growth algorithm model to mine the frequent item sets in the item set database.
[0218] Step 508: Calculate indicators such as the confidence level and lift of the association relationship.
[0219] Specifically, form an association relationship based on the frequent item sets and downed features, and calculate indicators such as the confidence level, lift, and imbalance rate of this association relationship. The calculation methods of these indicators can refer to the content of the above formulas (1), (2), and (3).
[0220] Step 510: Determine whether the indicator threshold is met. If so, execute step 512; if not, execute the end.
[0221] Specifically, define an indicator threshold for these indicators, and compare the confidence level, lift, and imbalance rate with this indicator threshold. When both the confidence level and lift are greater than or equal to the indicator threshold, and the imbalance rate is less than or equal to the indicator threshold, execute step 512; otherwise, execute the end.
[0222] Step 512: Generate the association relationship of the downed features.
[0223] Specifically, extract all association relationships where the confidence level, lift, and imbalance rate are all greater than or equal to the indicator threshold.
[0224] Among them, the operation and maintenance platform application downtime prediction rule means that: First, after obtaining the feature correlation relationships through frequent item set and frequent sequence mining, the two obtained correlation relationships are merged and configured to obtain the operation and maintenance rules for downtime prediction.
[0225] Second, the generated operation and maintenance rules are dry-run through A / B test, and the operation and maintenance evaluation algorithm (methods such as hypothesis testing) is used to evaluate whether the rule is significantly effective to obtain the operation and maintenance evaluation result.
[0226] Finally, when the operation and maintenance evaluation result meets the actual requirements, the operation and maintenance rules for downtime prediction are applied. The application method can be: monitoring the events in the log data of the cloud server. When it is determined that the abnormal event characteristics generated during the operation of the cloud server match the operation and maintenance rules, that is, when the logical conditions between the characteristics configured by the operation and maintenance rules are met, at this time it is determined that the operation and maintenance rules are hit or triggered, and there is a risk of downtime for this cloud server. Therefore, the ECS of the physical machine with a risk of downtime and sensitive customers can be hot migrated to achieve the effects of downtime prediction and downtime avoidance. Among them, the characteristics configured by the operation and maintenance rules can be all or part of the characteristics in the feature set A containing downtime characteristics, and / or all or part of the hardware feature items in the set C. And the logical conditions between the characteristics configured by the operation and maintenance rules can be logical relationships of AND / OR.
[0227] The method for determining the fault detection rule provided in this specification provides a solution for the downtime prediction rule based on frequent item set and frequent sequence mining, which is used for the prediction of downtime phenomena. Once the method predicts that the ECS will have a downtime behavior, the ECS is hot migrated from the current host to another healthier host, thus avoiding the occurrence of downtime. This solution is based on the method of frequent item set and frequent sequence mining, and analyzes the operation and maintenance rules for downtime prediction from the feature data of the ECS, which can reduce the downtime probability of downtime-sensitive customers.
[0228] Compared with the above solution that provides a hardware-based downtime prediction, the method for determining a fault detection rule provided in this specification determines a downtime prediction rule based on ECS features. By mining frequent item sets and frequent sequence patterns, features related to downtime are discovered, and operation and maintenance rules for downtime prediction are configured based on these features. This idea of generating a downtime prediction rule based on ECS feature data can discover the correlation relationship of downtime prediction based on the mining of frequent item sets and frequent sequences. The process is as follows: First, frequent item sets are mined from hardware features, and correlation relationships for downtime prediction are generated based on the frequent item sets. Second, frequent sequences are mined from the feature time series representing downtime, and confidence, lift, and imbalance rate are calculated for the mined sequence models to generate correlation relationships for downtime prediction. Finally, the correlation relationships generated in the above two steps are combined to configure the final operation and maintenance rules for downtime prediction, realizing the prediction and avoidance of downtime, thereby reducing the downtime rate of customer ECSs, and performing hot migration of ECSs in advance when the rules are hit to avoid potential downtime occurrences.
[0229] See Figure 6 , Figure 6 FIG. shows a schematic diagram of a system for determining a fault detection rule according to an embodiment of this specification. The system for determining a fault detection rule includes a cloud server 602 and a monitoring platform 604. Among them,
[0230] The cloud server 602 is configured to provide the operation data generated during operation to the monitoring platform 604;
[0231] The monitoring platform 604 is configured to determine the operation data of the failed cloud server; process the operation data through an operation data processing rule to obtain a fault feature, an abnormal time series feature, and an abnormal module feature. Among them, the abnormal module feature is a feature corresponding to a data processing module in the data processing object; determine a time series fault correlation rule based on the fault feature and the abnormal time series feature, and determine a module fault correlation rule based on the fault feature and the abnormal module feature; determine a fault detection rule based on the time series fault correlation rule and the module fault correlation rule.
[0232] Among them, the monitoring platform can be understood as a platform for determining a fault detection rule based on the operation data provided by the cloud server 602.
[0233] The cloud server in the fault detection rule determination system provided by the embodiments of this specification can provide operation data to the monitoring platform. The monitoring platform can process the operation data of the faulty data processing object through operation data processing rules to obtain fault characteristics, abnormal timing characteristics, and abnormal module characteristics, and determine timing fault association rules based on the fault characteristics and abnormal timing characteristics, and determine module fault association rules based on the fault characteristics and abnormal module characteristics. Finally, based on the timing fault association rules and module fault association rules, a fault detection rule that can accurately detect the fault problems of data processing objects is determined, thereby avoiding the problem of interruption of the data processing process caused by the failure of data processing objects, and enabling data processing objects to meet the actual application requirements of various application scenarios.
[0234] The above is a schematic solution of a fault detection rule determination system according to this embodiment. It should be noted that the technical solution of this fault detection rule determination system and the technical solution of the above fault detection rule determination method belong to the same concept. For the details not described in the technical solution of the fault detection rule determination system, reference can be made to the description of the technical solution of the above fault detection rule determination method.
[0235] See Figure 7 , Figure 7 shows a flowchart of a fault detection method according to an embodiment of this specification. This fault detection method is applied to a detection end and specifically includes the following steps.
[0236] Step 702: Configure the obtained fault detection rules, where the fault detection rules include timing fault association rules and module fault association rules;
[0237] Step 704: Determine the operation data of the cloud server in the host and determine abnormal timing characteristics and abnormal module characteristics based on the operation data;
[0238] Step 706: When it is determined that the abnormal timing characteristics satisfy the timing fault association rules and the abnormal module characteristics satisfy the module fault association rules, determine that the cloud server has a fault risk and perform hot migration on the cloud server in the host.
[0239] Among them, the detection end can be understood as a server that monitors the fault risk of the cloud server in the host. The detection end can be a server. For example, the detection end can be an operation and maintenance platform for the fault risk of the cloud server in the host.
[0240] Specifically, the fault detection method provided in this specification can obtain a fault detection rule and configure the fault detection rule locally. Among them, the fault detection rule includes a timing fault association rule and a module fault association rule. It should be noted that the fault detection rule can be the fault detection rule determined by the above-mentioned fault detection rule determination method.
[0241] After configuring the fault detection rule, it is necessary to monitor the running data of the cloud server running in the host, analyze the running data, and determine whether the abnormal timing characteristics and abnormal module characteristics meet the fault detection rule when determining the abnormal timing characteristics and abnormal module characteristics from the running data. It should be noted that for the step of determining the abnormal timing characteristics and abnormal module characteristics from the running data, reference can be made to the step of processing the running data through the running data processing rule to obtain the abnormal timing characteristics and abnormal module characteristics in the above-mentioned fault detection rule determination method, which will not be elaborated here.
[0242] By matching the abnormal timing characteristics and abnormal module characteristics with the fault detection rule, when it is determined that the abnormal timing characteristics conform to the timing fault association rule and the abnormal module characteristics meet the module fault association rule, it is determined that there is a fault risk for the cloud server running in the host. Therefore, it is necessary to perform hot migration on the cloud server to avoid problems such as service unavailability caused by the failure of the cloud server.
[0243] For example, the fault detection rule can be a downtime prediction rule (i.e., an operation and maintenance rule), and the detection end is the operation and maintenance platform. Based on this, in the fault detection method provided in this specification, the operation and maintenance platform can obtain the downtime prediction rule and apply the downtime prediction rule. Then the operation and maintenance platform monitors the log data of the cloud server in the host and extracts abnormal timing events (i.e., abnormal timing characteristics) and abnormal module events (i.e., abnormal timing characteristics) from the log data. Match the abnormal timing events and abnormal module events with the downtime prediction rule. When it is determined that the abnormal event characteristics such as the abnormal timing events and abnormal module events generated during the operation of the cloud server are consistent with the downtime prediction rule, that is, when the logical conditions between the abnormal event characteristics configured by the downtime prediction rule are met, at this time, it is determined that the downtime prediction rule is hit or triggered, and there is a risk of downtime for this cloud server. Therefore, the ECS of the host (i.e., the physical machine) with a risk of downtime and sensitive customers to downtime can be hot migrated to achieve the effects of downtime prediction and downtime avoidance. Thus, based on the fault detection rule, the prediction and avoidance of downtime are achieved, thereby reducing the downtime probability of the customer's ECS and avoiding the occurrence of downtime.
[0244] The fault detection method provided in this specification for the detection end can configure fault detection rules and monitor the operation data of cloud servers in the host. When it is determined that there are abnormal timing characteristics and abnormal module characteristics in the operation data, the abnormal timing characteristics and abnormal module characteristics are matched with the timing fault association rules. When it is determined that the abnormal timing characteristics meet the timing fault association rules and the abnormal module characteristics meet the module fault association rules, it is accurately determined that there is a fault risk in the cloud server, so as to perform a hot migration operation on the cloud server in the host, avoid the risk of cloud server failure, and avoid problems such as service unavailability caused by the failure of the cloud server.
[0245] Corresponding to the above method embodiment, this specification also provides an embodiment of a fault detection rule determination device. Figure 8 The structure diagram of a fault detection rule determination device provided by an embodiment of this specification is shown. As Figure 8 shown, the device includes:
[0246] A data determination module 802, configured to determine the operation data of the failed data processing object;
[0247] A feature determination module 804, configured to process the operation data through an operation data processing rule to obtain a fault feature, an abnormal timing feature, and an abnormal module feature, where the abnormal module feature is a feature corresponding to a data processing module in the data processing object;
[0248] A first rule determination module 806, configured to determine a timing fault association rule based on the fault feature and the abnormal timing feature, and determine a module fault association rule based on the fault feature and the abnormal module feature;
[0249] A second rule determination module 808, configured to determine a fault detection rule based on the timing fault association rule and the module fault association rule.
[0250] Optionally, the feature determination module 804 is further configured to:
[0251] Determine the abnormal data feature corresponding to the failed data processing object based on the operation data;
[0252] Obtain the fault feature and the abnormal timing feature by performing timing feature screening processing and timing feature mining processing on the abnormal data feature;
[0253] Obtain the fault feature and the abnormal module feature by performing module feature screening processing and module feature mining processing on the abnormal data feature.
[0254] Optionally, the feature determination module 804 is further configured to:
[0255] Determine the fault feature based on the abnormal data feature, perform time series feature screening processing on the abnormal data feature, and obtain candidate time series features;
[0256] Input the candidate time series features into a time series feature mining module for processing to determine the abnormal time series features among the candidate time series features.
[0257] Optionally, the feature determination module 804 is further configured to:
[0258] Determine the fault time corresponding to the processed object of the failed data based on the abnormal data feature;
[0259] Determine a feature time interval based on the fault time and a preset time series feature threshold;
[0260] Obtain the candidate time series features corresponding to the feature time interval from the abnormal data features.
[0261] Optionally, the feature determination module 804 is further configured to:
[0262] Input the candidate time series features into a frequent feature sequence mining model, perform frequent feature sequence mining on the candidate time series features through the frequent feature sequence mining model, and obtain the frequent feature sequences in the candidate time series features, where the frequent feature sequences include frequent features, and the frequent features are features that appear more than a preset number threshold in the frequent feature sequences;
[0263] Use the frequent feature sequences as the abnormal time series features.
[0264] Optionally, the first rule determination module 806 is further configured to:
[0265] Determine the time series correlation data between the fault feature and the abnormal time series features;
[0266] When it is determined that the time series correlation data meets a preset correlation condition, determine the time series fault correlation rule between the fault feature and the abnormal time series features.
[0267] Optionally, the first rule determination module 806 is further configured to:
[0268] Construct a candidate time series fault correlation rule based on the fault feature and the abnormal time series features;
[0269] Perform confidence calculation based on the fault feature and the abnormal time series features to obtain a time series confidence level;
[0270] Calculate the lift based on the abnormal data characteristics, the abnormal time series characteristics, and the fault characteristics to obtain the time series lift;
[0271] Determine the fault feature support of the fault feature and the time series feature support of the abnormal time series feature, and calculate the imbalance rate based on the fault feature support and the time series feature support to obtain the time series imbalance rate;
[0272] Use the time series confidence, the time series lift, and the time series imbalance rate as the time series association data of the candidate time series fault association rule.
[0273] Optionally, the first rule determination module 806 is further configured to:
[0274] In the case where it is determined that the time series confidence is greater than or equal to the preset association parameter threshold, the time series lift is greater than or equal to the preset association parameter threshold, and the time series imbalance rate is less than or equal to the preset association parameter threshold, use the candidate time series fault association rule as the time series fault association rule between the fault feature and the abnormal time series feature.
[0275] Optionally, the feature determination module 804 is further configured to:
[0276] Determine the fault feature based on the abnormal data characteristics, and perform module feature screening processing on the abnormal data characteristics to obtain candidate module features;
[0277] Input the candidate module features into the module feature mining module for processing to determine the abnormal module features in the candidate module features.
[0278] Optionally, the feature determination module 804 is further configured to:
[0279] Perform duplicate removal processing on the candidate module features to obtain the candidate module features after duplicate removal;
[0280] Input the candidate module features after duplicate removal into the frequent feature item set mining model for processing, and perform frequent feature item set mining on the candidate module features after duplicate removal through the frequent feature item set mining model to obtain the frequent feature item set, where the initial frequent feature item set includes frequent module features, and the frequent module features are module features whose occurrence times exceed the preset number threshold;
[0281] Use the frequent feature item set as the abnormal module feature.
[0282] Optionally, the first rule determination module 806 is further configured to:
[0283] Determine the module association data between the fault feature and the abnormal module feature;
[0284] When it is determined that the module association data satisfies the preset association condition, determine the module fault association rule between the fault feature and the abnormal module feature.
[0285] Optionally, the first rule determination module 806 is further configured to:
[0286] Construct a candidate module fault association rule based on the fault feature and the abnormal module feature;
[0287] Calculate the confidence based on the fault feature and the abnormal module feature to obtain the module confidence;
[0288] Calculate the lift based on the abnormal data feature, the abnormal module feature, and the fault feature to obtain the module lift;
[0289] Determine the fault feature support of the fault feature and the module feature support of the abnormal module feature, and calculate the imbalance rate based on the fault feature support and the module feature support to obtain the module imbalance rate;
[0290] Use the module confidence, the module lift, and the module imbalance rate as the module association data of the candidate module fault association rule.
[0291] Optionally, the first rule determination module 806 is further configured to:
[0292] When it is determined that the module confidence is greater than or equal to the preset association parameter threshold, the module lift is greater than or equal to the preset association parameter threshold, and the module imbalance rate is less than or equal to the preset association parameter threshold, use the candidate module fault association rule as the module fault association rule between the fault feature and the abnormal module feature.
[0293] Optionally, the feature determination module 804 is further configured to:
[0294] Determine the abnormal operation data of the failed data processing object from the operation data through the abnormal data screening rule;
[0295] Extract features from the abnormal operation data to determine the abnormal data feature corresponding to the failed data processing object.
[0296] Optionally, the second rule determination module 808 is further configured to:
[0297] Merge and configure the timing fault association rule and the module fault association rule to obtain the fault detection rule.
[0298] The fault detection rule determination device provided by the embodiments of this specification can process the operation data of a faulty data processing object through an operation data processing rule, obtain fault features, abnormal timing features, and abnormal module features, determine a timing fault association rule based on the fault features and the abnormal timing features, and determine a module fault association rule based on the fault features and the abnormal module features. Finally, based on the timing fault association rule and the module fault association rule, a fault detection rule that can accurately detect the fault problem of the data processing object is determined, thereby avoiding the problem of interruption of the data processing process caused by the fault of the data processing object, and enabling the data processing object to meet the actual application requirements of various application scenarios.
[0299] The above is a schematic solution of a fault detection rule determination device according to this embodiment. It should be noted that the technical solution of this fault detection rule determination device and the technical solution of the above fault detection rule determination method belong to the same concept. For the details not described in detail in the technical solution of this fault detection rule determination device, reference can be made to the description of the technical solution of the above fault detection rule determination method.
[0300] Corresponding to the above method embodiments, this specification also provides an embodiment of a fault detection device. This fault detection device is applied to a detection end and includes:
[0301] A rule configuration module, configured to configure the obtained fault detection rule, where the fault detection rule includes a timing fault association rule and a module fault association rule;
[0302] A feature determination module, configured to determine the operation data of a cloud server in a host and determine an abnormal timing feature and an abnormal module feature based on the operation data;
[0303] A fault determination module, configured to determine that there is a fault risk for the cloud server in the host and perform hot migration on the cloud server in the host when it is determined that the abnormal timing feature satisfies the timing fault association rule and the abnormal module feature satisfies the module fault association rule.
[0304] The fault detection device applied to the detection end provided in this specification can configure fault detection rules and monitor the operation data of the cloud servers in the host. When it is determined that there are abnormal timing characteristics and abnormal module characteristics in the operation data, the abnormal timing characteristics and abnormal module characteristics are matched with the timing fault association rules, and when it is determined that the abnormal timing characteristics meet the timing fault association rules and the abnormal module characteristics meet the module fault association rules, the fault risk of the cloud server is accurately determined, so as to perform a live migration operation on the cloud server in the host, avoid the risk of cloud server failure, and avoid problems such as service unavailability caused by the failure of the cloud server.
[0305] The above is a schematic solution of a fault detection device in this embodiment. It should be noted that the technical solution of this fault detection device and the technical solution of the above fault detection method belong to the same concept. For the details not described in detail in the technical solution of the fault detection device, reference can be made to the description of the technical solution of the above fault detection method.
[0306] Figure 9 The structural block diagram of a computing device 900 provided according to an embodiment of this specification is shown. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.
[0307] The computing device 900 further includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interfaces (e.g., network interface controller (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0308] In one embodiment of the present specification, the above components of the computing device 900 and Figure 9 other components not shown therein may also be connected to each other, for example, via a bus. It should be understood that Figure 9 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0309] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.
[0310] Wherein, the processor 920 is configured to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above-mentioned fault detection rule determination method are implemented.
[0311] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-mentioned method for determining a fault detection rule belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above-mentioned method for determining a fault detection rule.
[0312] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned method for determining a fault detection rule.
[0313] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-mentioned method for determining a fault detection rule belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above-mentioned method for determining a fault detection rule.
[0314] An embodiment of this specification also provides a computer program, which, when executed on a computer, causes the computer to execute the steps of the above-mentioned method for determining a fault detection rule.
[0315] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above-mentioned method for determining a fault detection rule belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above-mentioned method for determining a fault detection rule.
[0316] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0317] The computer instructions include computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0318] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.
[0319] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0320] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for determining a fault detection rule, comprising: Determine the operation data of the failed data processing object; Process the operation data through an operation data processing rule to obtain a fault feature, an abnormal timing feature, and an abnormal module feature, wherein the abnormal module feature is a feature corresponding to a data processing module in the data processing object; Determine a timing fault association rule based on the fault feature and the abnormal timing feature, and determine a module fault association rule based on the fault feature and the abnormal module feature; Determine a fault detection rule based on the timing fault association rule and the module fault association rule.
2. The method for determining a fault detection rule according to claim 1, wherein the process of processing the operation data through an operation data processing rule to obtain a fault feature, an abnormal timing feature, and an abnormal module feature comprises: Determine an abnormal data feature corresponding to the failed data processing object based on the operation data; Obtain the fault feature and the abnormal timing feature by performing timing feature screening processing and timing feature mining processing on the abnormal data feature; Obtain the fault feature and the abnormal module feature by performing module feature screening processing and module feature mining processing on the abnormal data feature.
3. The method for determining a fault detection rule according to claim 2, wherein the process of obtaining the fault feature and the abnormal timing feature by performing timing feature screening processing and timing feature mining on the abnormal data feature comprises: Determine the fault feature based on the abnormal data feature, and perform timing feature screening processing on the abnormal data feature to obtain candidate timing features; Input the candidate timing features into a timing feature mining module for processing to determine the abnormal timing features among the candidate timing features.
4. The method for determining a fault detection rule according to claim 3, wherein the process of performing timing feature screening processing on the abnormal data feature to obtain candidate timing features comprises: Determine the fault time corresponding to the failed data processing object based on the abnormal data feature; Determine a feature time interval based on the fault time and a preset timing feature threshold; Obtain the candidate timing features corresponding to the feature time interval from the abnormal data features.
5. The method for determining a fault detection rule according to claim 3, wherein the process of inputting the candidate timing features into a timing feature mining module for processing to determine the abnormal timing features among the candidate timing features comprises: Input the candidate timing features into a frequent feature sequence mining model, and perform frequent feature sequence mining on the candidate timing features through the frequent feature sequence mining model to obtain frequent feature sequences among the candidate timing features, wherein the frequent feature sequences include frequent features, and the frequent features are features that appear more than a preset number threshold in the frequent feature sequences; Use the frequent feature sequences as the abnormal timing features.
6. The method for determining a fault detection rule according to claim 2, wherein the determining of the temporal fault association rule based on the fault feature and the abnormal temporal feature includes: Determining the temporal association data between the fault feature and the abnormal temporal feature; When it is determined that the temporal association data satisfies a preset association condition, determining the temporal fault association rule between the fault feature and the abnormal temporal feature.
7. The method for determining a fault detection rule according to claim 6, wherein the determining of the temporal association data between the fault feature and the abnormal temporal feature includes: Based on the fault feature and the abnormal temporal feature, constructing a candidate temporal fault association rule; Calculating a temporal confidence based on the fault feature and the abnormal temporal feature; Calculating a temporal lift based on the abnormal data feature, the abnormal temporal feature, and the fault feature; Determining the fault feature support of the fault feature and the temporal feature support of the abnormal temporal feature, and calculating an imbalance rate based on the fault feature support and the temporal feature support to obtain a temporal imbalance rate; Taking the temporal confidence, the temporal lift, and the temporal imbalance rate as the temporal association data of the candidate temporal fault association rule.
8. The method for determining a fault detection rule according to claim 7, wherein the determining of the temporal fault association rule between the fault feature and the abnormal temporal feature when it is determined that the temporal association data satisfies a preset association condition includes: When it is determined that the temporal confidence is greater than or equal to a preset association parameter threshold, the temporal lift is greater than or equal to the preset association parameter threshold, and the temporal imbalance rate is less than or equal to the preset association parameter threshold, taking the candidate temporal fault association rule as the temporal fault association rule between the fault feature and the abnormal temporal feature.
9. The method for determining a fault detection rule according to claim 2, wherein the obtaining of the fault feature and the abnormal module feature by performing module feature screening processing and module feature mining processing on the abnormal data feature includes: Determining the fault feature based on the abnormal data feature, and performing module feature screening processing on the abnormal data feature to obtain candidate module features; Inputting the candidate module features into a module feature mining module for processing to determine the abnormal module features in the candidate module features.
10. The method for determining a fault detection rule according to claim 9, wherein the inputting of the candidate module features into a module feature mining module for processing to determine the abnormal module features in the candidate module features includes: Performing duplicate removal processing on the candidate module features to obtain the candidate module features after duplicate removal; Input the deduplicated candidate module features into a frequent feature item set mining model for processing. Mine the frequent feature item set from the deduplicated candidate module features through the frequent feature item set mining model to obtain a frequent feature item set, where the initial frequent feature item set includes frequent module features, and the frequent module features are module features whose occurrence times exceed a preset times threshold; Use the frequent feature item set as the abnormal module features.
11. The method for determining a fault detection rule according to claim 2, wherein determining the module fault association rule based on the fault features and the abnormal module features includes: Determine the module association data between the fault features and the abnormal module features; When it is determined that the module association data satisfies a preset association condition, determine the module fault association rule between the fault features and the abnormal module features.
12. The method for determining a fault detection rule according to claim 11, wherein determining the module association data between the fault features and the abnormal module features includes: Based on the fault features and the abnormal module features, construct a candidate module fault association rule; Calculate the confidence degree based on the fault features and the abnormal module features to obtain the module confidence degree; Calculate the lift degree based on the abnormal data features, the abnormal module features, and the fault features to obtain the module lift degree; Determine the fault feature support degree of the fault features and the module feature support degree of the abnormal module features, and calculate the imbalance rate based on the fault feature support degree and the module feature support degree to obtain the module imbalance rate; Use the module confidence degree, the module lift degree, and the module imbalance rate as the module association data of the candidate module fault association rule.
13. The method for determining a fault detection rule according to claim 12, wherein when it is determined that the module association data satisfies a preset association condition, determining the module fault association rule between the fault features and the abnormal module features includes: When it is determined that the module confidence degree is greater than or equal to a preset association parameter threshold, the module lift degree is greater than or equal to the preset association parameter threshold, and the module imbalance rate is less than or equal to the preset association parameter threshold, use the candidate module fault association rule as the module fault association rule between the fault features and the abnormal module features.
14. The method for determining a fault detection rule according to claim 2, wherein determining the abnormal data features corresponding to the already-faulty data processing object based on the operation data includes: Determine the abnormal operation data of the already-faulty data processing object from the operation data through the abnormal data screening rule; Extract features from the abnormal operation data to determine the abnormal data features corresponding to the already-faulty data processing object.
15. The method for determining a fault detection rule according to claim 1, wherein determining the fault detection rule based on the time-series fault association rule and the module fault association rule includes: Merge and configure the timing fault association rule and the module fault association rule to obtain the fault detection rule.
16. A system for determining a fault detection rule, the system comprising a cloud server and a monitoring platform, wherein the cloud server is configured to provide the operation data generated during operation to the monitoring platform; the monitoring platform is configured to determine the operation data of the failed cloud server; process the operation data through an operation data processing rule to obtain a fault feature, an abnormal timing feature, and an abnormal module feature, wherein the abnormal module feature is a feature corresponding to a data processing module in the data processing object; determine a timing fault association rule based on the fault feature and the abnormal timing feature, and determine a module fault association rule based on the fault feature and the abnormal module feature; determine a fault detection rule based on the timing fault association rule and the module fault association rule.
17. A fault detection method applied to a detection end, comprising: configure the obtained fault detection rule, wherein the fault detection rule includes a timing fault association rule and a module fault association rule; determine the operation data of the cloud server in the host and determine an abnormal timing feature and an abnormal module feature based on the operation data; when it is determined that the abnormal timing feature satisfies the timing fault association rule and the abnormal module feature satisfies the module fault association rule, determine that the cloud server has a fault risk, and perform hot migration on the cloud server in the host.
18. A computing device, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the fault detection rule determination method according to any one of claims 1 to 15 or the fault detection method according to any one of claim 17 are implemented.
19. A computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the fault detection rule determination method according to any one of claims 1 to 15 or the fault detection method according to any one of claim 17 are implemented.