Method and apparatus for determining batch faults, computer storage medium, and electronic device
By obtaining single-body fault information and configuration information, performing configuration dimension expansion and frequent item set mining, the problem of complexity of batch fault positioning is solved, fast and accurate fault positioning and early warning is achieved, and the operation and maintenance efficiency of the data center is improved.
Patent Information
- Application Number
- CN202010121380.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-02-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-02-26
AI Technical Summary
In the prior art, batch fault positioning is complex, and it is difficult to quickly and accurately locate and handle faults in large-scale fault scenarios, resulting in insufficiency in data center operation and maintenance.
By obtaining single-unit fault information and configuration information, the configuration dimension expansion is carried out, the single-unit fault dimension data set is constructed, and the data set for batch failure is determined by using frequent item set mining and failure rate comparison, the alarm is issued and false judgment is carried out.
It reduces the complexity of batch fault positioning, improves the accuracy and speed of fault positioning, reduces the risk of misjudgment, and ensures the stable operation of the data center.
Smart Images

Figure CN113312197B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and device for determining batch faults. This application also relates to a batch fault warning system, a computer storage medium, and an electronic device. Background Art
[0002] With the development of cloud computing and big data, the scale of data centers has become increasingly large, and a large number of servers have been purchased and deployed. To process big data, there are a large number of applications, a large number of servers, and a large number of components. During the operation of the data center, there is a possibility of faults occurring. The current fault forms can include single-point faults and batch faults.
[0003] A single-point fault refers to a fault that occurs in a single independent application, server, or component in the data center. Single-point faults can all be masked through fault tolerance technology.
[0004] A batch fault refers to a fault that occurs in a large range of service devices or software applications. For example, a fault that occurs in any one or more of a large number of applications, a large number of servers, and a large number of components within the same time period or within devices provided by the same vendor. Moreover, many faults occur in specific services, specific computer rooms, and specific manufacturers. Therefore, fault location has become extremely complex. In the complex scenarios where faults occur, simple software fault tolerance technology cannot handle the faults. Summary of the Invention
[0005] This application provides a method for determining batch faults to solve the problem of the complexity of batch fault location in the prior art.
[0006] This application provides a method for determining batch faults, including:
[0007] Obtain single-point fault information and configuration information for describing the service devices in the data center;
[0008] According to the configuration information, expand the single-point fault information in terms of configuration dimensions to obtain a set of single-point fault dimension data;
[0009] According to the set of single-point fault dimension data and the set batch fault judgment conditions, determine the set of batch fault data.
[0010] In some embodiments, the obtaining of the single-point fault information includes:
[0011] Obtain the single-point fault information of a single entity monitored in the data center.
[0012] In some embodiments, it further includes:
[0013] Format the single - unit fault information to obtain a single - unit fault work order;
[0014] The obtaining of the configuration information used to describe the data center service equipment includes:
[0015] According to the single - unit fault work order, obtain the configuration information used to describe the data center service equipment in the configuration management database, where the configuration management database stores the configuration information describing entities in the network environment.
[0016] In some embodiments, the expanding the single - unit fault information in the configuration dimension according to the configuration information to obtain a single - unit fault dimension data set includes:
[0017] Determine the configuration dimension according to the configuration items in the configuration information;
[0018] Construct the single - unit fault dimension data set according to the configuration dimension and the single - unit fault information.
[0019] In some embodiments, it further includes:
[0020] Determine a candidate fault dimension data set according to the correlation analysis between the single - unit fault dimension data sets;
[0021] The determining of the data set of batch faults according to the single - unit fault dimension data set and the set batch - fault judgment conditions includes:
[0022] Determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch - fault judgment conditions.
[0023] In some embodiments, the determining of the candidate fault dimension data set according to the correlation analysis between the single - unit fault dimension data sets includes:
[0024] Perform frequent item set mining on the single - unit fault dimension data set;
[0025] Determine the frequent item sets whose occurrence frequencies meet the occurrence - frequency requirements in the frequent item set range as the candidate fault dimension data set.
[0026] In some embodiments, the determining of whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch - fault judgment conditions includes:
[0027] Calculate the corresponding failure rate within the candidate fault dimension data set;
[0028] Compare the failure rate with the set failure rate baseline value. If the failure rate is greater than or equal to the failure rate baseline value, it is determined that there are batch failures in the candidate failure dimension data set corresponding to the failure rate.
[0029] In some embodiments, it further includes:
[0030] When the failure rate is compared with the failure rate baseline value, if the failure rate is less than the failure rate baseline value, it is determined that there are no batch failures in the candidate failure dimension data set corresponding to the failure rate.
[0031] In some embodiments, it further includes:
[0032] When it is determined that the candidate failure dimension data set is a data set of batch failures, a batch failure alarm is issued.
[0033] In some embodiments, it further includes:
[0034] Perform misjudgment detection on the determined data set of batch failures.
[0035] In some embodiments, the performing misjudgment detection on the determined data set of batch failures includes:
[0036] Pull the black box logs in the network environment;
[0037] According to the data in the black box logs, detect whether there is a misjudgment in the determined data set of batch failures.
[0038] The present application also provides a device for determining batch failures, including:
[0039] An acquisition unit, configured to acquire single - entity failure information and configuration information for describing data center service devices;
[0040] An expansion unit, configured to perform configuration - dimension expansion on the single - entity failure information according to the configuration information to obtain a single - entity failure dimension data set;
[0041] A determination unit, configured to determine a data set of batch failures according to the single - entity failure dimension data set and set batch - failure judgment conditions.
[0042] The present application also provides a method for monitoring batch failures, including:
[0043] Collect the single - entity failure information of the data center through a deployed monitoring module for monitoring the data center;
[0044] Send the collected single - entity failure information to the monitoring service management center.
[0045] In some embodiments, the monitoring module deployed for monitoring the data center collects the single - point failure information of the data center, including:
[0046] Through the configuration of the monitoring service management center for the monitoring module, the monitoring module configured to collect the single - point failure information is deployed in the data center.
[0047] In some embodiments, the monitoring module deployed for monitoring the data center collects the single - point failure information of the data center, including:
[0048] The monitoring module for collecting the single - point failure information of the data center is deployed on the servers in the data center.
[0049] This application also provides a monitoring device for batch failures, including:
[0050] A collection unit for collecting the single - point failure information of the data center through the deployed monitoring module for monitoring the data center;
[0051] A sending unit for sending the collected single - point failure information to the monitoring service management center.
[0052] This application also provides a fault warning system, including: a data center and a monitoring service management center; wherein, the data center is used to collect single - point failure information; the monitoring service management center is used to perform configuration - dimension expansion on the single - point failure information according to the obtained single - point failure information and the obtained configuration information for describing the service devices in the data center to obtain a set of single - point failure dimension data; and determine a set of batch - failure data according to the set of single - point failure dimension data and the set batch - failure judgment conditions.
[0053] In some embodiments, the fault warning system includes: deploying a monitoring module on the servers in the data center to monitor the single - point failure information in the data center.
[0054] In some embodiments, the fault warning system includes: the monitoring service management center issues a batch - failure alarm according to the determined set of batch - failure data.
[0055] This application also provides a computer storage medium for storing the data generated by the network platform and the programs for processing the data generated by the network platform;
[0056] When the program is read and executed, it performs the following steps:
[0057] Obtain single - point failure information and configuration information for describing the service devices in the data center;
[0058] According to the configuration information, perform configuration dimension expansion on the single-point fault information to obtain a set of single-point fault dimension data;
[0059] According to the set of single-point fault dimension data and the set batch fault judgment conditions, determine the set of batch fault data;
[0060] Alternatively, perform the following steps:
[0061] Collect the single-point fault information of the data center through the deployed monitoring module for monitoring the data center;
[0062] Send the collected single-point fault information to the monitoring service management center.
[0063] This application also provides an electronic device, including:
[0064] A processor;
[0065] A memory for storing a program for processing data generated by the network platform, and when the program is read and executed by the processor, the following steps are performed:
[0066] Obtain single-point fault information and configuration information for describing the service devices of the data center;
[0067] According to the configuration information, perform configuration dimension expansion on the single-point fault information to obtain a set of single-point fault dimension data;
[0068] According to the set of single-point fault dimension data and the set batch fault judgment conditions, determine the set of batch fault data;
[0069] Alternatively, perform the following steps:
[0070] Collect the single-point fault information of the data center through the deployed monitoring module for monitoring the data center;
[0071] Send the collected single-point fault information to the monitoring service management center.
[0072] Compared with the prior art, this application has the following advantages:
[0073] A method for determining batch faults provided by this application can expand the obtained single faults in terms of configuration dimensions through the obtained configuration information describing the data center service devices, obtain a set of single-fault dimension data, and then determine a set of batch-fault data according to the set of single-fault dimension data and the set batch-fault judgment conditions. It can be seen that this application expands the single-fault information in dimensions according to the configuration information describing the data center service devices to obtain an expanded set of single-fault dimension data, and then finds a set of hot data with faults in the set of single-fault dimension data according to the set of single-fault dimension data and the set batch-fault judgment conditions. These sets of hot data are the set of batch-fault data, thereby reducing the complexity of batch-fault location.
[0074] In addition, in the embodiments of this application, the determined set of batch-fault data is also detected to avoid the possibility of misjudging batch faults and reduce the risks existing in dealing with batch faults. Brief Description of the Drawings
[0075] Figure 1 is a flowchart of an embodiment of a method for determining batch faults provided by this application;
[0076] Figure 2 is a schematic structural diagram of an embodiment of a device for determining batch faults provided by this application;
[0077] Figure 3 is a schematic system architecture diagram of an embodiment of a fault warning system provided by this application;
[0078] Figure 4 is a flowchart of an embodiment of a method for monitoring batch faults provided by this application;
[0079] Figure 5 is a schematic structural diagram of an embodiment of a device for monitoring batch faults provided by this application. Detailed Embodiments
[0080] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific embodiments disclosed below.
[0081] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The descriptive methods used in this application and the appended claims, such as "a kind of", "the first", and "the second", etc., are not intended to limit the quantity or the order of precedence, but are used to distinguish the same type of information from each other.
[0082] Based on the above description of the background art, the prior art will be further elaborated in combination with the specific application scenarios of the method for determining batch faults provided in this application. Currently, the processing of massive data can be completed through a data center. Therefore, the application scenario of the embodiments of this application can use the data center as the application environment. Of course, it is not limited to this scenario of the data center. Any environment for processing massive data can achieve the technical purpose of this application. The data center needs to run all day long, and it is inevitable that faults will occur. How to quickly find the cause of the fault and eliminate the fault is the most direct manifestation of the operation and maintenance efficiency of the data center. Once a fault occurs in the data center, it will cause huge economic losses to the data center. However, when the data center faces massive data, due to the certain complexity of the environment where the massive data is located, when dealing with massive data processing, once a large-scale fault occurs, due to the complexity of the fault occurrence, it is difficult to find the cause of the fault in a short time. Therefore, to ensure the normal operation of the data center, it is necessary to have a certain prediction of large-scale faults (i.e., batch faults), that is, to discover and then process them. Currently, the prior art does not have effective measures for locating batch faults. It only performs fault processing when a fault occurs in a small range. For example, when the number of faulty devices exceeds a set threshold, a fault alarm is issued. When a large-scale fault occurs, due to the complex fault environment, it is difficult for the monitoring to cope with the fault location in a complex scenario. For this purpose, this application provides a method for determining batch faults, which can locate batch faults in the complex scenario of massive data, so as to give an early warning in advance, avoid the inability to handle the massive data processing scenario due to the outbreak of batch problems, and prevent the data processing from being paralyzed.
[0083] The following will introduce in detail a method for determining batch faults provided in this application. Please refer to Figure 1 as shown in Figure 1 is a flowchart of an embodiment of a method for determining batch faults provided in this application.
[0084] As Figure 1 shown, the method for determining batch faults provided in the embodiments of this application includes:
[0085] Step S101: Obtain single-fault information and configuration information for describing the service devices of the data center.
[0086] First, the nouns in step S101 are explained. In this embodiment, a single-fault can be understood as a fault that occurs in an independent hardware device or component or an independent software application product. For example, a fault that occurs in a CPU, a fault that occurs in a memory, etc. The single-fault information is the fault information describing an independent hardware device or an independent software application product. For example, **the component cannot be accessed.
[0087] In this embodiment, the configuration information used to describe the data center service equipment can be obtained through a Configuration Management Database (CMDB). The CMDB can be understood as storing and managing various configuration information of devices in the IT architecture. It is closely connected to all service support and service delivery processes, supports the operation of these processes, gives play to the value of the configuration information, and at the same time depends on the relevant processes to ensure the accuracy of the data. The CMDB includes entities and the configuration information for the entities. Among them, the entity can be understood as a configuration item. The entity can include hardware devices, such as network devices, storage devices, security devices, computer room devices, network ports, etc., as well as sub-configuration items of the devices, that is, the configuration items can be hierarchically set. The configuration information can be understood as the attribute information of the configuration item. For example, the configuration information can be the device name, serial number, model, product line, application group, production number, capacity, interface rate, etc.
[0088] In the specific implementation process of obtaining the single entity failure information and the data in the CMDB in step S101, there is no specific limitation on the acquisition sequence. It is possible to obtain the single entity failure information first and then obtain the data in the CMDB; it is also possible to obtain the data in the CMDB first and then obtain the single entity failure information; it is also possible to obtain the single entity failure information and the data in the CMDB separately.
[0089] In the application embodiment, the obtaining of the single entity failure information can specifically be obtaining the single entity failure information of a single entity monitored by the data center.
[0090] The data center can be understood as a specific network of globally collaborative devices used to transfer, accelerate, display, calculate, and store data information on the Internet network infrastructure. In this embodiment, the single entity failure information is obtained through the monitoring of a single entity by the data center. It can be understood that the data center includes a large number of entities, so the monitored single entity failure information can come from multiple entities.
[0091] To facilitate the computer to process the monitored data, therefore, formatting operations are performed on the single entity failure information of the single entity monitored in the obtained data to obtain a single entity failure work order. The single entity failure work order is used to describe the formatted information of the single entity failure.
[0092] In this embodiment, when obtaining the data in the CMDB, the data in the CMDB can be obtained according to the single entity failure work order. The specific obtaining method can be completed through the interface (API) between the data center and the CMDB.
[0093] The configuration information of an entity can be obtained from the configuration management database, which usually includes all entities involved in the data service process, and thus the configuration information of each entity can be obtained. Therefore, through the configuration information in the configuration management database, the entity information in the data center for massive data processing can be obtained, and furthermore, the specific information of all single-point failures in the data center can be obtained.
[0094] Step S102: According to the configuration information, perform configuration dimension expansion on the single-point failure information to obtain a set of single-point failure dimension data.
[0095] The purpose of step S102 is to expand the obtained single-point failure information in high dimensions to more comprehensively obtain the specific failure content involved in the single-point failure information.
[0096] Therefore, the specific implementation process of step S102 may include:
[0097] Step S102-1: Determine the configuration dimension according to the configuration items in the configuration information;
[0098] Step S102-2: Construct a set of single-point failure dimension data according to the configuration dimension and the single-point failure information.
[0099] Based on step S101, it can be known that the configuration information is an attribute description of an entity (configuration item). Therefore, the configuration dimension may include at least one of the following dimensions:
[0100] Entity model, entity product line, application group of the entity, firmware version of the entity, component model of the entity, production number of the entity, serial number of the entity, interface rate of the entity, capacity of the entity.
[0101] The purpose of step S102-2 is to expand the single-point failure information in the configuration dimension, thereby constructing a set of single-point failure dimension data for the single-point failure information. For a vivid understanding, the following example can be referred to:
[0102] The single-point failure information may be that the storage device fails to store data, the network device fails to access, etc. After formatting, it may be storage device, storage failure; network device, access failure. According to the single-point failure information and the data in the configuration management database, the constructed set of single-point failure dimension data may include relevant information of the storage device and the network device in dimensions such as entity model dimension, entity product line dimension, application group dimension of the entity, firmware version dimension of the entity, component model dimension of the entity, production number dimension of the entity, serial number dimension of the entity, interface rate dimension of the entity, capacity dimension of the entity, etc. That is to say, according to the single-point failure information, a multi-dimensional data set for multiple failure information can be constructed.
[0103] It should be noted that, in this embodiment, the obtained single - unit fault information can be real - time or obtained periodically. The servers in the data center obtain the single - unit fault information. Usually, there are multiple servers in the data center. Therefore, each server obtains the single - unit fault information it monitors. How the data center specifically obtains the single - unit fault information will be specifically described in the subsequent fault warning system.
[0104] Step S103: Determine the data set of batch faults according to the single - unit fault dimension data set and the set batch - fault judgment conditions.
[0105] The purpose of step S103 is to find out the data set of batch faults in the constructed single - unit dimension data set.
[0106] The specific implementation process of step S103 may include:
[0107] Calculate the corresponding failure rate in the single - unit fault dimension data set; compare the failure rate with the set failure - rate baseline value, and determine the data set of batch faults in the single - unit fault dimension data set according to the comparison result. The specific determination of the data set of batch faults will be described in detail below.
[0108] In order to narrow the scope of determining batch faults, this embodiment of the present application may further include:
[0109] Step S10 + 1: Determine the candidate fault dimension data set according to the correlation analysis between the single - unit fault dimension data sets.
[0110] The purpose of step S10 + 1 is to analyze the correlation relationship between the single - unit fault dimension data sets, screen out the candidate fault dimension data sets, so as to narrow the scope of the data set of batch faults. The specific implementation process may include:
[0111] Step S10 + 11: Mine frequent item sets for the single - unit fault dimension data sets;
[0112] Step S10 + 12: Determine the frequent item sets whose occurrence frequencies meet the occurrence - frequency requirements in the frequent item - set range as the candidate fault dimension data sets.
[0113] The item sets in the frequent item sets are sets of several items. For example, the set of configuration dimensions in this embodiment can be regarded as an item set. Find out the sets whose support degrees are greater than or equal to the minimum support degree (min_sup) from these item sets. Among them, the support degree refers to the frequency of a certain set appearing in all transactions. Frequent item - set mining is the basis for many important data - mining tasks such as association rules, correlation analysis, causal relationships, sequential item sets, local periodicity, and episode segments.
[0114] The mining of frequent itemsets can adopt algorithms such as Apriori, FP-growth, and FP-Tree. Taking the FP-growth algorithm as an example for an overview:
[0115] Step a: Scan the data set of single-fault dimension, count the fault dimensions, and count the number of times of the fault dimensions.
[0116] Step b: Set the minimum support according to requirements. For example, the minimum support is 2.
[0117] Step c: Sort the statistical data in step a. The descending order can be adopted to sort the data set of single-fault dimensions after statistics. If the number of times a fault dimension appears is less than 2, it is deleted.
[0118] Step d: Build an FP-tree based on step 3, and mine frequent itemsets based on the built FP-tree.
[0119] The above content is only an overview of mining frequent itemsets using the FP-growth algorithm.
[0120] Finally, the data set after excluding those with the number of occurrences less than 2 can be determined as the candidate fault dimension data set.
[0121] Based on the above content, it can be known that the requirement for the occurrence frequency can be the requirement for the number of times a fault occurs. For example, the support degree of 2 above. Of course, the value of the support degree can also be set according to the actual situation. When the statistical fault dimension data is less than the value of the support degree, the fault dimension is thrown out, that is, the possibility of a batch fault occurring for this fault dimension can be ignored.
[0122] It can be seen that the correlation analysis between the data sets of single-fault dimensions can be carried out by using the method of mining frequent itemsets, so as to narrow the scope of batch fault location, that is, exclude the fault dimension information with a relatively small probability of batch fault occurrence, which significantly improves both the complexity reduction of location and the location processing speed.
[0123] Based on the above candidate fault dimension data set, the specific implementation process of step S103 can also be:
[0124] Step S301-1: Calculate the failure rate corresponding to the candidate fault dimension data set;
[0125] Step S301-2: Compare the failure rate with the set failure rate baseline value. If the failure rate is greater than or equal to the failure rate baseline value, it is determined that there is a batch fault in the candidate fault dimension data set corresponding to the failure rate.
[0126] In this embodiment, the failure rate can be the ratio of the number of failures in the candidate failure dimension data set to the number of devices that meet the candidate failure dimension data set.
[0127] The baseline value of the failure rate can be a reference value set according to industry standards and operation experience. Of course, it can also be a threshold value set according to actual requirements.
[0128] Of course, a range can be set during judgment. For example, when the failure rate of the calculated candidate failure dimension data set is greater than N times the baseline value of the failure rate, it is determined that there are batch failures in the candidate failure dimension data set corresponding to the failure rate. Among them, N times can be adjusted according to actual requirements, and the specific value of N can be determined according to the situation of batch failures.
[0129] It also includes:
[0130] When the failure rate is compared with the baseline value of the failure rate, if the failure rate is less than the baseline value of the failure rate, it is determined that there are no batch failures in the candidate failure dimension data set corresponding to the failure rate.
[0131] In this embodiment, it also includes:
[0132] When it is determined that the candidate failure dimension data set is a data set of batch failures, a batch failure alarm is issued so as to be able to give an early warning for the discovered batch failures.
[0133] It can be understood that after the data set of batch failures is determined, there may be misjudgments. Therefore, this embodiment may also include:
[0134] Perform misjudgment detection on the determined data set of batch failures. Specifically, it can be to pull the black box log of the network environment, and detect whether there are misjudgments in the determined data set of batch failures according to the data in the black box log. For example: determine misjudgments based on firmware kernel data, that is, the black box log can be the kernel data of the firmware, and of course it can also be other data contents.
[0135] The above is the description process of an embodiment of a method for determining batch failures provided by this application. It can be seen that in this embodiment, by performing high-dimensional expansion on the obtained single-failure information, a single-failure dimension data set in multiple dimensions is obtained, and then through the correlation analysis between the single-failure dimension data sets, the candidate failure dimension data set is screened out, and then the hot dimension of batch failures is determined within the candidate failure dimension data set, thereby narrowing the scope of determining batch failures, and greatly reducing the complexity of locating batch problems by using the powerful computing power of the data center.
[0136] The above is a specific description of an embodiment of a method for determining batch faults provided by this application. Corresponding to the embodiment of the method for determining batch faults provided above, this application also discloses an embodiment of a device for determining batch faults. Please refer to Figure 2 , since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0137] As Figure 2 shown, Figure 2 is a schematic structural diagram of an embodiment of a device for determining batch faults provided by this application. The device includes:
[0138] An acquisition unit 201, configured to acquire single - entity fault information and configuration information for describing data center service devices;
[0139] Specifically, the acquisition unit 201 is configured to acquire single - entity fault information of a single entity monitored by the data center; and acquire through a Configuration Management Database (CMDB). The Configuration Management Database can be understood as storing and managing various configuration information of devices in the IT architecture. It is closely associated with all service support and service delivery processes, supports the operation of these processes, gives play to the value of configuration information, and at the same time depends on relevant processes to ensure the accuracy of data. The Configuration Management Database includes entities and configuration information for the entities. Here, the entity can be understood as a configuration item. The entity may include hardware devices, such as: network devices, storage devices, security devices, computer room devices, network ports, etc., and sub - configuration items of the devices, that is, the configuration items can be hierarchically set. The configuration information can be understood as the attribute information of the configuration item. For example: the configuration information can be device name, serial number, model, product line, application group, production number, capacity, interface rate, etc. Specifically, refer to the specific description of step S101 above, and details will not be repeated here.
[0140] It further includes: a formatting unit, configured to perform a formatting operation on the single - entity fault information to obtain a single - entity fault work order.
[0141] When the acquisition unit 201 acquires data in the Configuration Management Database, specifically, it may acquire the data in the Configuration Management Database according to the single - entity fault work order.
[0142] An extension unit 202, configured to perform configuration - dimension extension on the single - entity fault information according to the configuration information to obtain a set of single - entity fault dimension data;
[0143] Specifically, the extension unit 202 includes: a configuration - dimension determination subunit and a construction subunit;
[0144] The configuration dimension determination subunit is configured to determine a configuration dimension according to configuration items in the configuration information.
[0145] The construction subunit is configured to construct the single-fault dimension data set according to the configuration dimension and the single-fault information.
[0146] The determination unit 203 is configured to determine a data set of batch faults according to the single-fault dimension data set and a set batch fault determination condition.
[0147] It further includes: an analysis unit, specifically configured to determine a candidate fault dimension data set according to an association analysis between the single-fault dimension data sets.
[0148] The analysis unit includes: a mining subunit and a determination subunit;
[0149] The mining subunit is configured to perform frequent item set mining on the single-fault dimension data set;
[0150] The determination subunit is configured to determine, as the candidate fault dimension data set, the frequent item sets in the frequent item set range whose occurrence frequencies meet the occurrence frequency requirements.
[0151] The determination unit 203 is specifically configured to determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and a set batch fault determination condition.
[0152] The determination unit 203 includes: a calculation subunit and a comparison subunit;
[0153] The calculation subunit is configured to calculate a failure rate corresponding to the candidate fault dimension data set;
[0154] The comparison subunit is configured to compare the failure rate with a set failure rate baseline value. If the failure rate is greater than or equal to the failure rate baseline value, it is determined that there are batch faults in the candidate fault dimension data set corresponding to the failure rate.
[0155] The comparison subunit is further specifically configured to, when comparing the failure rate with the failure rate baseline value, if the failure rate is less than the failure rate baseline value, it is determined that there are no batch faults in the candidate fault dimension data set corresponding to the failure rate.
[0156] This device embodiment further includes:
[0157] An alarm unit, configured to issue a batch fault alarm when it is determined that the candidate fault dimension data set is a data set of batch faults.
[0158] The embodiment of the device further includes:
[0159] A detection unit for performing misjudgment detection on the data set of the batch faults determined by the determination unit 203.
[0160] The detection unit may include: a pulling subunit and a detection subunit;
[0161] The pulling subunit is used to pull the black box logs in the network environment;
[0162] The detection subunit is used to detect whether there is a misjudgment in the data set of the determined batch faults according to the data in the black box logs.
[0163] The above is a summary description of the embodiment of the device for determining batch faults provided by the present application. The specific process can refer to the description of the embodiment of the method for determining batch faults, which will not be elaborated here.
[0164] Based on the above content, the present application further provides a fault warning system. Please refer to Figure 3 as shown in Figure 3 which is a schematic diagram of the system architecture of an embodiment of the fault warning system provided by the present application. The system includes:
[0165] A data center 301 and a monitoring service management center 302; wherein, the data center 301 is used to collect single-fault information; the monitoring service management center 302 is used to perform configuration dimension expansion on the single-fault information according to the obtained single-fault information and the data in the obtained configuration management database to obtain a data set of single-fault dimensions; and determine a data set of batch faults according to the data set of single-fault dimensions and the set batch fault judgment conditions.
[0166] In this embodiment, the data center 301 can collect single-fault information by deploying a fault monitoring module on the server of the data center, and the monitoring module uses the agent technology to implement the monitoring of single faults. The monitoring service management center 302 can be responsible for deploying agents, configuring agent operation policies and monitoring contents, etc., and can issue batch fault alarms according to the determined data set of batch faults, as well as perform misjudgment detection on batch faults.
[0167] Based on the above content, from the perspective of fault generation, the present application further provides an embodiment of a method for monitoring batch faults, as Figure 4 shown in Figure 4 which is a flowchart of an embodiment of the method for monitoring batch faults provided by the present application. This embodiment of the monitoring method includes:
[0168] Step S401: Collect the single-fault information of the data center through the deployed monitoring module for the data center;
[0169] The purpose of step S401 is to monitor the operation status of service devices in the data center in real time. That is, when a service device has an abnormal operation, the monitoring module deployed in the data center will collect corresponding single-fault information.
[0170] In this embodiment, the data center can be understood as a specific network of devices for global collaboration, used to transfer, accelerate, display, calculate, and store data information on the Internet network infrastructure. Single-fault information refers to the fault information that occurs in a certain independent application, independent server, or independent component, etc. in the data center, that is, it includes at least one of software fault information and hardware fault information.
[0171] The specific implementation process of step S401 is to deploy a monitoring module (agent) for monitoring fault information on the service devices in the data center. The monitoring module can collect the fault information that appears on the service devices. In this embodiment, the monitoring module deployed on the data center service devices can be configured through the monitoring service management center, and the configured monitoring module is deployed in the data center service devices.
[0172] In this embodiment, corresponding monitoring modules can be deployed on all service devices in the data center. Of course, they can also be deployed according to actual monitoring requirements.
[0173] Step S402: Send the collected single-fault information to the monitoring service management center.
[0174] The purpose of step S402 is that the monitoring module will send the single-fault information collected by monitoring to the monitoring service management center for corresponding processing by the monitoring service management center.
[0175] Correspondingly, the present application also provides a monitoring device for batch faults, such as Figure 5 shown Figure 5 is a schematic structural diagram of an embodiment of a monitoring device for batch faults provided by the present application. This embodiment of the monitoring device includes:
[0176] A collection unit 501, configured to collect the single-fault information of the data center through a deployed monitoring module for monitoring the data center; for the specific implementation process of the collection unit 501, reference can be made to the descriptions of the above steps S101 - step S103 and steps S401 - step S402, which will not be elaborated here.
[0177] A sending unit 501, configured to send the collected single - unit fault information to a monitoring service management center. Similarly, the specific implementation process of the sending unit 501 can refer to the descriptions of the above steps S101 - step S103 and steps S401 - step S402, which will not be elaborated here.
[0178] Based on the above, the present application also provides a computer storage medium, which is used to store data generated by a network platform and a program for processing the data generated by the network platform;
[0179] When the program is read and executed, the following steps are performed:
[0180] Obtain single - unit fault information and configuration information used to describe the service devices in the data center;
[0181] According to the configuration information, perform configuration - dimension expansion on the single - unit fault information to obtain a set of single - unit fault - dimension data;
[0182] According to the set of single - unit fault - dimension data and the set batch - fault judgment conditions, determine the set of batch - fault data;
[0183] Or, perform the following steps:
[0184] Collect the single - unit fault information of the data center through a deployed monitoring module for the data center;
[0185] Send the collected single - unit fault information to the monitoring service management center.
[0186] Based on the above, the present application also provides an electronic device, including:
[0187] A processor;
[0188] A memory, configured to store a program for processing data generated by a network platform. When the program is read and executed by the processor, the following steps are performed:
[0189] Obtain single - unit fault information and configuration information used to describe the service devices in the data center;
[0190] According to the configuration information, perform configuration - dimension expansion on the single - unit fault information to obtain a set of single - unit fault - dimension data;
[0191] According to the set of single - unit fault - dimension data and the set batch - fault judgment conditions, determine the set of batch - fault data;
[0192] Or, perform the following steps:
[0193] Collect the single - unit fault information of the data center through a deployed monitoring module for the data center;
[0194] Send the collected single - unit fault information to the monitoring service management center.
[0195] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0196] The memory may include non - permanent memory in the form of computer - readable media, random access memory (RAM), and / or non - volatile memory such as read - only memory (ROM) or flash RAM. Memory is an example of computer - readable media.
[0197] 1. Computer - readable media includes permanent and non - permanent, removable and non - removable media that can store information by any method or technology. The information can be computer - readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase - change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read - only memory (ROM), electrically erasable programmable read - only memory (EEPROM), flash memory or other memory technologies, compact disc read - only memory (CD - ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non - transitory medium that can store information accessible by a computing device. As defined herein, computer - readable media does not include transitory media such as modulated data signals and carrier waves.
[0198] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk memory, CD - ROM, optical memory, etc.) containing computer - usable program code.
[0199] Although the present application is disclosed above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.
Claims
1. A method for determining batch faults, characterized in that, Including: Obtaining single entity failure information and configuration information for describing data center service devices; According to the configuration information, performing configuration dimension expansion on the single entity failure information to obtain a set of single entity failure dimension data; Determining a candidate failure dimension data set based on the correlation analysis between the single entity failure dimension data sets; Determining a data set of batch failures according to the single entity failure dimension data set and a set batch failure judgment condition, including: determining whether the candidate failure dimension data set is a data set of batch failures according to the candidate failure dimension data set and the set batch failure judgment condition.
2. The method for determining batch faults according to claim 1, wherein The obtaining of the single entity failure information includes: Obtaining the single entity failure information of a single entity monitored by the data center.
3. The method for determining batch faults according to claim 1 or 2, characterized in that Also including: Performing a formatting operation on the single entity failure information to obtain a single entity failure work order; The obtaining of the configuration information for describing data center service devices includes: According to the single entity failure work order, obtaining the configuration information for describing data center service devices in the configuration management database, where the configuration management database stores configuration information describing entities in the network environment.
4. The method for determining batch faults according to claim 1, characterized in that, The performing of configuration dimension expansion on the single entity failure information according to the configuration information to obtain a set of single entity failure dimension data includes: Determining configuration dimensions according to configuration items in the configuration information; Constructing the set of single entity failure dimension data according to the configuration dimensions and the single entity failure information.
5. The method for determining batch faults according to claim 1, wherein, The determining of the candidate failure dimension data set based on the correlation analysis between the single entity failure dimension data sets includes: Performing frequent item set mining on the single entity failure dimension data sets; Determining the frequent item sets whose occurrence frequencies meet the occurrence frequency requirements in the frequent item set range as the candidate failure dimension data set.
6. The method for determining batch faults according to claim 1, characterized in that, The determining of whether the candidate failure dimension data set is a data set of batch failures according to the candidate failure dimension data set and a set batch failure judgment condition includes: Calculating the corresponding failure rate within the candidate failure dimension data set; Comparing the failure rate with a set failure rate baseline value, and if the failure rate is greater than or equal to the failure rate baseline value, determining that there are batch failures in the candidate failure dimension data set corresponding to the failure rate.
7. The method for determining batch faults according to claim 6, wherein Also including: When comparing the failure rate with the failure rate baseline value, if the failure rate is less than the failure rate baseline value, determining that there are no batch failures in the candidate failure dimension data set corresponding to the failure rate.
8. The method for determining batch faults according to claim 1, wherein Also including: When determining that the candidate failure dimension data set is a data set of batch failures, issuing a batch failure alarm.
9. The method for determining batch faults according to claim 1 or 8, characterized in that, Also including: Performing misjudgment detection on the determined data set of batch failures.
10. The method for determining batch faults according to claim 9, wherein The performing of misjudgment detection on the determined data set of batch failures includes: Pulling black box logs in the network environment; Detecting whether there is a misjudgment in the determined data set of batch failures according to the data in the black box logs.
11. An apparatus for determining batch faults, characterized in that, Including: An obtaining unit for obtaining single entity failure information and configuration information for describing data center service devices; An expansion unit, configured to expand the single - unit fault information in terms of configuration dimensions according to the configuration information, so as to obtain a set of single - unit fault dimension data; An analysis unit, specifically configured to determine a candidate fault dimension data set according to the correlation analysis among the single - unit fault dimension data sets A determination unit, configured to determine a data set of batch faults according to the single - unit fault dimension data set and a set batch - fault judgment condition, including: determining whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch - fault judgment condition.
12. A monitoring method for batch faults, characterized in that, including: Collect the single - unit fault information of the data center through a deployed monitoring module for monitoring the data center; Send the collected single - unit fault information to the monitoring service management center; The single - unit fault information is used to enable the monitoring service management center to expand the single - unit fault information in terms of configuration dimensions to obtain a set of single - unit fault dimension data; determine a candidate fault dimension data set according to the correlation analysis among the single - unit fault dimension data sets; and determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch - fault judgment condition.
13. The monitoring method for batch faults according to claim 12, characterized in that, The step of collecting the single - unit fault information of the data center through a deployed monitoring module for monitoring the data center includes: Deploy, through the configuration of the monitoring module by the monitoring service management center, the configured monitoring module for collecting the single - unit fault information in the data center.
14. The monitoring method for batch faults according to claim 12, characterized in that, The step of collecting the single - unit fault information of the data center through a deployed monitoring module for monitoring the data center includes: Deploy, on the servers of the data center, the monitoring module for collecting the single - unit fault information of the data center.
15. A monitoring device for batch faults, characterized in that, including: A collection unit, configured to collect the single - unit fault information of the data center through a deployed monitoring module for monitoring the data center; A sending unit, configured to send the collected single - unit fault information to the monitoring service management center; The single - unit fault information is used to enable the monitoring service management center to expand the single - unit fault information in terms of configuration dimensions to obtain a set of single - unit fault dimension data; determine a candidate fault dimension data set according to the correlation analysis among the single - unit fault dimension data sets; and determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch - fault judgment condition.
16. A fault warning system, characterized in that, including: A data center and a monitoring service management center; wherein, the data center is used to collect single-fault information; the monitoring service management center is used to perform configuration dimension expansion on the single-fault information according to the obtained single-fault information and the obtained configuration information for describing the service devices in the data center, to obtain a set of single-fault dimension data; determine a candidate fault dimension data set according to the correlation analysis between the sets of single-fault dimension data; determine a candidate fault dimension data set according to the correlation analysis between the sets of single-fault dimension data; and determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch fault judgment conditions.
17. The fault warning system according to claim 16, wherein Including: A server deployment monitoring module in the data center to monitor the single-fault information in the data center.
18. The fault warning system according to claim 16, wherein Including: The monitoring service management center issues a batch fault alarm according to the determined data set of batch faults.
19. A computer storage medium for storing data generated by a network platform and a program for processing the data generated by the network platform; When the program is read and executed, the following steps are performed: Obtain single-fault information and configuration information for describing the service devices in the data center; Perform configuration dimension expansion on the single-fault information according to the configuration information to obtain a set of single-fault dimension data; Determine a candidate fault dimension data set according to the correlation analysis between the sets of single-fault dimension data; Determine the data set of batch faults according to the set of single-fault dimension data and the set batch fault judgment conditions, including: determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch fault judgment conditions; Or, perform the following steps: Collect the single-fault information of the data center through a deployed monitoring module for monitoring the data center; Send the collected single-fault information to the monitoring service management center; the single-fault information is used to enable the monitoring service management center to perform configuration dimension expansion on the single-fault information to obtain a set of single-fault dimension data; determine a candidate fault dimension data set according to the correlation analysis between the sets of single-fault dimension data; and determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch fault judgment conditions.
20. An electronic device, including: A processor; A memory for storing a program for processing data generated by a network platform. When the program is read and executed by the processor, the following steps are performed: Obtain single-fault information and configuration information for describing the service devices in the data center; Perform configuration dimension expansion on the single-fault information according to the configuration information to obtain a set of single-fault dimension data; Determine a candidate fault dimension data set according to the correlation analysis between the sets of single-fault dimension data; Determine the data set of batch faults according to the single-fault dimension data set and the set batch fault judgment conditions, including: determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch fault judgment conditions; Alternatively, perform the following steps: Collect the single-fault information of the data center through the deployed monitoring module for monitoring the data center; Send the collected single-fault information to the monitoring service management center; the single-fault information is used to enable the monitoring service management center to perform configuration dimension expansion on the single-fault information to obtain a single-fault dimension data set; determine a candidate fault dimension data set according to the correlation analysis between the single-fault dimension data sets; determine whether the candidate fault dimension data set is a data set of batch faults according to the candidate fault dimension data set and the set batch fault judgment conditions.
Citation Information
Patent Citations
Dynamic correlation fault mining method based on business paths and frequency matrixes
CN107579844A
Batch log abnormal data alarm method and device
CN108737170A