Cluster data anomaly detection method and electronic device
By determining the mapping relationship and grouping information between application clusters and storage resource pools in the cloud platform, aggregating key storage indicators, and generating threshold strategies for anomaly detection, the problem of difficulty in identifying anomalies due to performance degradation of storage systems is solved, and abnormal applications or nodes can be quickly located.
Patent Information
- Application Number
- CN202211151012.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-09-21
AI Technical Summary
In cloud or virtualization platforms, storage systems may experience performance degradation due to excessive load leading to resource contention, but it can be difficult to identify which application cluster or data cluster is causing the anomaly.
By determining the mapping relationship between application clusters and storage resource pools, grouping information is generated, and key storage indicators are aggregated based on the grouping information to obtain threshold strategies for anomaly detection.
In big data cluster scenarios, it can quickly locate the target application or node that causes storage system anomalies, improving the speed of problem identification and diagnosis.
Smart Images

Figure CN115509853B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a cluster data anomaly detection method and an electronic device. BACKGROUND
[0002] In the operation and maintenance activities of a storage system in a cloud platform or a virtualization platform, resource competition of the storage system may occur due to excessive load generated by some applications or systems, and thus the performance of the storage platform is reduced. For example, for a big data cluster, the load fluctuation of a single node is limited, but the cumulative load of the cluster is excessive, and observation is more difficult, so that when the performance of the storage platform is abnormal, it is not possible to identify which application cluster or data cluster causes the abnormality. SUMMARY
[0003] Therefore, the embodiments of the present application aim to provide a cluster data anomaly detection method and an electronic device.
[0004] To achieve the above object, the technical scheme of the present application is as follows:
[0005] According to an aspect of the present application, a cluster data anomaly detection method is provided, comprising:
[0006] determining a mapping relationship between application clusters and storage resource pools based on configuration information of the application clusters;
[0007] generating grouping information of different application clusters in different storage resource pools based on the mapping relationship;
[0008] aggregating key storage indicators of different application clusters based on the grouping information, to obtain threshold strategies of different application clusters in different storage resource pools;
[0009] performing anomaly detection on the performance of a target object in a corresponding time window based on the threshold strategies, the target object including at least one of each application cluster or each application server in each application cluster.
[0010] In the above scheme, the anomaly detection on the performance of the target object in the corresponding time window based on the threshold strategies includes one of the following:
[0011] performing anomaly detection on specific storage indicators of each application cluster in a historical time window based on the threshold strategies, the specific storage indicators including at least one of input / output bandwidth flow and the number of read / write operations per second between each application cluster and nodes in each cluster;
[0012] perform anomaly detection on node correlation storage indicators of each application cluster in a historical time window in a corresponding time window based on the threshold strategy; the node correlation storage indicators include at least one of input / output bandwidth traffic between nodes in each cluster, the number of traffic nodes between nodes in each cluster, and the number of links between nodes in each cluster.
[0013] In the above solution, the anomaly detection on the node correlation storage indicators of each application cluster in the historical time window in the corresponding time window based on the threshold strategy comprises:
[0014] processing data of the node correlation storage indicators of the nodes in each cluster in the historical time window to obtain indicator data in each sliding window;
[0015] obtaining an indicator correlation coefficient of different application clusters based on the nodes in each cluster and the indicator data;
[0016] determining an application cluster corresponding to an indicator correlation coefficient satisfying a target condition as an abnormal application cluster based on the threshold strategy, the target condition indicating that the indicator correlation coefficient between the nodes in the cluster is weak.
[0017] In the above solution, the method further comprises:
[0018] outputting an alarm information if the detection result indicates that at least one target object has performance anomaly.
[0019] In the above solution, the detection result indicating that at least one target object has performance anomaly comprises:
[0020] obtaining feature aggregation data of each target object in a historical time window;
[0021] determining that the at least one target object has performance anomaly if a target value of the feature aggregation data exceeds a boundary threshold value;
[0022] The target value is used to reflect the data characteristics of the feature aggregation data.
[0023] In the above solution, the aggregation of the key storage indicators of different application clusters based on the grouping information to obtain the threshold strategy of different application clusters in different storage resource pools comprises:
[0024] obtaining historical load data of different application clusters in different storage resource pools based on the grouping information;
[0025] performing aggregation calculation on the historical load data to obtain a boundary threshold value of different application clusters in different storage resource pools;
[0026] generating the threshold strategy based on the boundary threshold value.
[0027] In the above solution, before the performance of the target object is detected for abnormality in the corresponding time window based on the threshold strategy, the method further comprises:
[0028] obtaining a start time of the performance anomaly of the storage resource pool;
[0029] detecting the historical data of each application cluster in the time window corresponding to the start time based on the boundary threshold in the threshold strategy; or detecting the historical data of each application cluster in the time window corresponding to a period of time before the start time.
[0030] In the above solution, the method further comprises:
[0031] performing performance anomaly detection on each application cluster according to the sorting of the index correlation coefficients of the different application clusters;
[0032] presenting the monitoring index data in the application cluster range corresponding to each index correlation coefficient according to the sorting result.
[0033] In the above solution, the method further comprises:
[0034] outputting a change curve graph of the index correlation coefficient.
[0035] According to another aspect of the present application, an electronic device is provided, wherein the electronic device comprises:
[0036] a determination unit configured to determine a mapping relationship between an application cluster and a storage resource pool based on configuration information of the application cluster;
[0037] a generation unit configured to generate grouping information of different application clusters in different storage resource pools based on the mapping relationship;
[0038] an aggregation unit configured to aggregate key storage indexes of different application clusters based on the grouping information, and determine a threshold strategy of different application clusters in different storage resource pools;
[0039] a detection unit configured to detect the performance of a target object for abnormality in a corresponding time window based on the threshold strategy, the target object comprising at least one of each application cluster or each application server in each application cluster.
[0040] The cluster data anomaly detection method and the electronic device provided by the application generate grouping information of different application clusters in different storage resource pools by applying the mapping relationship between the application clusters and the storage resource pools; the key storage indicators of different application clusters are aggregated based on the grouping information, the threshold strategy of different application clusters in different storage resource pools is determined, and the performance of a target object is detected for anomaly in a corresponding time window based on the threshold strategy. In this way, in a big data cluster scenario, when the performance of a single cloud hard disk fluctuates slightly and the storage platform is seriously affected by the competition of cluster resources, the target application or node causing the anomaly of the storage system can be quickly located. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 The flow implementation schematic of the cluster data anomaly detection method in the application Figure One ;
[0042] Figure 2 The flow implementation schematic of the threshold strategy generation method in the application
[0043] Figure 3 The flow implementation schematic of the cluster data anomaly detection method in the application Figure Two ;
[0044] Figure 4 The structural composition schematic of the electronic device in the application Figure One ;
[0045] Figure 5 The structural composition schematic of the electronic device in the application Figure Two . DETAILED DESCRIPTION
[0046] To make the objectives, technical solutions and advantages of the application clearer, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application. The embodiments in the application and the features in the embodiments can be combined with each other in a non-conflicting manner. The steps shown in the flowchart of the drawings can be executed in a computer system such as a group of computer executable instructions. Moreover, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0047] As described above, in a big data cluster scenario, resource competition may occur in a storage system due to excessive load generated by some applications or systems, resulting in a performance decline of the storage platform. However, the application that causes the abnormality cannot be identified when the performance of the storage platform is abnormal. According to the scheme provided in the application, the key storage indicators of different application clusters are aggregated based on the grouping information of different application clusters in different storage resource pools, and the threshold strategy of different application clusters in different storage resource pools is obtained. Based on the threshold strategy, the performance of the target object is detected in the corresponding time window. In this way, when the performance of a single cloud hard disk fluctuates slightly and the storage platform is seriously affected by cluster resource competition, the target application or node that causes the abnormality of the storage system can be quickly located.
[0048] The technical scheme of the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0049] Figure 1 Flow implementation diagram of the cluster data anomaly detection method in the application Figure One The method can be applied to an electronic device, which can be a server and a client device. As shown in Figure 1 The method includes the following steps.
[0050] In step 101, the mapping relationship between the application cluster and the storage resource pool is determined based on the configuration information of the application cluster.
[0051] Here, the electronic device can obtain the configuration information of the application cluster through a cloud platform or virtualization platform interface (or cloud platform or virtualization platform database), a storage system interface, and a monitoring database. Based on the configuration information, the mapping relationship between the application cluster and the storage resource pool can be determined.
[0052] Specifically, the electronic device can determine the first correspondence relationship between the storage resource pool and the delivery disk (or delivery volume) through the storage system interface, and determine the second correspondence relationship between the delivery disk (or delivery volume) and the physical machine or virtual machine through the cloud platform or virtualization platform interface (or database), that is, the delivery disk (or delivery volume) is deployed on which physical machine or virtual machine. Based on the first correspondence relationship and the second correspondence relationship, the electronic device can determine the mapping relationship between the cloud hard disk in the application cluster and the storage resource pool.
[0053] Since multiple application clusters can run on one platform, the electronic device can also obtain monitoring data through the monitoring database, and determine the third correspondence relationship between different applications, different databases, or different data platform clusters and the cloud hard disk based on the monitoring data.
[0054] The electronic device can determine a mapping relationship between the application clusters and the storage resource pools based on the first correspondence relationship, the second correspondence relationship and the third correspondence relationship.
[0055] In step 102, grouping information of different application clusters in different storage resource pools is generated based on the mapping relationship.
[0056] In this application, the electronic device can group different application clusters in different storage resource pools based on the mapping relationship between different application clusters and storage resource pools, and generate grouping information of different application clusters in different storage resource pools.
[0057] Here, the grouping information includes but is not limited to first mapping information of each application cluster and each cluster node (such as Elasticsearch (ES), an ES node is a Lucene-based search server), second mapping information of each cluster node and disk, and third mapping information of disk and storage resource pool.
[0058] Based on the first mapping information, the electronic device can determine which node each application runs on, based on the second mapping information, it can determine which disk / volume each node has, and based on the third mapping information, it can determine which storage resource pool each disk / volume is delivered by.
[0059] In this application, the electronic device can also determine the current distribution state information and the historical distribution state information of each node in each application cluster according to the grouping information. Based on the historical distribution state information and the current distribution state information, it can be determined which nodes are expanded and which nodes are originally present.
[0060] In step 103, key storage indicators of different application clusters are aggregated based on the grouping information to obtain threshold strategies of different application clusters in different storage resource pools.
[0061] In this application, the key storage indicators include but are not limited to input / output (I / O, Input / Output) bandwidth and the number of read / write operations per second (IOPS, Input / Output Operations Per Second).
[0062] In one implementation, the electronic device can obtain historical load data of different application clusters in different storage resource pools based on the grouping information of different application clusters in different storage resource pools. By aggregating and calculating the historical load data, the boundary threshold of different application clusters in different storage resource pools can be obtained. The threshold strategy is generated based on the boundary threshold.
[0063] For example, the electronic device can obtain the I / O throughput of each node in each application cluster, and then accumulate the I / O throughput of each node in each application cluster to obtain the total I / O throughput of each application cluster in different storage resource pools, and take the total I / O throughput as the boundary threshold of different application clusters in different storage resource pools, so as to realize the aggregated calculation of the I / O throughput of different application clusters in different storage resource pools.
[0064] In step 104, the performance of the target object is detected for abnormality in the corresponding time window based on the threshold strategy, and the target object includes at least one of each application cluster or each application server in each application cluster.
[0065] In one implementation, the electronic device can detect the specific storage indicators of each application cluster in the historical time window for abnormality in the corresponding time window based on the threshold strategy.
[0066] Here, the specific storage indicators include at least one of the I / O bandwidth flow and the IOPS between each application cluster and each node in the cluster. The electronic device can obtain the I / O bandwidth and the IOPS between each application cluster and each node in the cluster through the gateway. Then the specific storage indicators are detected for abnormality based on the percentile or based on the standard deviation method.
[0067] In another implementation, the electronic device can also detect the node-related storage indicators of each application cluster in the historical time window for abnormality in the corresponding time window based on the threshold strategy.
[0068] Here, the node-related storage indicators include at least one of the input / output bandwidth flow between each node in the cluster, the number of flow nodes between each node in the cluster, and the number of links between each node in the cluster.
[0069] Here, the flow abnormality detection between the nodes in the cluster can be understood as the detection of the east-west flow between the nodes in the cluster.
[0070] For example, a big data cluster such as Hadoop or Elasticsearch runs on a cloud platform or a virtualization platform. When data rebalancing is performed within the cluster, the performance of any single disk may fluctuate within the normal monitoring range, but the performance of the storage platform may be abnormally increased.
[0071] In this application, if the storage resource pool (or storage platform) is abnormal, and the detection result indicates that there is an abnormality between the nodes in the cluster in the corresponding time window, or the threshold detection is manually labeled as abnormal, it is recorded as a suspected risk, and then an alarm or an abnormality report can be output. Thus, engineers can quickly locate the abnormal point and speed up the diagnosis speed of the abnormal point.
[0072] The application groups different application clusters by applying the mapping relationship of the cluster and the storage resource pool, aggregates the key storage indicators of different application clusters according to the grouping information of different application clusters in different storage resource pools, and obtains the threshold strategy of different application clusters in different storage resource pools; and based on the threshold strategy, the performance of the target object is detected in the corresponding time window. In this way, in the big data cluster scenario, when the performance of a single cloud hard disk fluctuates slightly and the storage platform is seriously affected by the competition of cluster resources, the problem identification speed can be accelerated, and the target application or node causing the abnormality of the storage system can be quickly located.
[0073] In the application, when the electronic device performs abnormality detection on the node correlation storage indicators of each application cluster in the historical time window based on the threshold strategy in the corresponding time window, the data of the node correlation storage indicators of the nodes in each cluster in the historical time window can also be processed to obtain the indicator data in each sliding window; then the indicator correlation coefficients of different application clusters are obtained based on the nodes in each cluster and the indicator data; the indicator correlation coefficient graph is established based on the indicator correlation coefficients; and the application cluster corresponding to the indicator correlation coefficient that meets the target condition in the indicator correlation coefficient graph is determined as an abnormal application cluster based on the threshold strategy.
[0074] Here, the target condition represents that the indicator correlation coefficients between the nodes in the cluster are weak.
[0075] Here, the indicator correlation coefficient can be the correlation coefficient between the same indicators or related indicators of multiple nodes. The indicator correlation coefficient can include two levels:
[0076] 1. Different nodes, same indicators;
[0077] 2. Same or different nodes, related indicators;
[0078] It should be emphasized that the indicator correlation coefficient represents the relationship between multiple indicators, rather than the correlation coefficient of one indicator.
[0079] In the application, the abnormality detection is performed on the traffic correlation between nodes, and the process is specifically as follows:
[0080] a) Smooth processing is performed on the observation data (such as I / O bandwidth, IOPS, etc.) of each node, such as resampling, exponential smoothing, kernel function smoothing, or convolution smoothing.
[0081] b) The data after the smoothing processing is subjected to sliding window processing to obtain each sliding window. As an optional step, the mean of the obtained sliding window is obtained.
[0082] c) Construct each node and index data into a two-dimensional data frame, one data frame for each index, respectively for the node and the node value under different sliding windows.
[0083] d) Perform correlation analysis on the matrix of c), for example using Pearson correlation coefficient. Further obtain the correlation coefficient of the correlation index between different nodes. Obtain the correlation coefficient of the load of different nodes, and use the percentile record correlation coefficient anomaly detection boundary and value. And according to the relationship between different nodes, construct a correlation graph.
[0084] e) Using dynamic threshold, remove the nodes with weak correlation from the correlation graph, and only keep the nodes with high correlation. (The effect of this step is to optimize the computing power).
[0085] In this application, the electronic device can also sort the performance anomaly detection of each application cluster according to the index correlation coefficient of different application clusters; and present the monitoring index data in the application cluster range corresponding to the index correlation coefficient according to the sorting result.
[0086] In this application, the electronic device can also output the change curve graph of the index correlation coefficient.
[0087] In this application, the electronic device can also obtain feature aggregation data of each target object in a historical time window; if the target value of the feature aggregation data exceeds the boundary threshold, it is determined that the at least one target object has performance anomaly; the target value is used to reflect the data characteristics of the feature aggregation data.
[0088] For example, the target value can be the peak value, median value, mean value or valley value of the feature aggregation data. It can also be data with geometric characteristics, such as slope, curvature, radius, etc.
[0089] In this application, if a storage platform anomaly occurs, the electronic device triggers to perform cluster performance anomaly detection. Here, the electronic device can also obtain the start time of performance anomaly of the storage resource pool before performing anomaly detection on the performance of the target object in the corresponding time window based on the threshold strategy; perform anomaly detection on the historical data of each application cluster in the time window corresponding to the start time based on the boundary threshold in the threshold strategy; or perform anomaly detection on the historical data of each application cluster in the time window corresponding to a period of time before the start time. If the detection result indicates that the current data index correlation of the relevant node breaks through the dynamic threshold, it is marked as abnormal. And at this time, the situation of excessive node traffic may occur, or the rebalance situation may occur.
[0090] In this application, if the detection result indicates that at least one target object has performance anomaly, the electronic device outputs an alarm information.
[0091] The abnormality detection index here includes performance index after aggregation, and also includes the number of nodes with abnormality appearing simultaneously, the number of links with abnormal flow between nodes (rebalance scenario), storage platform cache, storage platform CPU and the like.
[0092] If the cluster consumption cloud disk abnormality and the storage platform abnormality occur at the same time or within a limited time window, then abnormality a and abnormality b == true, indicating that the cluster abnormality may cause platform abnormality, and then an alarm or a report is triggered.
[0093] The cluster data abnormality detection method provided in the application obtains the storage system resource pool used by the delivery disk through a cloud platform interface or a database and a storage system interface. Through a monitoring system, different application, database or big data platform cluster boundaries are identified. Different clusters are grouped on different storage resource pools, and key storage index aggregation calculation is performed based on the grouping information. The abnormality detection is performed on the aggregation calculation results within a period of time, or the threshold range is manually labeled. If the storage platform has an abnormality, and there is a combined abnormality detection or manually labeled threshold detection abnormality within the time window, it is recorded as a suspected risk for alarm or output of an abnormality report. When the single cloud disk performance fluctuation is not large, and the storage platform is seriously affected by the cluster resource competition, the problem identification is accelerated, and the end-to-end problem solving efficiency is improved. Moreover, the method provided in the application has low algorithm complexity and high reliability.
[0094] Figure 2 The generation method flowchart of the threshold strategy in the application is shown as Figure 2 The method comprises the following steps:
[0095] Step 201, obtaining the mapping relationship between the cloud disk and the resource pool through a cloud platform or a virtualization platform and a storage system interface;
[0096] Step 202, identifying different application, different database or different big data platform cluster boundaries through a configuration management database (CMDB), a cluster configuration database or a monitoring database;
[0097] Step 203, grouping different application clusters on different storage resource pools to generate grouping information of the cloud disk consumption of different resource pools of different application clusters;
[0098] Here, the grouping information includes the current distribution state information and the distribution history record information of the cloud disk of different application clusters.
[0099] Step 204, aggregating the historical load data under different storage resource pools based on the grouping information;
[0100] Step 205, taking the historical data extreme value in the aggregation result as a dynamic threshold boundary;
[0101] Step 206, determining the dynamic threshold of different clusters in different storage resource pools according to the dynamic threshold boundary.
[0102] Here, the dynamic threshold can also be referred to as a threshold strategy.
[0103] Here, the dynamic threshold range can also be manually annotated according to the aggregation result.
[0104] In this application, the electronic device typically includes the following aggregation indicators during the aggregation process:
[0105] I / O bandwidth and read-write IOPS.
[0106] Figure 3 The method flow for detecting data anomalies in clusters in this application is shown in the following figure: Figure Two As shown in the figure, the method includes: Figure 3
[0107] Step 301, determining that a performance anomaly exists in the storage platform;
[0108] Here, the storage platform can also be referred to as a storage resource pool.
[0109] Step 302, identifying the start time of the storage platform anomaly based on the CMDB, cluster configuration database, or monitoring database;
[0110] Step 303, obtaining historical grouping information of cloud disk consumption of different storage resource pools of different application clusters;
[0111] Here, the historical grouping information of cloud disk consumption of different storage resource pools of different application clusters can be obtained based on the CMDB, cluster configuration database, or monitoring database.
[0112] Step 304, obtaining each feature aggregation data.
[0113] Here, the aggregation data includes historical aggregation data of cloud disks of different application clusters within a period of time before the start time of the platform anomaly.
[0114] Step 305, performing dynamic threshold detection according to the threshold strategy;
[0115] Here, not only the anomaly occurrence time needs to be detected, but also the historical data within a time window before the anomaly occurrence time needs to be detected.
[0116] Step 306, if the storage platform anomaly and the cluster anomaly occur at the same time, or the cluster anomaly occurs within a time window before the platform anomaly time, an alarm is triggered;
[0117] Step 307: Output an exception report or an alarm message indicating an exception.
[0118] Here, the alarm information for abnormal messages can be reports or messages connected to other business systems. The output format of the alarm information is not limited, as long as it can indicate that the storage platform or cluster is abnormal.
[0119] Figure 4 This is a schematic diagram of the structural composition of the electronic device in this application. Figure One ,like Figure 4 As shown, the electronic device includes:
[0120] The determining unit 401 is used to determine the mapping relationship between the application cluster and the storage resource pool based on the configuration information of the application cluster;
[0121] The generation unit 402 is used to generate grouping information of different application clusters in different storage resource pools based on the mapping relationship;
[0122] Aggregation unit 403 is used to aggregate key storage indicators of different application clusters based on the grouping information, and determine the threshold strategy of different application clusters in different storage resource pools.
[0123] The detection unit 404 is used to perform anomaly detection on the performance of the target object within a corresponding time window based on the threshold strategy. The target object includes at least one of each application cluster or each application server in each application cluster.
[0124] In the preferred embodiment, the detection unit 404 is specifically used to perform anomaly detection on specific storage metrics of each application cluster within a historical time window based on the threshold strategy within the corresponding time window. The specific storage metrics include at least one of the input / output bandwidth traffic between each application cluster and nodes within each cluster, and the number of read / write operations per second; and / or
[0125] Based on the threshold strategy, anomaly detection is performed on the node correlation storage indicators of each application cluster within the corresponding time window. The node correlation storage indicators include at least one of the following: input / output bandwidth traffic between nodes in each cluster, number of traffic nodes between nodes in each cluster, and number of links between nodes in each cluster.
[0126] In a preferred embodiment, the electronic device further includes a processing unit 405 and a setup unit 406;
[0127] The processing unit 405 is used to process the data of the node correlation storage index of each node in the cluster within the historical time window to obtain the index data in each sliding window.
[0128] The determining unit 401 is further configured to obtain an index correlation coefficient of different application clusters based on the nodes in each cluster and the index data.
[0129] The establishing unit 406 is configured to establish an index correlation coefficient graph based on the index correlation coefficients.
[0130] The determining unit 401 is further configured to determine, based on the threshold strategy, an application cluster corresponding to an index correlation coefficient that meets a target condition in the index correlation coefficient graph as an abnormal application cluster, where the target condition indicates that the index correlation coefficients between the nodes in the cluster are weak.
[0131] In a preferred implementation, the output unit 407 is configured to output an alarm information if the detection result indicates that the at least one target object has performance anomaly.
[0132] In a preferred implementation, the electronic device further includes:
[0133] The obtaining unit 408 is configured to obtain feature aggregation data of each target object in a historical time window.
[0134] The determining unit 401 is further configured to determine that the at least one target object has performance anomaly if a target value of the feature aggregation data exceeds a boundary threshold value.
[0135] The target value is used to reflect a data characteristic of the feature aggregation data.
[0136] In a preferred implementation, the determining unit 401 is further configured to obtain historical load data of different application clusters in different storage resource pools based on the grouping information.
[0137] The aggregating unit 403 is specifically configured to perform aggregation calculation on the historical load data to obtain a boundary threshold value of different application clusters in different storage resource pools.
[0138] The generating unit 402 is configured to generate the threshold strategy based on the boundary threshold value.
[0139] In a preferred implementation, the obtaining unit 408 is further configured to obtain a start time of performance anomaly of the storage resource pool.
[0140] The detecting unit 404 is specifically configured to perform anomaly detection on historical data of each application cluster in a time window corresponding to the start time based on a boundary threshold value in the threshold strategy, or perform anomaly detection on historical data of each application cluster in a time window corresponding to a period of time before the start time.
[0141] In a preferred implementation, the electronic device further includes:
[0142] The sorting unit 409 is configured to sort the application clusters according to the index correlation coefficients of the different application clusters.
[0143] The display unit 410 is configured to present the monitoring index data in the application cluster range corresponding to each index correlation coefficient according to the sorting result.
[0144] In a preferred implementation, the output unit 407 is further configured to output a change curve of the index correlation coefficient.
[0145] The change curve facilitates the operation and maintenance engineers to accelerate decision-making.
[0146] By the scheme provided in the present application, the key storage indexes of different application clusters are aggregated based on the grouping information of different application clusters in different storage resource pools, and the threshold strategy of different application clusters in different storage resource pools is obtained; and the performance of a target object is detected for abnormality based on the threshold strategy within a corresponding time window. In this way, when the performance of a single cloud hard disk fluctuates slightly, and the storage platform is seriously affected by the resource competition of the cluster and has a serious performance risk, the target application or node causing the abnormality of the storage system can be quickly located.
[0147] It should be noted that the electronic device provided in the above embodiments is only used as an example to illustrate the division of the above program modules, and in actual applications, the above processing can be completed by different program modules according to needs, that is, the internal structure of the electronic device is divided into different program modules to complete all or part of the above processing. In addition, the electronic device provided in the above embodiments and the cluster data abnormality detection method provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.
[0148] The electronic device provided in the present application also includes a processor and a memory for storing a computer program capable of running on the processor,
[0149] When the processor runs the computer program, it performs any of the method steps of the above processing method.
[0150] Figure 5 The electronic device in the present application is a structural composition of the electronic device Figure Two The electronic device 500 can be a mobile phone, a computer, a digital broadcast terminal, an information transceiver device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like. Figure 5The electronic device 500 shown includes at least one processor 501, a memory 502, at least one network interface 504, and a user interface 503. The various components in the electronic device 500 are coupled together by a bus system 505. It is understood that the bus system 505 is used for communicating data between the components. The bus system 505 includes a data bus, a power bus, a control bus, and a state signal bus for communicating data, power, control, and status information, respectively. However, for the sake of clarity, only the data bus is shown in Figure 5 FIG. 5.
[0151] The user interface 503 can include a display, a keyboard, a mouse, a trackball, a click wheel, a keypad, a button, a touchpad, a touchscreen, or the like.
[0152] It can be understood that the memory 502 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory 502 described in the embodiments of the present application is intended to include, but is not limited to, these and any other suitable type of memory.
[0153] The memory 502 in the embodiments of the present application is used to store various types of data to support the operation of the electronic device 500. Examples of these data include: any computer programs used for operation on the electronic device 500, such as an operating system 5021 and an application program 5022; contact data; phonebook data; messages; pictures; audio; and the like. The operating system 5021 contains various system programs, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks. The application program 5022 can contain various application programs, such as a media player (Media Player), a browser (Browser), and the like, for implementing various application services. The program implementing the method of the embodiments of the present application can be contained in the application program 5022.
[0154] The method disclosed in the embodiments of the present application can be applied in the processor 501 or implemented by the processor 501. The processor 501 can be an integrated circuit chip having a processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit or the instruction in the form of software in the processor 501. The processor 501 described above can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, and the like. The processor 501 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, and the like. In combination with the steps of the method disclosed in the embodiments of the present application, the above-mentioned method can be directly embodied as a hardware coding processor to execute, or be executed by a combination of hardware and software modules in the coding processor. The software module can be located in the storage medium, and the storage medium is located in the memory 502. The processor 501 reads the information in the memory 502 and combines the hardware to complete the steps of the above-mentioned method.
[0155] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, micro controllers (MCUs), microprocessors (Microprocessors), or other electronic elements for executing the aforementioned methods.
[0156] In an exemplary embodiment, the embodiments of the present application further provide a computer readable storage medium, for example, the memory 502 including a computer program, which can be executed by the processor 501 of the electronic device 500 to complete the steps of the aforementioned methods. The computer readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.; or can be various devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0157] A computer readable storage medium having a computer program stored thereon, which, when executed by a processor, performs any of the steps of the above processing methods.
[0158] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. The above described device embodiments are merely illustrative, for example, the division of units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling, or direct coupling, or communication connection between the various components can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0159] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on a plurality of network units; some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0160] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.
[0161] The features disclosed in the several product embodiments provided by the present application can be combined arbitrarily without conflict to obtain new product embodiments.
[0162] The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0163] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A cluster data anomaly detection method, comprising: determining a mapping relationship between application clusters and storage resource pools based on configuration information of the application clusters; generating grouping information of different application clusters in different storage resource pools based on the mapping relationship; aggregating key storage indicators of different application clusters based on the grouping information to determine threshold strategies of different application clusters in different storage resource pools; performing anomaly detection on performance of a target object in a corresponding time window based on the threshold strategies, the target object including at least one of each application cluster or each application server in each application cluster; the aggregating of the key storage indicators of different application clusters based on the grouping information to obtain the threshold strategies of different application clusters in different storage resource pools comprises: obtaining historical load data of different application clusters in different storage resource pools based on the grouping information; performing aggregation calculation on the historical load data to obtain boundary thresholds of different application clusters in different storage resource pools; and generating the threshold strategies based on the boundary thresholds.
2. The method of claim 1, wherein, the performing of the anomaly detection on the performance of the target object in the corresponding time window based on the threshold strategies comprises one of: performing anomaly detection on specific storage indicators of each application cluster in a historical time window in the corresponding time window based on the threshold strategies, the specific storage indicators including at least one of input / output bandwidth traffic and the number of read / write operations per second between each application cluster and nodes in each cluster; and performing anomaly detection on node-related storage indicators of each application cluster in the historical time window in the corresponding time window based on the threshold strategies, the node-related storage indicators including at least one of input / output bandwidth traffic between nodes in each cluster, the number of traffic nodes between nodes in each cluster, and the number of links between nodes in each cluster.
3. The method of claim 2, wherein, the performing of the anomaly detection on the node-related storage indicators of each application cluster in the historical time window in the corresponding time window based on the threshold strategies comprises: processing data of the node-related storage indicators of nodes in each cluster in the historical time window to obtain indicator data in each sliding window; obtaining indicator correlation coefficients of different application clusters based on the nodes in each cluster and the indicator data; and determining, based on the threshold strategies, an application cluster corresponding to an indicator correlation coefficient that meets a target condition as an abnormal application cluster, the target condition representing that the indicator correlation coefficients between nodes in the cluster are weak.
4. The method of claim 1, wherein, The method further comprises: if a detection result represents that at least one target object has performance anomaly, outputting alarm information.
5. The method of claim 4, wherein, the detection result representing that at least one target object has performance anomaly comprises: obtaining feature aggregation data of each target object in a historical time window; if a target value of the feature aggregation data exceeds a boundary threshold, determining that the at least one target object has performance anomaly; the target value is used to reflect data characteristics of the feature aggregation data.
6. The method of claim 1, wherein, Before the performing of the anomaly detection on the performance of the target object in the corresponding time window based on the threshold strategies, the method further comprises: obtaining a start time of performance anomaly of the storage resource pool; and The abnormality detection is performed on historical data of each application cluster in a time window corresponding to the start time based on a boundary threshold in the threshold strategy; or the abnormality detection is performed on historical data of each application cluster in a time window corresponding to a time before the start time.
7. The method of claim 3, wherein, The method further includes: sequencing performance abnormality detection of each application cluster according to the index correlation coefficients of the different application clusters; monitoring index data in the application cluster range corresponding to each index correlation coefficient is presented according to the sequencing result.
8. The method of claim 3, wherein, The method further includes: a change curve of the index correlation coefficient is output.
9. An electronic device, comprising: The electronic device includes: a determination unit configured to determine a mapping relationship between an application cluster and a storage resource pool based on configuration information of the application cluster; a generation unit configured to generate grouping information of different application clusters in different storage resource pools based on the mapping relationship; an aggregation unit configured to aggregate key storage indexes of different application clusters based on the grouping information, and determine a threshold strategy of different application clusters in different storage resource pools; a detection unit configured to perform abnormality detection on performance of a target object in a corresponding time window based on the threshold strategy, the target object including at least one of each application cluster or each application server in each application cluster; the determination unit is further configured to obtain historical load data of different application clusters in different storage resource pools based on the grouping information; the aggregation unit is further configured to perform aggregation calculation on the historical load data to obtain a boundary threshold of different application clusters in different storage resource pools; the generation unit is further configured to generate the threshold strategy based on the boundary threshold.
Citation Information
Patent Citations
Method and device for evaluating cluster performance
CN108874640A
Object storage service management method and electronic equipment
CN111212111A