Operation risk monitoring method, device and equipment for distributed service cluster
By grouping and calculating the abnormal distribution ratio of various types of operating data of distributed service clusters, the problem that traditional methods are difficult to accurately predict risks in complex dynamic environments is solved, and accurate risk prediction of distributed service clusters is achieved.
Patent Information
- Application Number
- CN202510446598.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional service node risk prediction methods are difficult to accurately predict potential risks in complex dynamic environments, and are limited by changes in the number of service nodes.
By obtaining various types of operation data in the distributed service cluster, grouping service nodes, determining the proportion of abnormal distribution of each packet, calculating the index value and abnormal score, and finally predicting the operation risk based on the total abnormal score.
It realizes the quantification of abnormal situations in distributed service clusters, and can accurately predict operational risks when the number of service nodes changes dynamically, improving the adaptability and accuracy of abnormal prediction in complex dynamic environments.
Smart Images

Figure CN120223574A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method, apparatus, and device for monitoring the operation risks of a distributed service cluster. Background Art
[0002] In modern distributed systems and cloud computing environments, the health status of service nodes directly affects the stability and performance of the system. Traditional service node risk prediction methods mainly rely on threshold monitoring and rule engines, and these methods often have difficulty accurately predicting potential risks when facing complex dynamic environments. Summary of the Invention
[0003] To solve the problems existing in the prior art, the embodiments of this specification provide a method, apparatus, and device for monitoring the operation risks of a distributed service cluster, which obtain multiple types of operation data of each service node in the distributed service cluster, then group each service node according to each type of operation data to obtain multiple groups under each type of operation data, each group includes several service nodes, then determine the trained abnormal distribution ratio corresponding to the operation data characteristics of each group, calculate the index value corresponding to each group according to the abnormal distribution ratio, then sum the index values of each group corresponding to each type of operation data respectively to obtain the abnormal score corresponding to the corresponding data of each category, and finally sum the abnormal scores of all categories to obtain the total abnormal score, and predict the operation risk of the distributed service cluster according to the total abnormal score, realizing the quantification of the abnormal situation of the distributed service cluster. Even if the number of computing nodes in the distributed service cluster changes dynamically, the operation risk of the distributed service cluster can be accurately predicted.
[0004] The specific technical solutions of the embodiments of this specification are as follows:
[0005] On the one hand, the embodiments of this specification provide a method for monitoring the operation risks of a distributed service cluster, and the method includes:
[0006] Obtain the operation data of multiple data types of each service node in the distributed service cluster;
[0007] Group each service node according to the operation data of each data type to obtain multiple groups under each data type, and each group includes several service nodes;
[0008] Determine the trained abnormal distribution ratio corresponding to the operation data characteristics of each group, where the abnormal distribution ratio is obtained after training according to the historical operation data of each data type of each service node in the distributed service cluster and the historical abnormal situations of each service node;
[0009] Calculate the index value corresponding to each group according to the abnormal distribution ratio;
[0010] For each data type, calculate the abnormal score corresponding to the data type according to the index values of each group under the data type;
[0011] Sum up the abnormal scores of all data types to obtain the total abnormal score;
[0012] Predict the operation risk of the distributed service cluster according to the total abnormal score.
[0013] Further, grouping each service node according to the operation data of each data type to obtain multiple groups under each data type further includes:
[0014] For each data type, match the operation data of each service node of the data type with the predetermined data range of the data type, and divide the service node into the group corresponding to the predetermined data range whose operation data is successfully matched.
[0015] Further, the operation data characteristic is the predetermined data range;
[0016] The steps of training the abnormal distribution ratio corresponding to the operation data characteristic of each group under a data type include:
[0017] Obtain the historical operation data of each service node of the data type in the previous time period in the distributed service cluster;
[0018] Group each service node according to the historical operation data and the predetermined data range to obtain multiple groups;
[0019] Calculate the predicted number of abnormal service nodes in each group according to the initial abnormal distribution ratio corresponding to the operation data characteristic of each group and the number of service nodes in the corresponding group;
[0020] After the distributed service cluster passes through the previous time period, obtain the actual abnormal results of each service node in the distributed service cluster;
[0021] Calculate the actual number of abnormal service nodes in each group according to the actual abnormal results;
[0022] Judge whether the difference between the actual number of abnormal service nodes and the predicted number of abnormal service nodes in all groups meets the accuracy requirement;
[0023] If there is a group where the difference does not meet the accuracy requirement, adjust the initial abnormal distribution ratio of this group according to the actual number of abnormal service nodes and the predicted number of abnormal service nodes, and when the distributed service cluster runs to the next time period, take the next time period as the previous time period, and repeat the step of obtaining the historical operation data of multiple data types of each service node in the distributed service cluster in the previous time period;
[0024] If the differences of all groups meet the accuracy requirement, take the initial abnormal distribution ratio of each group as the trained abnormal distribution ratio of each group.
[0025] Further, the formula for calculating the metric value corresponding to each group according to the abnormal distribution ratio is:
[0026]
[0027] where P i represents the metric value of the i-th group, and δ i represents the abnormal distribution ratio of the i-th group.
[0028] Further, the formula for calculating the anomaly score corresponding to this data type according to the metric values of each group under this data type is:
[0029]
[0030] where Score j represents the anomaly score corresponding to the j-th data type, N represents the number of groups under the j-th data type, and ω i represents the weight of the i-th group under the j-th data type. where M i represents the number of service nodes in the i-th group, M 总 represents the total number of service nodes in the distributed service cluster, and P i is the metric value of the i-th group under the j-th data type.
[0031] Further, the formula for obtaining the total anomaly score by summing the anomaly scores of all data types is:
[0032]
[0033] where Score 总 represents the total anomaly score, K represents the number of data types, represents the predetermined weight corresponding to the j-th data type.
[0034] Further, the method further includes:
[0035] Judge whether the index value corresponding to each group is 0 respectively;
[0036] If so, filter out the index value of this group.
[0037] Furthermore, the method further includes:
[0038] Calculate the prediction ability of each data type according to the following formula:
[0039]
[0040] where Q j represents the prediction ability of the jth data type, and N represents the number of groups under the jth data type;
[0041] Judge whether the prediction ability is less than the threshold, and if so, filter out this data type.
[0042] On the other hand, the embodiments of the present specification also provide an operation risk monitoring device for a distributed service cluster, and the device includes:
[0043] An operation data acquisition unit, configured to acquire operation data of multiple data types of each service node in the distributed service cluster;
[0044] A grouping unit, configured to group each service node according to the operation data of each data type, obtain multiple groups under each data type, and each group includes several service nodes;
[0045] An abnormal distribution ratio determination unit, configured to determine the trained abnormal distribution ratio corresponding to the operation data characteristics of each group, where the abnormal distribution ratio is obtained after training according to the historical operation data of each data type of each service node in the distributed service cluster and the historical abnormal conditions of each service node;
[0046] An index value calculation unit, configured to calculate the index value corresponding to each group according to the abnormal distribution ratio;
[0047] An abnormal score calculation unit, configured to calculate the abnormal score corresponding to each data type according to the index values of each group under the data type;
[0048] An abnormal total score calculation unit, configured to sum the abnormal scores of all data types to obtain an abnormal total score;
[0049] An abnormal prediction unit, configured to predict the operation risk of the distributed service cluster according to the abnormal total score.
[0050] On the other hand, an embodiment of this specification also provides a computer device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the above-mentioned method is implemented.
[0051] Using the embodiments of this specification, service nodes are grouped according to the operation data of each service node in a distributed service cluster, and each group of operation data has the same operation data characteristics. The embodiments of this specification train the abnormal distribution ratio of the service nodes corresponding to each operation data characteristic, calculate the index value of each group using the abnormal distribution ratio, and finally calculate the total abnormal score of the distributed service cluster, and predict the operation risk of the distributed service cluster according to the total abnormal score. Compared with the traditional method of using threshold monitoring or rule engines to determine whether there are abnormalities in a distributed service cluster, the method of the embodiments of this specification is not limited by the change in the number of service nodes in the distributed service cluster, improving the adaptability of abnormal prediction in a complex dynamic environment, and the abnormal distribution ratio is continuously iteratively updated, thus avoiding prediction deviation and improving prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0053] Figure 1 The figure shows a schematic flowchart of a method for monitoring the operation risk of a distributed service cluster in an embodiment of this specification;
[0054] Figure 2 The figure shows a schematic flowchart of training the abnormal distribution ratio corresponding to the operation data characteristics of each group under a data type in an embodiment of this specification;
[0055] Figure 3 The figure shows a schematic structural diagram of a device for monitoring the operation risk of a distributed service cluster in an embodiment of this specification;
[0056] Figure 4 The figure shows a schematic structural diagram of a computer device in an embodiment of this specification.
[0057]
Description of the Reference Numerals
[0058] 301, Operation data acquisition unit;
[0059] 302, Grouping unit;
[0060] 303, Abnormal distribution ratio determination unit;
[0061] 304. Index value calculation unit;
[0062] 305. Abnormal score calculation unit;
[0063] 306. Total abnormal score calculation unit;
[0064] 307. Abnormal prediction unit;
[0065] 402. Computer device;
[0066] 404. Processor;
[0067] 406. Memory;
[0068] 408. Driving mechanism;
[0069] 410. Input / output module;
[0070] 412. Input device;
[0071] 414. Output device;
[0072] 416. Presentation device;
[0073] 418. Graphical user interface;
[0074] 420. Network interface;
[0075] 422. Communication link;
[0076] 424. Communication bus. Detailed implementation manners
[0077] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the embodiments of this specification.
[0078] It should be noted that in the description, claims and drawings of the embodiments of this specification, the terms "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the embodiments of this specification described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0079] It should be noted that in the technical solutions of the embodiments of this specification, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0080] It should be noted that in the embodiments of this specification, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solutions of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0081] In order to solve the problems existing in the prior art, the embodiments of this specification provide a method for monitoring the operation risks of a distributed service cluster, as Figure 1 shown, the method includes:
[0082] Step 101: Obtain the operation data of multiple data types of each service node in the distributed service cluster;
[0083] Step 102: Group each service node according to the operation data of each data type to obtain multiple groups under each data type, and each group includes several service nodes;
[0084] Step 103: Determine the trained abnormal distribution ratio corresponding to the operation data characteristics of each group, where the abnormal distribution ratio is obtained after training based on the historical operation data of each data type of each service node in the distributed service cluster and the historical abnormal conditions of each service node;
[0085] Step 104: Calculate the index value corresponding to each group according to the abnormal distribution ratio;
[0086] Step 105: For each data type, calculate the abnormal score corresponding to this data type according to the index values of each group under this data type;
[0087] Step 106: Sum up the anomaly scores of all data types to obtain the total anomaly score;
[0088] Step 107: Predict the operation risk of the distributed service cluster based on the total anomaly score.
[0089] Using the embodiments of this specification, the service nodes in the distributed service cluster are grouped according to the operation data of each service node. Each group of operation data has the same operation data characteristics. The embodiments of this specification train the anomaly distribution ratio of the service nodes corresponding to each operation data characteristic, use the anomaly distribution ratio to calculate the metric value of each group, and finally calculate the total anomaly score of the distributed service cluster. The operation risk of the distributed service cluster is predicted based on the total anomaly score. Compared with the traditional method of using threshold monitoring or rule engines to determine whether there are anomalies in the distributed service cluster, the method of the embodiments of this specification is not limited by the change in the number of service nodes in the distributed service cluster, improves the adaptability of anomaly prediction in complex dynamic environments, and the anomaly distribution ratio is continuously iteratively updated, thereby avoiding prediction deviation and improving prediction accuracy.
[0090] In the embodiments of this specification, the data types can be divided according to the experience of the staff. For example, CPU usage rate, memory usage rate, network latency, disk I / O, request response time, etc. are not limited in the embodiments of this specification.
[0091] The data collection frequency can be dynamically adjusted according to actual needs. The embodiments of this specification obtain the operation data of multiple data types within a period of time. It should be noted that the operation data is the average value within this time period.
[0092] In addition, the obtained operation data can also be cleaned and variable screening can be performed, including handling missing values, outliers, and performing necessary data conversions to ensure data quality.
[0093] The anomaly results predicted by the embodiments of this specification can indicate whether faults such as downtime will occur in the distributed service cluster under the current operating conditions. The distributed service cluster includes multiple service nodes. If a service node does not experience faults such as downtime during this time period, the operation data of the service node can be obtained. However, although the operation data can be obtained, this does not mean that the service node will not have operation anomalies under the current operating conditions. The downtime fault of the distributed service cluster occurs in an instant, but the reasons for causing the downtime fault accumulate over time.
[0094] For example, if the CPU usage rate remains too high, if the CPU usage rate of a certain service node continuously exceeds 90% within a period of time, then in this operating condition, the possibility of the service node having operation risk is relatively high.
[0095] However, the running data of only one data type may not be sufficient to illustrate the possibility of running risks. Therefore, in the embodiments of this specification, the running risks are predicted by combining the running data of multiple data types.
[0096] According to an embodiment of this specification, grouping each service node according to the running data of each data type to obtain multiple groups under each data type further includes:
[0097] For each data type, match the running data of this data type of each service node with the predetermined data range of this data type, and divide the service node into the group corresponding to the predetermined data range whose running data is successfully matched.
[0098] For example, multiple ranges of CPU usage rates are preset, including 10% - 30%, 30% - 50%, 50% - 70%, 70% - 80%, 80% - 90%, 90% - 100%. Then, match the comprehensive CPU usage rate of each service node within a certain time period obtained with the above ranges, and divide the service nodes belonging to the same range into one group.
[0099] It should be noted that when predicting the running risks of other distributed service clusters or multiple time periods of the same distributed service cluster, for the running data of the same data type, the same predetermined data range can be used for grouping.
[0100] Then determine the trained abnormal distribution ratio corresponding to the running data characteristics of each group. The running data characteristics can be the predetermined data range. The abnormal distribution ratio can represent the proportion of service nodes expected to have abnormalities within the predetermined data range.
[0101] As Figure 2 shown, the steps of training the abnormal distribution ratio corresponding to the running data characteristics of each group under one data type include:
[0102] Step 201: Obtain the historical running data of this data type of each service node in the previous time period of the distributed service cluster;
[0103] Step 202: Group each service node according to the historical running data and the predetermined data range to obtain multiple groups;
[0104] Step 203: Calculate the predicted number of abnormal service nodes in each group according to the initial abnormal distribution ratio corresponding to the running data characteristics of each group and the number of service nodes in the corresponding group;
[0105] Step 204: After the distributed service cluster has passed through the previous time period, obtain the actual abnormal results of each service node in the distributed service cluster;
[0106] Step 205: Calculate the actual number of abnormal service nodes in each group according to the actual abnormal result.
[0107] Step 206: Determine whether the difference between the actual number of abnormal service nodes and the predicted number of abnormal service nodes in all groups meets the accuracy requirement.
[0108] Step 207: If there is a group in which the difference does not meet the accuracy requirement, adjust the initial abnormal distribution ratio of this group according to the actual number of abnormal service nodes and the predicted number of abnormal service nodes, and when the distributed service cluster runs to the next time period, use the next time period as the previous time period, and repeat the step of obtaining the historical operation data of multiple data types of each service node in the distributed service cluster in the previous time period.
[0109] Step 208: If the differences in all groups meet the accuracy requirement, use the initial abnormal distribution ratio of each group as the trained abnormal distribution ratio of each group.
[0110] In the embodiments of this specification, the staff preset, based on experience, what proportion of service nodes will be abnormal within each predetermined data range to obtain the initial abnormal distribution ratio. However, this initial abnormal distribution ratio may not be accurate enough, so it needs to be continuously iteratively optimized.
[0111] Specifically, first obtain the historical operation data of the previous time period, then group the historical operation data according to the predetermined data to obtain multiple groups, and then calculate the predicted number of abnormal service nodes in each group according to the initial abnormal distribution ratio preset for each group.
[0112] After the distributed service cluster goes through the previous time period, obtain the actual abnormal results of each service node in the distributed service cluster. The actual abnormal result is whether the service node has an abnormality.
[0113] Then, the actual number of abnormal service nodes in each group can be calculated according to the actual abnormal results of each service node and the group to which each service node belongs.
[0114] Then determine whether the difference between the actual number of abnormal service nodes and the predicted number of abnormal service nodes in the group meets the accuracy requirement (this accuracy requirement can be a quantity threshold). If not, it means that the preset abnormal distribution ratio for this group is unreasonable.
[0115] Therefore, the initial abnormal distribution ratio of the group is adjusted according to the actual number of abnormal service nodes and the predicted number of abnormal service nodes. For example, if the predicted number of service abnormal nodes is greater than the actual number of abnormal service nodes, it means that the initial abnormal distribution ratio is too large, and the initial abnormal distribution ratio can be decreased according to a predetermined step size. Conversely, the initial abnormal distribution ratio can be increased according to a predetermined step size.
[0116] Because there is a pattern between the result of whether each service node in the distributed service cluster is abnormal and its running data, through continuous adjustment, the optimal abnormal distribution ratio can ultimately be obtained.
[0117] Subsequently, the index value corresponding to the group is calculated according to the abnormal distribution ratio. The formula is:
[0118]
[0119] Among them, P i represents the index value of the i-th group, and δ i represents the abnormal distribution ratio of the i-th group.
[0120] Among them, 1 - δ i represents the distribution ratio of normal service nodes in the i-th group.
[0121] If the distribution ratio of normal service nodes is greater than the distribution ratio of abnormal service nodes, the index value is greater than 0, and if the distribution ratio of normal service nodes exceeds the distribution ratio of abnormal service nodes by more, the positive index value is larger.
[0122] If the distribution ratio of normal service nodes is less than the distribution ratio of abnormal service nodes, the index value is less than 0, and if the distribution ratio of normal service nodes is less than the distribution ratio of abnormal service nodes by more, the absolute value of the negative index value is larger.
[0123] If the distribution ratio of normal service nodes is equal to the distribution ratio of abnormal service nodes, the index value is equal to 0. In this case, the group has no prediction significance, so the index value of this group is filtered out.
[0124] It should be noted that if the distribution ratio of abnormal service nodes is equal to 0, it means that it is predicted that there are no abnormal service nodes in this group, then this group has no prediction significance.
[0125] According to an embodiment of this specification, the method further includes:
[0126] Calculate the prediction ability of each data type according to the following formula:
[0127]
[0128] Among them, Qj represents the prediction ability of the j-th data type, and N represents the number of groups under the j-th data type;
[0129] Determine whether the prediction ability is less than a threshold, and if so, filter out this data type.
[0130] Then calculate the anomaly score corresponding to this data type according to the metric values of each group under this data type. The formula is:
[0131]
[0132] where, Score j represents the anomaly score corresponding to the j-th data type, N represents the number of groups under the j-th data type, and ω i represents the weight of the i-th group under the j-th data type, where, M i represents the number of service nodes in the i-th group, and M 总 represents the total number of service nodes in the distributed service cluster, and P i is the metric value of the i-th group under the j-th data type.
[0133] In the embodiments of the present specification, the anomaly score is directly related to the number of service nodes in the group. The larger the number of service nodes in a certain group, the larger the weight ω. Conversely, if the number of service nodes in a certain group is smaller, the weight ω is smaller.
[0134] Therefore, if a certain group has a negative metric value P i and its absolute value is large, but if the number of service nodes in this group is very small, then the weight of this group is smaller, so the absolute value of the product result of the weight and the metric value of this group is smaller.
[0135] If the metric value of a certain group is P i and is positive and large, but if the number of service nodes in this group is very small, then the weight of this group is smaller, so the product result of the weight and the metric value of this group is smaller.
[0136] Therefore, the magnitude of the anomaly score depends not only on the magnitude of the metric value but also on the number of service nodes in the group.
[0137] For example, for a group with a CPU utilization rate of 90% - 100%, the probability of abnormal service nodes may be relatively high in this CPU utilization range, that is, the abnormal distribution ratio of this group is relatively large. If the distribution ratio of normal service nodes in this group is less than the distribution ratio of abnormal service nodes, and the distribution ratio of normal service nodes is much less than the distribution ratio of abnormal service nodes, then the metric value of this group is less than 0 and its absolute value is large.
[0138] However, if the number of service nodes in this group is small, then the weight ω of this group i is also small. After multiplying the weight and the index value, the result is negative and the absolute value is also small, which will not make the anomaly score more inclined to be anomalous (the anomaly score is negative and the larger the absolute value, the greater the possibility of being predicted as anomalous).
[0139] Conversely, if the number of service nodes in this group is large, then the weight ω of this group i is also large. After multiplying the weight and the index value, the result is negative and the absolute value is also large, which will make the anomaly score more inclined to be anomalous.
[0140] After calculating the anomaly scores for each data type, sum up the anomaly scores for all data types to obtain the total anomaly score. The formula is as follows:
[0141]
[0142] where Score 总 represents the total anomaly score, K represents the number of data types, represents the predetermined weight corresponding to the jth data type.
[0143] Among them, the predetermined weight corresponding to each data type can be set by the staff based on experience, and the embodiments of this specification do not make limitations.
[0144] It can be understood that if the anomaly score corresponding to a certain data type is negative and the absolute value is larger, it means that the running data of this data type predicts a greater possibility that the distributed service cluster is anomalous.
[0145] Conversely, if the anomaly score corresponding to a certain data type is positive and the larger it is, it means that the running data of this data type predicts a smaller possibility that the distributed service cluster is anomalous.
[0146] Therefore, after summing up the anomaly scores for all data types to obtain the total anomaly score in the embodiments of this specification, if the total anomaly score is negative and the absolute value is larger, then the possibility of predicting that the distributed service cluster is anomalous is greater.
[0147] Conversely, if the total anomaly score is positive and the larger it is, then the possibility of predicting that the distributed service cluster is anomalous is smaller.
[0148] Therefore, the staff can set multiple threshold intervals corresponding to the total anomaly score based on experience or experiments, and set the running risk levels corresponding to each threshold interval, and obtain the running risk level of the distributed service cluster according to the total anomaly score and the threshold intervals.
[0149] The staff can also optimize and adjust the number of service nodes in the distributed service cluster according to the obtained operation risk level, so as to reduce the operation risk of the distributed service cluster.
[0150] In addition, the embodiments of this specification can also use the ROC curve or the AUC curve to monitor the stability of the prediction. If the stability does not meet the requirements, it is necessary to retrain the abnormal distribution ratio.
[0151] Based on the same inventive concept, the embodiments of this specification also provide an operation risk monitoring device for a distributed service cluster, as Figure 3 shown, including:
[0152] An operation data acquisition unit 301, configured to acquire operation data of multiple data types of each service node in the distributed service cluster;
[0153] A grouping unit 302, configured to group each service node according to the operation data of each data type, obtain multiple groups under each data type, and each group includes several service nodes;
[0154] An abnormal distribution ratio determination unit 303, configured to determine the trained abnormal distribution ratio corresponding to the operation data characteristics of each group, where the abnormal distribution ratio is obtained after training according to the historical operation data of each data type of each service node in the distributed service cluster and the historical abnormal conditions of each service node;
[0155] An index value calculation unit 304, configured to calculate the index value corresponding to each group according to the abnormal distribution ratio;
[0156] An abnormal score calculation unit 305, configured to calculate the abnormal score corresponding to each data type according to the index values of each group under the data type for each data type;
[0157] An abnormal total score calculation unit 306, configured to sum the abnormal scores of all data types to obtain an abnormal total score;
[0158] An abnormal prediction unit 307, configured to predict the operation risk of the distributed service cluster according to the abnormal total score.
[0159] Further, grouping each service node according to the operation data of each data type to obtain multiple groups under each data type further includes:
[0160] For each data type, match the operation data of each service node of the data type with the predetermined data range of the data type, and divide the service node into the group corresponding to the predetermined data range whose operation data is successfully matched.
[0161] Further, the operating data characteristic is the predetermined data range;
[0162] The steps of training the abnormal distribution ratio corresponding to the operating data characteristics of each group under a data type include:
[0163] Obtain the historical operating data of this data type of each service node in the distributed service cluster in the previous time period;
[0164] Group each service node according to the historical operating data and the predetermined data range to obtain multiple groups;
[0165] Calculate the predicted number of abnormal service nodes in each group according to the initial abnormal distribution ratio corresponding to the operating data characteristics of each group and the number of service nodes in the corresponding group;
[0166] After the distributed service cluster has passed the previous time period, obtain the actual abnormal results of each service node in the distributed service cluster;
[0167] Calculate the actual number of abnormal service nodes in each group according to the actual abnormal results;
[0168] Judge whether the difference between the actual number of abnormal service nodes and the predicted number of abnormal service nodes in all groups meets the accuracy requirement;
[0169] If there is a group where the difference does not meet the accuracy requirement, adjust the initial abnormal distribution ratio of the group according to the actual number of abnormal service nodes and the predicted number of abnormal service nodes, and when the distributed service cluster runs to the next time period, use the next time period as the previous time period, and repeat the step of obtaining the historical operating data of multiple data types of each service node in the distributed service cluster in the previous time period;
[0170] If the differences in all groups meet the accuracy requirement, take the initial abnormal distribution ratios of each group as the trained abnormal distribution ratios of each group.
[0171] Further, the formula for calculating the index value corresponding to each group according to the abnormal distribution ratio is:
[0172]
[0173] where P i represents the index value of the i-th group, and δ i represents the abnormal distribution ratio of the i-th group.
[0174] Further, the formula for calculating the abnormal score corresponding to this data type according to the index values of each group under this data type is:
[0175]
[0176] Among them, Score j represents the anomaly score corresponding to the j-th data type, N represents the number of groups under the j-th data type, and ω i represents the weight of the i-th group under the j-th data type. Among them, M i represents the number of service nodes in the i-th group, and M 总 represents the total number of service nodes in the distributed service cluster, and P i is the metric value of the i-th group under the j-th data type.
[0177] Furthermore, by summing the anomaly scores of all data types, the formula for obtaining the total anomaly score is:
[0178]
[0179] Among them, Score 总 represents the total anomaly score, K represents the number of data types, represents the predetermined weight corresponding to the j-th data type.
[0180] Furthermore, the device is further configured to:
[0181] respectively determine whether the metric value corresponding to each group is 0;
[0182] If so, filter out the metric value of this group.
[0183] Furthermore, the device is further configured to:
[0184] calculate the prediction ability of each data type according to the following formula:
[0185]
[0186] Among them, Q j represents the prediction ability of the j-th data type, and N represents the number of groups under the j-th data type;
[0187] judge whether the prediction ability is less than the threshold, and if so, filter out this data type.
[0188] The beneficial effects obtained by the above device are the same as those obtained by the above method, and are not elaborated in the embodiments of this specification.
[0189] As Figure 4 shown is the structural schematic diagram of the computer device in the embodiments of this specification, and the computer device in the embodiments of this specification can run the method in the embodiments of this specification.
[0190] The computer device 402 may include one or more processors 404, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 402 may also include any memory 406 for storing any kind of information such as code, settings, data, etc. By way of non-limiting example, for instance, the memory 406 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical discs, etc. More generally, any storage resource may store information using any technology.
[0191] Further, any storage resource may provide volatile or non-volatile retention of information.
[0192] Further, any storage resource may represent a fixed or removable component of the computer device 402. In one case, when the processor 404 executes the associated instructions stored in any storage resource or combination of storage resources, the computer device 402 may perform any operation of the associated instructions. The computer device 402 also includes one or more drive mechanisms 408 for interacting with any storage resource, such as a hard disk drive system, an optical disc drive system, etc.
[0193] The computer device 402 may also include an input / output module 410 (I / O) for receiving various inputs (via the input device 412) and for providing various outputs (via the output device 414). One particular output mechanism may include a presentation device 416 and an associated graphical user interface (GUI) 418. In other embodiments, the input / output module 410 (I / O), the input device 412, and the output device 414 may also be excluded and the computer device may only be a computer device in a network. The computer device 402 may also include one or more network interfaces 420 for exchanging data with other devices via one or more communication links 422. One or more communication buses 424 couple the components described above together.
[0194] The communication link 422 may be implemented in any manner, for example, via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 422 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0195] The embodiments of this specification also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0196] The embodiments of this specification also provide a computer-readable instruction. When a processor executes the instruction, the program therein causes the processor to execute the above method.
[0197] It should be understood that in the various embodiments of this specification, the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this specification.
[0198] It should also be understood that in the embodiments of this specification, the term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the embodiments of this specification generally represents an "or" relationship between the associated objects before and after.
[0199] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this specification can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.
[0200] Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of this specification.
[0201] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0202] In several embodiments provided by the embodiments of this specification, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways.
[0203] For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0204] In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can also be an electrical, mechanical, or other form of connection.
[0205] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units.
[0206] Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this specification.
[0207] In addition, in each of the embodiments of this specification, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0208] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0209] Based on such an understanding, the technical solution of the embodiments of this specification, in essence, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of this specification.
[0210] The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0211] Specific embodiments are applied in the embodiments of this specification to elaborate on the principles and implementation manners of the embodiments of this specification. The description of the above embodiments is only used to help understand the method and its core idea of the embodiments of this specification; at the same time, for those of ordinary skill in the art, based on the idea of the embodiments of this specification, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the embodiments of this specification.
Claims
1. A method for monitoring the operation risk of a distributed service cluster, characterized in that: The method comprises: Obtain the operating data of multiple data types of each service node in the distributed service cluster; Grouping each service node according to the running data of each data type to obtain multiple groups under each data type, each group including a number of service nodes; Determine a trained anomaly distribution ratio corresponding to the respective operating data characteristics of each group, wherein the anomaly distribution ratio is obtained after training based on the historical operating data of each data type of each service node in the distributed service cluster and the historical anomaly situation of each service node; Calculate the index value corresponding to each group according to the abnormal distribution ratio; For each data type, calculate the anomaly score corresponding to the data type according to the indicator values of each group under the data type; Sum the anomaly scores of all data types to get the total anomaly score; The operation risk of the distributed service cluster is predicted according to the total abnormality score.
2. The method according to claim 1, characterized in that The service nodes are grouped according to the operation data of each data type, and the multiple groups under each data type are obtained, further including: For each data type, the running data of each service node of the data type is matched with a predetermined data range of the data type, and the service node is divided into a group corresponding to the predetermined data range whose running data successfully matches.
3. The method according to claim 2, characterized in that The operating data characteristic is the predetermined data range; The steps of training the abnormal distribution ratio corresponding to the running data characteristics of each group under a data type include: Obtaining historical operation data of the data type of each service node in the distributed service cluster in the previous period; Grouping each service node according to the historical operation data and the predetermined data range to obtain a plurality of groups; Calculate the predicted number of abnormal service nodes for each group according to the initial abnormal distribution ratio corresponding to the operating data characteristics of each group and the number of service nodes in the corresponding group; After the distributed service cluster has passed a previous period, obtaining actual abnormal results of each service node in the distributed service cluster; Calculate the actual number of abnormal service nodes in each group according to the actual abnormal result; Determine whether the difference between the actual number of abnormal service nodes and the predicted number of abnormal service nodes of all groups meets the accuracy requirement; If there is a group whose difference does not meet the accuracy requirement, the initial abnormal distribution ratio of the group is adjusted according to the actual number of abnormal service nodes and the predicted number of abnormal service nodes, and when the distributed service cluster runs to a later time period, the later time period is used as the previous time period, and the step of obtaining the historical operation data of multiple data types of each service node in the distributed service cluster in the previous time period is repeated; If the differences of all groups meet the accuracy requirement, the initial abnormal distribution ratio of each group is used as the trained abnormal distribution ratio of each group.
4. The method according to claim 1, characterized in that The formula for calculating the index value corresponding to each group according to the abnormal distribution ratio is: Among them, P i represents the index value of the i-th group, δ i Represents the abnormal distribution ratio of the i-th group.
5. The method according to claim 4, characterized in that The formula for calculating the anomaly score corresponding to the data type according to the indicator value of each group under the data type is: Among them, Score j represents the anomaly score corresponding to the j-th data type, N represents the number of groups under the j-th data type, ω i represents the weight of the i-th group under the j-th data type, Among them, M i represents the number of service nodes in the i-th group, M 总 represents the total number of service nodes in the distributed service cluster, P i The index value of the i-th group under the j-th data type.
6. The method according to claim 5, characterized in that The formula for summing up the anomaly scores of all data types and obtaining the total anomaly score is: Among them, Score 总 represents the total abnormal score, K represents the number of data types, Indicates the predetermined weight corresponding to the j-th data type.
7. The method according to claim 4, characterized in that The method further comprises: Determine whether the indicator value corresponding to each group is 0; If so, filter out the indicator value of the group.
8. The method according to claim 4, characterized in that The method further comprises: The predictive power of each data type is calculated according to the following formula: Among them, Q j represents the prediction ability of the j-th data type, and N represents the number of groups under the j-th data type; It is determined whether the prediction ability is less than a threshold, and if so, the data type is filtered out.
9. A distributed service cluster operation risk monitoring device, characterized in that: The device comprises: An operation data acquisition unit, used to acquire operation data of multiple data types of each service node in the distributed service cluster; A grouping unit, used to group each service node according to the operation data of each data type, to obtain multiple groups under each data type, each group including a number of service nodes; An abnormal distribution ratio determination unit is used to determine a trained abnormal distribution ratio corresponding to the operation data characteristics of each group, wherein the abnormal distribution ratio is obtained after training based on the historical operation data of each data type of each service node in the distributed service cluster and the historical abnormal conditions of each service node; An index value calculation unit, used to calculate the index value corresponding to each group according to the abnormal distribution ratio; An anomaly score calculation unit, used to calculate, for each data type, an anomaly score corresponding to the data type according to the indicator values of each group under the data type; The anomaly total score calculation unit is used to sum the anomaly scores of all data types to obtain the anomaly total score; The abnormality prediction unit is used to predict the operation risk of the distributed service cluster according to the total abnormality score.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.