Cloud server anomaly detection method
By subdividing the business cycle into multiple minimal business scenarios and using the K-Means clustering algorithm for abnormal detection, the problem of high computing power and resource burden in the cloud server abnormal detection method is solved, the processing efficiency and detection accuracy are improved, and the stability of cloud services is ensured.
Patent Information
- Application Number
- CN202510098034.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-10
AI Technical Summary
The existing cloud server abnormality detection methods have high computing power and resource burdens, resulting in low processing efficiency and complex machine learning models that are difficult to locate and interpret abnormalities, reducing detection accuracy.
By subdividing the complete business cycle into multiple minimal business scenarios, the historical operation indicator data of all applications in the cloud server is collected, and the K-Means clustering algorithm is used for classification, and abnormal detection is performed according to the distribution characteristics of each indicator data in each business scenario of each application.
It reduces the computing power burden of cloud server abnormal detection, improves processing efficiency, improves the reliability and fault tolerance of cloud servers, and ensures the stability and continuity of cloud services.
Smart Images

Figure CN120123923A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an anomaly detection method, in particular to an anomaly detection method for cloud servers. Background Art
[0002] A cloud server (Elastic Compute Service, abbreviated as ECS) is a simple, efficient, secure and reliable computing service that provides the ability to elastically scale, enabling developers to directly deploy applications on it without having to concern themselves with the underlying hardware. However, despite the broad application prospects of cloud server hosting services, their availability remains a major challenge for cloud service providers.
[0003] Cloud servers often experience various anomalies that affect their availability. Software and hardware failures or human errors may cause cloud servers to crash, thereby affecting user experience and service quality. Multiple cloud services are deployed on the same physical machine to share resources and kernels, and there is a potential anomaly risk that cloud services may exploit kernel vulnerabilities to affect the confidentiality and availability of the physical machine and the cloud platform. To ensure high service availability and improve the disaster recovery performance of the cloud environment, the industry generally adopts large-scale distributed cluster systems. When a device or component fails, the system can quickly transfer the affected service to other normally operating devices or components to achieve rapid recovery and reconstruction. Therefore, whether it is a public cloud or a private cloud, improving the self-detection, self-reconfiguration, and self-recovery capabilities of the cloud environment is a major issue faced by operation and maintenance personnel.
[0004] The monitoring data of cloud servers themselves and hosting services is huge and diverse. It is unrealistic to rely on operation and maintenance personnel to observe traditionally. Although monitoring tools can assist in screening key performance indicator (KPI) data and setting threshold comparisons, they lack universality, rely on operation and maintenance personnel, and have no predictive ability. In recent years, intelligent operation and maintenance has become an important approach. It uses artificial intelligence algorithms to analyze historical KPI data and automatically identify anomaly types, and uses predictive algorithms to diagnose future data, issue early warnings for repair, and avoid failures.
[0005] Among various KPI data in cloud environment monitoring, most are time series data. With the rise of artificial intelligence in recent years, many anomaly detection methods for time series data in cloud environments based on supervised artificial intelligence have emerged. The large amount of data generated by the cloud environment also provides a data basis for training and optimizing artificial intelligence algorithm models. However, supervised artificial intelligence algorithms have problems in data acquisition. The KPI data in the cloud environment is imbalanced in categories and naturally has no labels, and it is impossible to construct a large number of labeled data sets for supervised artificial intelligence algorithms to use.
[0006] Therefore, there has been a lot of research on anomaly detection algorithms in the cloud environment of unsupervised artificial intelligence. For example, log anomaly detection based on deep learning uses LSTM to model and learn patterns from log content. Another example is the active learning anomaly detection algorithm based on a tree structure, which receives user feedback labels for active learning. However, most of these algorithms are too complex. The large unsupervised artificial intelligence algorithm models bring a relatively large computing power and resource burden to the real-time anomaly detection module deployed on the cloud server, thus reducing the processing efficiency of the cloud server. At the same time, the black-box anomaly detection algorithms also bring certain difficulties to locating, explaining, and fixing anomalies, thereby reducing the accuracy of anomaly detection. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a method for anomaly detection of a cloud server, which reduces the computing power and resource burden for anomaly detection in the cloud server, thereby improving the reliability and fault tolerance of the cloud server.
[0008] The technical solution adopted by the present invention to solve the above technical problems is as follows: A method for anomaly detection of a cloud server includes the following steps:
[0009] Step 1): During a preset complete business cycle, divide it at intervals of a preset minimum scenario cycle to obtain multiple consecutive minimum business scenarios with a time length of the preset minimum scenario cycle.
[0010] Step 2): Collect the index data of all application historical operations running in the cloud server. According to the collected index data, obtain the average value of each index data in the time interval corresponding to each minimum business scenario for each application. Take the average value of all index data in each minimum business scenario as the minimum scenario index corresponding to this minimum business scenario, and form a data set with the minimum scenario indexes corresponding to all minimum business scenarios in each application.
[0011] Step 3): Use the K-Means clustering algorithm to cluster each data set to obtain multiple business scenarios corresponding to each application.
[0012] Step 4): Obtain the distribution characteristics of each index data in each business scenario of each application. Perform anomaly detection on the cloud server according to the distribution characteristics, set the normal interval of the detection result according to the distribution characteristics. If the mean value of the index data within the set detection time period falls within the normal interval of the detection result, it is determined that the index data is normal. If the mean value of the index data within the set detection time period falls outside the normal interval of the detection result, it is determined that the index data is abnormal and marked as an abnormal point. Set the anomaly detection sensitivity, and perform statistics every interval of the duration of the set detection time period. When the total number of abnormal points is more than the set anomaly detection sensitivity, it is determined that an application anomaly is detected, and the anomaly detection of the cloud server is completed.
[0013] Compared with the prior art, the advantages of the present invention are as follows: by subdividing the complete business cycle into multiple minimum business scenarios, collecting the historical operation index data of all applications in the cloud server, obtaining the average value of each index data of each application in each minimum business scenario according to the index data, and taking the average value of all index data in each minimum business scenario as the minimum scenario index, forming a data set with all minimum scenario indexes in each application, using the K-Means clustering algorithm to classify each data set, obtaining multiple business scenarios corresponding to each application, and realizing the efficient detection of cloud server anomalies according to the distribution characteristics of each index data in each business scenario of each application. Compared with complex machine learning models, the present invention uses a relatively simple statistical method for real-time anomaly detection, greatly reducing the computing power burden, improving the processing efficiency of the cloud server, thus enhancing the reliability and fault tolerance of the cloud server and ensuring the stability and continuity of cloud services.
[0014] Further, the index data includes CPU utilization rate, memory utilization rate, disk utilization rate, read / write data volume, and network traffic.
[0015] Further, the specific operation process of step 3) is as follows:
[0016] Step 3-1: Randomly select K = 1 minimum scenario index from the f-th data set as the starting centroid, where 1 ≤ f ≤ H, and H represents the total number of applications in the cloud server;
[0017] Step 3-2: Calculate the Euclidean distance d between each minimum scenario index in the data set and each starting centroid aj , where, X Ta represents the a-th minimum scenario index, 1 ≤ a ≤ n, and n represents the total number of minimum scenario indexes in the data set, μ j represents the j-th starting centroid, 1 ≤ j ≤ K. According to the calculated Euclidean distance, assign each minimum scenario index in the data set to the set where the starting centroid with the closest Euclidean distance is located, obtaining K sets, and each set corresponds to a business scenario;
[0018] Step 3-3: For the K sets, calculate the average value of all minimum scenario indexes in each set, and take this average value as the new centroid μ j ’ where, 1 ≤ j ≤ K, C j represents the j-th set;
[0019] Step 3-4: If all new centroids {μ 1 ’, …, μ KIf none of them has changed, calculate the silhouette coefficient s of the a-th smallest scenario metric a , s a = (B a - A a ) / max(A a , B a ), and execute steps 3 - 5, where 1 ≤ a ≤ n, A a represents the average distance from the a-th smallest scenario metric to all the smallest scenario metrics in its set, and B a represents the average distance from the a-th smallest scenario metric to all the smallest scenario metrics not in its set; if the new centroid has changed, return to execute step 3 - 2;
[0020] Step 3 - 5: Output K clustering quantities {C 1 , C 2 , …, C K}, and the corresponding clustering silhouette coefficient S K = ∑s a / n;
[0021] Step 3 - 6: Let K = K + 1, return to execute step 3 - 1 until K = 10, obtain 10 clustering silhouette coefficients {S 1 , S 2 , …, S 10}, select the one with the largest value from all the clustering silhouette coefficients as the maximum silhouette coefficient, denoted as S, and denote the corresponding clustering quantity as k'; if S > 0.5, then the number of business scenarios k after clustering for this application is k' = k; if S ≤ 0.5, then the number of business scenarios k after clustering for this application is k = 1;
[0022] Step 3 - 7: Output the k business scenarios corresponding to the f-th application, where 1 ≤ f ≤ H, and H represents the total number of applications in the cloud server.
[0023] Furthermore, the specific operation process of step 4) is as follows:
[0024] Step 4 - 1: Take all the metric data within each business scenario of each application in the cloud server as input, perform normal distribution fitting on each metric data respectively, and output the maximum likelihood estimate value of the mean of each metric data and the maximum likelihood estimate value of the standard deviation of each metric data, denoted as μ m and σ m , where 1 ≤ m ≤ M, and M represents the total number of metric data;
[0025] Step 4 - 2: Perform anomaly detection on each business scenario according to μ m and σ m specifically as follows: Set the sampling frequency as F and the detection interval as tsample , perform anomaly detection once every time interval of Ft sample , calculate the mean value of each metric data in the current business scenario within the time period of Ft sample , and mark the metric data corresponding to the mean value falling outside the interval of (μ m - 3σ m , μ m + 3σ m ) as 3σ anomaly points, and mark the metric data corresponding to the mean value falling outside the interval of (μ m - 2σ m , μ m + 2σ m ) and within the interval of (μ m - 3σ m , μ m + 3σ m ) as 2σ anomaly points;
[0026] Step 4 - 3: Set the anomaly detection sensitivity and denote it as G, and perform statistics once every time interval of Ft sample . When the total number of 3σ anomaly points is more than G or the total number of 2σ anomaly points is more than 3G, mark that an application anomaly is detected, and complete the anomaly detection of the cloud server. It can clearly identify which metric data exceeds the normal range in a specific business scenario and is thus marked as an anomaly point, while those metric data falling within the normal interval are considered relatively normal. This not only improves the accuracy of anomaly detection but also provides clear instructions for the operation and maintenance personnel, enabling them to quickly identify and respond to potential problems, thereby taking timely measures to reduce potential service interruptions or performance degradation, ensuring the stable operation of the cloud server and service quality.
[0027] Furthermore, between step 3) and step 4), it further includes: repeating step 2) and step 3) every set update time to obtain new multiple business scenarios corresponding to each application. This can ensure that after the application update and iteration change, the division of the business scenarios still conforms to the functional characteristics of the updated application.
[0028] Furthermore, the set update time is 3 to 6 months. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is the overall flow diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The present invention will be further described in detail below with reference to the embodiments in the drawings.
[0031] As Figure 1 shown, a method for anomaly detection of a cloud server includes the following steps:
[0032] Step 1): Divide within a preset complete business cycle at intervals of a preset minimum scenario cycle to obtain multiple consecutive minimum business scenarios with a time length of the preset minimum scenario cycle;
[0033] Denote the i-th minimum business scenario as T i , which is located in the time interval [(i - 1)T min , iT min )], where, T min represents the preset minimum scenario cycle, and T represents the preset complete business cycle, both of which are set by the user himself. In this embodiment, without any configuration experience, T min = 20 min. Considering that most online cloud applications receive user inputs to perform corresponding tasks and there are relatively fixed workload peaks and valleys within a day, so T = 1 day is set;
[0034] In this setting, a complete business cycle contains 72 minimum scenario cycles, that is, T = 72T min , and each minimum scenario cycle can be described as T 1 , T 2 , …, T 72 .
[0035] In the cloud server periodic scenario division step, a time period that needs to be divided into a specific business scenario can be additionally added, and additional calculation of the scenario metric benchmark value and generation of scenario description configuration can be performed. This solution allows users to adjust the anomaly detection algorithm according to different specific usage scenarios. Users can adapt unique periodic characteristics for each specific application, and further make the anomaly detection algorithm adapt to cloud applications or cloud services with different application performance cycle lengths and no obvious periodic characteristics.
[0036] Step 2): Collect the historical operation metric data of all applications running in the cloud server. According to the collected metric data, obtain the average value of each metric data in the time interval corresponding to each minimum business scenario for each application. Take the average value of all metric data in each minimum business scenario as the minimum scenario metric corresponding to this minimum business scenario, and form a data set with the minimum scenario metrics corresponding to all minimum business scenarios in each application; the metric data includes CPU utilization rate, memory utilization rate, disk utilization rate, read / write data volume, and network traffic;
[0037] Group the metric data collected at the same time interval corresponding to the same minimum business scenario and from the same application into a set of data. Each set of data includes a time series within the preset minimum scenario cycle. Among them, the time series at time t is X t ={x 1 , x2 ,…,x M}, 0 ≤ t ≤ T, where M represents the total number of index data; aggregate each group of data to obtain the average value of each index data;
[0038] The n smallest business scenarios correspond to n smallest scenario indexes, and each of the smallest scenario indexes contains the average value of M index data; among them, the f-th data set X f ={X f1 , X f2 ,…, X fn}, 1 ≤ f ≤ H, where H represents the total number of applications in the cloud server.
[0039] Step 3): Use the K-Means clustering algorithm to cluster each data set to obtain multiple business scenarios corresponding to each application; the specific operation process is as follows:
[0040] Step 3-1: Randomly select K = 1 smallest scenario index from the f-th data set as the starting centroid, where 1 ≤ f ≤ H, and H represents the total number of applications in the cloud server.
[0041] Step 3-2: Calculate the Euclidean distance d aj , between each smallest scenario index in the data set and each starting centroid, where X Ta represents the a-th smallest scenario index, 1 ≤ a ≤ n, and n represents the total number of smallest scenario indexes in the data set, and μ j represents the j-th starting centroid, 1 ≤ j ≤ K. According to the calculated Euclidean distance, assign each smallest scenario index in the data set to the set where the starting centroid with the closest Euclidean distance is located, and obtain K sets, and each set corresponds to a business scenario.
[0042] Step 3-3: For the K sets, calculate the average value of all smallest scenario indexes in each set, and use this average value as the new centroid μ j ’ of the corresponding set, where 1 ≤ j ≤ K, and C j represents the j-th set.
[0043] Step 3-4: If all new centroids {μ 1 ’,…, μ K ’} have not changed, then calculate the silhouette coefficient s a of the a-th smallest scenario index, s a =(B a -A a ) / max(A a , B a ), and execute Step 3-5, where 1 ≤ a ≤ n, Aa represents the average distance from the a-th smallest scenario metric to all the smallest scenario metrics in its set, B a represents the average distance from the a-th smallest scenario metric to all the smallest scenario metrics not in its set; if the new centroid has changed, return to execute step 3-2.
[0044] Step 3-5: Output K clustering numbers {C 1 , C 2 , …, C K}, and the corresponding silhouette coefficient S K = ∑s a / n.
[0045] Step 3-6: Let K = K + 1, return to execute step 3-1 until K = 10, and obtain 10 silhouette coefficients {S 1 , S 2 , …, S 10}. Select the one with the largest value from all the silhouette coefficients as the maximum silhouette coefficient, denoted as S, and denote the corresponding clustering number as k'; if S > 0.5, then the number of business scenarios k after clustering for this application is k' ; if S ≤ 0.5, then the number of business scenarios k after clustering for this application is 1.
[0046] Step 3-7: Output the k business scenarios corresponding to the f-th application, where 1 ≤ f ≤ H, and H represents the total number of applications in the cloud server.
[0047] Step 4): Obtain the distribution characteristics of each metric data in each business scenario of each application, perform anomaly detection on the cloud server according to the distribution characteristics, set the normal interval of the detection result according to the distribution characteristics. If the mean of the metric data within the set detection time period falls within the normal interval of the detection result, then determine that the metric data is normal; if the mean of the metric data within the set detection time period falls outside the normal interval of the detection result, then determine that the metric data is abnormal and mark it as an abnormal point. Set the anomaly detection sensitivity, and perform statistics every time the duration of the set detection time period elapses. When the total number of abnormal points is more than the set anomaly detection sensitivity, determine that an application anomaly has been detected, and complete the anomaly detection of the cloud server; the distribution characteristics of each metric data include the maximum likelihood estimate value of the mean of each metric data and the maximum likelihood estimate value of the standard deviation of each metric data; the specific operation process is as follows:
[0048] Step 4-1: Take all the metric data within each business scenario of each application in the cloud server as input, perform normal distribution fitting on each metric data respectively, and output the maximum likelihood estimate value of the mean of each metric data and the maximum likelihood estimate value of the standard deviation of each metric data, denoted as μ m and σ m, where \(1\leq m\leq M\), and \(M\) represents the total number of index data.
[0049] Step 4-2: According to \(\mu\) m and \(\sigma\) m Perform anomaly detection on each business scenario. Specifically: Set the sampling frequency to \(F\) and the detection interval to \(t\) sample , and perform anomaly detection once every \(Ft\) sample time. Calculate the mean value of each index data in the current business scenario within the \(Ft\) sample time period. Mark the index data corresponding to the mean value falling outside the interval \((\mu\) m - 3\(\sigma\) m , \(\mu\) m + 3\(\sigma\) m ) as 3\(\sigma\) anomaly points, and mark the index data corresponding to the mean value falling outside the interval \((\mu\) m - 2\(\sigma\) m , \(\mu\) m + 2\(\sigma\) m ) and within the interval \((\mu\) m - 3\(\sigma\) m , \(\mu\) m + 3\(\sigma\) m ) as 2\(\sigma\) anomaly points; Both \(F\) and \(t\) sample can be set by the user himself. In this embodiment, \(F = 20s\), \(t\) sample = 13 times.
[0050] Step 4-3: Set the anomaly detection sensitivity and denote it as \(G\). Conduct a statistics once every \(Ft\) sample time. When the total number of 3\(\sigma\) anomaly points is more than \(G\) or the total number of 2\(\sigma\) anomaly points is more than \(3G\), mark that an application anomaly is detected, and complete the anomaly detection of the cloud server; \(G\) is set by the user himself. In this embodiment, \(G = 1\).
[0051] In this embodiment, between step 3) and step 4), it further includes: Every set update time, the set update time is arbitrarily selected within the range of 3 to 6 months, repeat steps 2) and 3) to obtain new multiple business scenarios corresponding to each application.
Claims
1. A cloud server anomaly detection method, characterized in that The following steps are involved: Step 1): within a preset complete business cycle, divide it into intervals of a preset minimum scenario cycle to obtain a plurality of minimum business scenarios whose continuous time length is the preset minimum scenario cycle; Step 2): Collect the historical running indicator data of all applications running in the cloud server, and obtain the average value of each indicator data of each application in the time interval corresponding to each minimum business scenario based on the collected indicator data, and use the average value of all indicator data in each minimum business scenario as the minimum scenario indicator corresponding to the minimum business scenario, and form a data set with the minimum scenario indicators corresponding to all minimum business scenarios in each application; Step 3): Use the K-Means clustering algorithm to cluster each data set to obtain multiple business scenarios corresponding to each application; Step 4): Obtain the distribution characteristics of each indicator data in each business scenario of each application, perform anomaly detection on the cloud server based on the distribution characteristics, set the normal range of detection results based on the distribution characteristics, if the mean value of the indicator data within the set detection time period falls within the normal range of detection results, then the indicator data is judged to be normal, if the mean value of the indicator data within the set detection time period falls outside the normal range of detection results, then the indicator data is judged to be abnormal and marked as an abnormal point, set the anomaly detection sensitivity, and perform statistics once every interval of the set detection time period. When the total number of abnormal points is more than the set anomaly detection sensitivity, it is judged that the application abnormality is detected, and the anomaly detection of the cloud server is completed.
2. A cloud server anomaly detection method according to claim 1, characterized in that In the step 2), the indicator data includes CPU utilization, memory utilization, disk utilization, read and write data volume and network traffic.
3. A cloud server anomaly detection method according to claim 1, characterized in that The specific operation process of step 3) is as follows: Step 3-1: Randomly select K = 1 minimum scenario indicators from the fth data set as the starting centroid, where 1≤f≤H, H represents the total number of applications in the cloud server; Step 3-2: Calculate the Euclidean distance d from each minimum scene index in the data set to each starting centroid aj , Among them, X Ta represents the ath minimum scene index, 1≤a≤n, n represents the total number of minimum scene indexes in the data set, μ j represents the jth starting centroid, 1≤j≤K. According to the calculated Euclidean distance, each minimum scenario indicator in the data set is assigned to the set with the starting centroid closest to it in Euclidean distance, and K sets are obtained, each set corresponding to a business scenario; Step 3-3: For K sets, calculate the average of all minimum scene indicators in each set and use the average as the new centroid μ of the corresponding set j ’ , Among them, 1≤j≤K, C j represents the jth set; Step 3-4: If all new centroids {μ1 ’ ,…,μ K ’ } have not changed, then calculate the silhouette coefficient s of the ath minimum scene index a ,s a =(B a -A a ) / max(A a ,B a ), and execute steps 3-5, where 1≤a≤n, A a represents the average distance from the ath minimum scene index to all the minimum scene indexes in its set, B a Represents the average distance from the ath minimum scene index to all minimum scene indexes in the non-belonging set; if the new centroid has changed, return to execute step 3-2; Step 3-5: Output the number of K clusters {C1, C2, …, C K }, and the corresponding clustering silhouette coefficient S K =∑s a / n; Step 3-6: Let K = K + 1, return to step 3-1, and continue until K = 10, obtaining 10 clustering silhouette coefficients {S1, S2, …, S 10 }, select the largest value from all clustering silhouette coefficients as the maximum silhouette coefficient, denoted as S, and the corresponding number of clusters is denoted as k ’ ; If S>0.5, the number of business scenarios after the application is clustered is k=k ’ , if S≤0.5, the number of business scenarios after the application is clustered is k=1; Step 3-7: Output k business scenarios corresponding to the f-th application, where 1≤f≤H, and H represents the total number of applications in the cloud server.
4. A cloud server anomaly detection method according to claim 1, characterized in that The specific operation process of step 4) is as follows: Step 4-1: Take all the indicator data in each business scenario of each application in the cloud server as input, fit each indicator data with a normal distribution, and output the maximum likelihood estimate of the mean of each indicator data and the maximum likelihood estimate of the standard deviation of each indicator data, which are denoted as μ m and σ m , where 1≤m≤M, M represents the total number of indicator data; Step 4-2: According to μ m and σ m Perform anomaly detection for each business scenario, specifically: set the sampling frequency to F and the detection interval to t sample , each interval Ft sample Perform anomaly detection once in the time to calculate the value of each indicator data in the current business scenario in Ft sample The mean value within the time period is placed in (μ m -3σ m , μ m +3σ m ) interval is marked as 3σ outliers, and the mean falls within (μ m -2σ m , μ m +2σ m ) interval and outside (μ m -3σ m , μ m +3σ m ) The corresponding indicator data within the interval is marked as a 2σ abnormal point; Step 4-3: Set the anomaly detection sensitivity and record it as G, each interval Ft sample When the total number of 3σ anomaly points is greater than G or the total number of 2σ anomaly points is greater than 3G, it is marked that an application anomaly is detected, and the anomaly detection of the cloud server is completed.
5. A cloud server anomaly detection method according to claim 1, characterized in that The step 3) and the step 4) further include: repeating the step 2) and the step 3) at every set update time to obtain multiple new business scenarios corresponding to each application.
6. A cloud server anomaly detection method according to claim 5, characterized in that The update period of the settings is 3 to 6 months.