Multi-connection clustering method for data center cluster
By repeatedly executing the clustering steps multiple times, the problem of inaccurate and unstable clustering results in data center cluster cluster analysis is solved, and more accurate and stable clustering results are achieved, the number of clusters can be automatically adjusted, and the resource consumption pattern of the data center is more flexible.
Patent Information
- Application Number
- CN202411793406.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The prior art has problems in the cluster analysis of data center clusters that are inaccurate, unstable and difficult to automatically adjust the number of clusters, especially when complex shapes and clusters of different sizes exist.
The method of repeatedly performing the clustering step is adopted, and the monitoring data is randomly selected for initial clustering, and the cluster center and data allocation are repeatedly calculated until the stop condition is met. Finally, the final clustering result is obtained based on the multiple clustering results.
It effectively reduces the uncertainty of initial value selection on clustering results, improves the accuracy and stability of clustering, can automatically adjust the number of clusters, and more flexibly and accurately reflects the real resource consumption pattern of the data center.
Smart Images

Figure CN119939289A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data center cluster applications, and in particular to a multi-connection clustering method for data center clusters. Background Art
[0002] In the era of big data, the scale of data center clusters continues to expand, and the complexity of their applications is also increasing. Traditional data processing and analysis methods are unable to cope with the monitoring data of massive applications, especially in the classification of large-scale applications, which has many challenges and limitations. At present, by using clustering algorithms to perform cluster analysis on the monitoring data of data center cluster applications, it is possible to effectively classify and organize massive application loads, thereby discovering monitoring data with similar load patterns, thereby optimizing cluster resource allocation patterns and improving cluster scheduling management efficiency; it can also be used for anomaly detection, identifying monitoring data with abnormal resource consumption, and performing targeted performance optimization, thereby improving the scheduling efficiency and resource utilization of data center cluster applications.
[0003] The most classic clustering algorithm is the K-means algorithm, whose core idea is to divide the data set into K clusters through an iterative process, so that the distance between each data point and the center point (centroid) of its cluster is minimized, and each cluster has similar characteristics. Although the K-means algorithm is simple and easy to understand, it requires the user to pre-specify the number of clusters K at the beginning. In practical applications, it is often difficult to determine the actual number of clusters of the data in advance. Choosing an inappropriate K value may lead to inaccurate clustering results or inconsistency with the actual situation; and the selection of the initial value in the K-means algorithm directly affects the convergence and clustering results of the K-means algorithm. If the initial value is not selected properly, the algorithm may fall into a local optimal solution and fail to find a global optimal solution, resulting in unstable and inaccurate clustering results. The same data may get different clustering results under different initial conditions. In addition, the K-means algorithm assumes that clusters are convex and similar in shape and size. It is less effective for clusters of complex shapes and different sizes. Clusters of complex shapes may be incorrectly segmented or merged. When the actual sizes of clusters vary greatly, small clusters may be ignored or misjudged as noise. Large clusters may be over-subdivided or contain irrelevant data points, thus failing to accurately reflect the actual distribution of the data. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a multi-connection clustering method for data center clusters, which can obtain accurate clustering results that conform to the actual situation.
[0005] The technical solution adopted by the present invention to solve the above technical problem is: a multi-connection clustering method for data center clusters, comprising the following steps:
[0006] Step ①, randomly selecting N monitoring data from the monitoring system of the data center cluster application to form a data set, and normalizing the data set to obtain a normalized data set;
[0007] Step ②, perform initial clustering on the normalized data set to obtain K clusters;
[0008] Step ③, for K clusters, obtain the average value of all monitoring data in each cluster, and use the average value as the new cluster center;
[0009] Step ④, obtain the Euclidean distance between each monitoring data in the normalized data set and the new cluster center, and assign each monitoring data in the normalized data set to the cluster where the new cluster center with the closest Euclidean distance is located according to the obtained Euclidean distance, and obtain K clusters;
[0010] Step ⑤, repeat steps ③ to ④ until the stop condition is met, and then execute step ⑥;
[0011] Step ⑥, numbering each cluster obtained in step ⑤, obtaining and recording the cluster number of each monitoring data;
[0012] Step ⑦, repeat steps ② to ⑥ until the number of repetitions reaches L times. Each monitoring data is given L cluster numbers. The time obtained according to the L cluster numbers of each monitoring data is aggregated into a multi-group in chronological order. The monitoring data corresponding to the same multi-group are classified into the same category, and Knew categories are obtained to complete the clustering of the data center cluster.
[0013] Compared with the prior art, the advantage of the present invention is that the method of repeatedly executing the clustering steps, i.e., step ③ to step ④, effectively reduces the uncertainty brought to the clustering results by the randomness of the initial value selection, and improves the accuracy and stability of the data center clustering; repeated execution can ensure that the algorithm explores the data space from multiple angles, and finally converges to a more stable clustering center, avoiding the local optimal solution problem that may be caused by a single run. The present invention aggregates multiple clustering results into a multi-tuple, and the K clusters obtained by the initial clustering do not represent the final clustering results, but gradually increase with multiple runs, which solves the disadvantage that the existing K-means algorithm needs to predetermine the initial value, and can obtain accurate and actual clustering results, and enhances the stability and availability of the data center clustering results; at the same time, this method can automatically adjust the number of clusters according to the actual data distribution, and more flexibly and accurately reflect the real resource consumption mode of the data center. In addition, the constructed data set is clustered into Knew clusters, and the monitoring data in each cluster has similar resource consumption, which is convenient for subsequent application category management and anomaly detection in the data center cluster. By building more stable and accurate clustering results, clusters with abnormal behaviors can be more easily identified, which is crucial for data center operations and maintenance because it helps to promptly discover and resolve potential problems, prevent system failures and security threats, and thus improve the overall system performance and security of the data center cluster.
[0014] Furthermore, in step ①, each of the monitoring data includes multiple indicator data;
[0015] The indicator data includes CPU, memory, and disk IO;
[0016] The specific operation of normalizing the data set to obtain the normalized data set is: using the minimum and maximum normalization method to scale each indicator data in the data set to the range of 0-1 to obtain the normalized data set. The minimum and maximum normalization method can eliminate the influence of different dimensions and different value ranges on the calculation method of data point distance, ensuring the comparability and fairness between different indicator data.
[0017] Furthermore, the specific operation process of step ② is as follows:
[0018] Step ②-1, randomly select K monitoring data from the normalized data set as the initial cluster center;
[0019] Step ②-2, obtain the Euclidean distance between each monitoring data in the normalized data set and the K initial cluster centers respectively, and assign each monitoring data in the normalized data set to the cluster where the initial cluster center with the closest Euclidean distance is located according to the obtained Euclidean distance, and obtain K clusters.
[0020] Furthermore, K ≥ 2.
[0021] Furthermore, in step ⑤, the stopping condition is that the number of repetitions reaches a preset number or the cluster allocation no longer changes.
[0022] Furthermore, the preset number of times is 20 times.
[0023] Furthermore, in step ⑦, L≥1; Knew≥K. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION
[0025] The present invention is further described in detail below with reference to the accompanying drawings.
[0026] A multi-connection clustering method for data center clusters includes the following steps:
[0027] Step ①, randomly select N monitoring data from the monitoring system of the data center cluster application to form a data set, and normalize the data set to obtain a normalized data set; each monitoring data includes multiple indicator data; the indicator data includes system resource utilization indicators such as CPU, memory, and disk IO;
[0028] The specific operation of normalizing the data set to obtain the normalized data set is as follows: using the minimum and maximum normalization method to scale each indicator data in the data set to the range of 0-1 to obtain the normalized data set;
[0029] Step ②, perform initial clustering on the normalized data set to obtain K clusters:
[0030] Step ②-1, randomly select K monitoring data from the normalized data set as the initial cluster center, K ≥ 2;
[0031] Step ②-2, obtain the Euclidean distance between each monitoring data in the normalized data set and the centers of the K initial clusters, and assign each monitoring data in the normalized data set to the cluster where the initial cluster center with the closest Euclidean distance is located according to the obtained Euclidean distance, to obtain K clusters;
[0032] Step ③, for K clusters, obtain the average value of all monitoring data in each cluster, and use the average value as the new cluster center;
[0033] Step ④, obtain the Euclidean distance between each monitoring data in the normalized data set and the new cluster center, and assign each monitoring data in the normalized data set to the cluster where the new cluster center with the closest Euclidean distance is located according to the obtained Euclidean distance, and obtain K clusters;
[0034] Step ⑤, repeat steps ③ to ④ until the stop condition is met, and then execute step ⑥; the stop condition is that the number of repetitions reaches the preset number of times or the cluster allocation does not change (that is, the average value calculated currently is the cluster center of the previous time); the preset number of times is 20;
[0035] Step ⑥, numbering each cluster obtained in step ⑤, obtaining and recording the cluster number of each monitoring data; wherein the clusters are numbered starting from number 0 until number K-1, for example: the cluster number of the first cluster is 0, then the cluster numbers of all monitoring data in the first cluster are 0;
[0036] Step ⑦, repeat steps ② to ⑥ until the number of repetitions reaches L (L ≥ 1) times. Each monitoring data is given L cluster numbers. The time obtained according to the L cluster numbers of each monitoring data is aggregated into a multi-group in chronological order. The monitoring data corresponding to the same multi-group are classified into the same category, and Knew (Knew ≥ K) categories are obtained to complete the clustering of the data center cluster.
[0037] Example: Randomly select N = 4 monitoring data from the monitoring system of the data center cluster monitoring data to form a data set, namely monitoring data a, monitoring data b, monitoring data c and monitoring data d; set K = 2, L = 3;
[0038] The L clusters obtained from monitoring data a are numbered as: 0, 1, 1; the L clusters obtained from monitoring data b are numbered as: 1, 0, 0; the L clusters obtained from monitoring data c are numbered as: 0, 1, 0; the L clusters obtained from monitoring data d are numbered as: 1, 0, 0;
[0039] After aggregation, the multi-tuple obtained from monitoring data a is (0, 1, 1), the multi-tuple obtained from monitoring data b is (1, 0, 0), the multi-tuple obtained from monitoring data c is (0, 1, 0), and the multi-tuple obtained from monitoring data d is (1, 0, 0);
[0040] Therefore, monitoring data b and monitoring data d are classified into the same category, monitoring data a is one category, and monitoring data c is one category, and Knew=3 categories are obtained.
Claims
1. A multi-connection clustering method for data center clusters, characterized by The following steps are involved: Step ①, randomly selecting N monitoring data from the monitoring system of the data center cluster application to form a data set, and normalizing the data set to obtain a normalized data set; Step ②, perform initial clustering on the normalized data set to obtain K clusters; Step ③, for K clusters, obtain the average value of all monitoring data in each cluster, and use the average value as the new cluster center; Step ④, obtain the Euclidean distance between each monitoring data in the normalized data set and the new cluster center, and assign each monitoring data in the normalized data set to the cluster where the new cluster center with the closest Euclidean distance is located according to the obtained Euclidean distance, and obtain K clusters; Step ⑤, repeat steps ③ to ④ until the stop condition is met, and then execute step ⑥; Step ⑥, numbering each cluster obtained in step ⑤, obtaining and recording the cluster number of each monitoring data; Step ⑦, repeat steps ② to ⑥ until the number of repetitions reaches L times. Each monitoring data is given L cluster numbers. The time obtained according to the L cluster numbers of each monitoring data is aggregated into a multi-group in chronological order. The monitoring data corresponding to the same multi-group are classified into the same category, and Knew categories are obtained to complete the clustering of the data center cluster.
2. A multi-connection clustering method for data center clusters according to claim 1, characterized in that In step ①, each of the monitoring data includes multiple indicator data; The indicator data includes CPU, memory, and disk IO; The specific operation of normalizing the data set to obtain the normalized data set is: using the minimum and maximum normalization method to scale each indicator data in the data set to the range of 0-1 to obtain the normalized data set.
3. A multi-connection clustering method for data center clusters according to claim 1, characterized in that The specific operation process of step ② is as follows: Step ②-1, randomly select K monitoring data from the normalized data set as the initial cluster center; Step ②-2, obtain the Euclidean distance between each monitoring data in the normalized data set and the K initial cluster centers respectively, and assign each monitoring data in the normalized data set to the cluster where the initial cluster center with the closest Euclidean distance is located according to the obtained Euclidean distance, and obtain K clusters.
4. A multi-connection clustering method for data center clusters according to claim 3, characterized in that K≥2。 5. A multi-connection clustering method for data center clusters according to claim 1, characterized in that In step ⑤, the stopping condition is that the number of repetitions reaches a preset number or the cluster allocation no longer changes.
6. A multi-connection clustering method for data center clusters according to claim 5, characterized in that The preset number of times is 20 times.
7. A multi-connection clustering method for data center clusters according to claim 1, characterized in that In step ⑦, L≥1; Knew≥K.
Citation Information
Patent Citations
Network flow time sequence prediction method based on distributed clustering
CN107067028A
Big data clustering method and device based on distributed structure
CN108717444A
Privacy information protection method based on K-means clustering
CN110233730A
Differential privacy k-means clustering method based on cluster similarity and transformation invariance
CN112364914A