A Server Cluster Deployment Monitoring System and Method Based on Big Data

Through the clustered deployment monitoring system of big data and machine learning, the flexibility and adaptability of server clustered deployment are solved, accurate prediction and dynamic monitoring of new requirements are achieved, and operation and maintenance efficiency is improved.

CN119003277BActive Publication Date: 2025-07-22上海市大数据中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411155426.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-07-22
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

In the prior art, server cluster deployment lacks flexibility and adaptability, deployment solutions are difficult to adapt to complex and changing needs, and the deployment effect feedback mechanism is imperfect, and the ability to continuously improve is lacking.

Method used

A clustered server deployment monitoring system based on big data is adopted, including record and screening modules, architecture prediction modules, component selection modules and monitoring and early warning modules. Using machine learning models and similarity analysis, high-quality deployment architectures are screened through historical records for prediction and component selection, and dynamic monitoring and early warning are carried out in combination with time series prediction.

Benefits of technology

Accurate prediction of new deployment requirements is achieved, manual intervention is reduced, operational processes are simplified, operation and maintenance efficiency is improved, and complex and changing environmental needs can be better adapted to complex and changing environmental needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119003277B_ABST
    Figure CN119003277B_ABST
Patent Text Reader

Abstract

The present invention discloses a monitoring system and method for server cluster deployment based on big data, which relates to the technical field of big data. The system includes a record screening module, an architecture prediction module, a component selection module, and a monitoring and warning module; the record screening module is used to collect and screen server cluster deployment records through index effects; the architecture prediction module is used to build a model according to deployment requirements and deployment architectures to classify and predict the corresponding deployment architectures for new deployment requirements; the component selection module is used to select components according to the similarity between deployment requirements and new deployment requirements under the same architecture; the monitoring and warning module is used to deploy servers in a cluster according to the selected architecture and components and perform monitoring and prediction; the present invention also provides a method for implementing the system. The present invention can effectively improve the situation that the deployment scheme in the prior art lacks flexibility and adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and specifically to a monitoring system and method for server cluster deployment based on big data. Background Art

[0002] Server cluster deployment is a way of combining multiple servers to work together. It can significantly improve performance, distribute the workload to each server, and speed up the processing speed; it can also enhance availability, and a failure of one server does not affect the overall service; in addition, cluster deployment is convenient for expansion, and new servers can be added as the business grows. Common types include load balancing, high availability, and high-performance computing clusters; factors such as hardware, network, and algorithms need to be considered during deployment, which is an important way to build a powerful IT infrastructure.

[0003] Server cluster deployment usually relies on expert experience and fixed rules for decision-making, which may lead to the lack of flexibility and adaptability of the deployment plan. Especially when facing complex and changing requirements, it is easy to generate inapplicable or inefficient deployment plans; in addition, in the prior art, the effect feedback mechanism after deployment is usually imperfect, it is difficult to guide future deployments through historical deployment effects, and it lacks the ability of continuous improvement. Summary of the Invention

[0004] The purpose of the present invention is to provide a monitoring system and method for server cluster deployment based on big data to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A monitoring system for server cluster deployment based on big data, the system includes a record screening module, an architecture prediction module, a component selection module, and a monitoring and warning module;

[0006] The record screening module is used to collect server cluster deployment records and screen server cluster deployment records through index effects; the architecture prediction module is used to build a model according to deployment requirements and deployment architectures to classify and predict the corresponding deployment architectures for new deployment requirements; the component selection module is used to select components according to the similarity between deployment requirements and new deployment requirements under the same architecture; the monitoring and warning module is used to deploy servers in a cluster according to the selected architecture and components and conduct monitoring and prediction;

[0007] The output end of the record screening module is connected to the input end of the architecture prediction module; the output end of the architecture prediction module is connected to the input end of the component selection module; the output end of the component selection module is connected to the input end of the monitoring and warning module.

[0008] The record screening module includes a record collection unit, an index classification unit, an effect evaluation unit, and a record screening unit;

[0009] The record collection unit is used to collect the historical server clustering deployment records and new deployment requirements; the metric classification unit is used to classify the collected metric data; the effect evaluation unit is used to evaluate the clustering deployment effect according to the metric data; the record screening unit is used to set a threshold and screen out the records whose effects are not less than the threshold;

[0010] The output end of the record collection unit is connected to the input end of the metric classification unit; the output end of the metric classification unit is connected to the input end of the effect evaluation unit; the output end of the effect evaluation unit is connected to the input end of the record screening unit; the output end of the record screening unit is connected to the input end of the architecture prediction module.

[0011] The architecture prediction module includes a model construction unit, a model training unit and a classification prediction unit;

[0012] The model construction unit is used to construct a model according to the deployment requirements and their corresponding deployment architectures; the model training unit is used to train the model; the classification prediction unit is used to classify and predict the new deployment requirements according to the trained model;

[0013] The output end of the model construction unit is connected to the input end of the model training unit; the output end of the model training unit is connected to the input end of the classification prediction unit; the output end of the classification prediction unit is connected to the input end of the component selection module.

[0014] The component selection module includes a similarity analysis unit and a component selection unit;

[0015] The similarity analysis unit is used to calculate the cosine similarity between the new deployment requirements and the deployment requirements of each record corresponding to the classified prediction architecture; the component selection unit is used to select the components corresponding to the record with the highest cosine similarity value as the components used for the new deployment requirements deployment;

[0016] The output end of the similarity analysis unit is connected to the input end of the component selection unit; the output end of the component selection unit is connected to the input end of the monitoring and warning module.

[0017] The monitoring and warning module includes a deployment unit, a monitoring unit and a warning unit;

[0018] The deployment unit is used to perform clustering deployment on the servers according to the classified prediction architecture and the selected components; the monitoring unit is used to regularly collect the metric data of the server cluster; the prediction unit is used to establish a time series prediction model according to the collected metric data for prediction and set a threshold for warning;

[0019] The output end of the deployment unit is connected to the input end of the monitoring unit; the output end of the monitoring unit is connected to the input end of the prediction unit.

[0020] A method for monitoring the clustering deployment of a server based on big data, the method comprising the following steps:

[0021] Step1. Collect server clustering deployment records from historical deployments and corresponding logs;

[0022] Step2. Set a threshold a, evaluate the deployment effect of each record, and filter out records whose deployment effect is not less than a;

[0023] Step3. For the filtered records, construct a machine learning model according to their requirements and the selected architecture, and use the model to predict new requirements;

[0024] Step4. Obtain the architecture components corresponding to the record with the highest similarity to the new deployment requirements from all records corresponding to the selected architecture as the components of the new deployment requirements;

[0025] S5. Perform clustering deployment on the server according to the obtained architecture and architecture components, monitor the load situation of the cluster and make predictions.

[0026] In step Step1, each of the server clustering deployment records includes: deployment requirements, deployment architecture, architecture components, and metric logs;

[0027] The deployment requirements refer to the business requirements or performance goals that the system or application needs to meet during the clustering deployment of the server; expressed as: [R1, R2, …, R r ; where r is a positive integer representing the number of deployment requirements, and R r represents the rth deployment requirement;

[0028] The deployment architecture refers to the overall design and structural arrangement for implementing the deployment requirements, including the organization method of the server cluster, the distribution method of services, and the network topology structure; expressed as: [S1, S2, …, S h ; where h is a positive integer representing the number of deployment architectures, and S h represents the hth deployment architecture;

[0029] The architecture components refer to the specific hardware and software parts that make up the deployment architecture; these components are the basic elements for actually constructing and running the deployment architecture; expressed as: [M1, M2, …, M m ; where m is a positive integer representing the number of architecture components, and M m represents the mth architecture component;

[0030] The index log refers to the record of the performance data and running status of the server cluster system collected over a period of time; the time is expressed as: [T1, T2, …, T t ; each collection of indexes is expressed as: [D1, D2, …, D d ; where t is a positive integer representing the number of times of collecting index data, and T t represents the time of collecting index data for the t-th time; d is a positive integer representing the number of indexes, and D d represents the d-th index;

[0031] A complete record of server cluster deployment is expressed as: [R1, R2, …, R r , S b , M1, M2, …, M m , [(T1)D1, D2, …, D d , [(T2)D1, D2, …, D d , …, [(T t )D1, D2, …, D d ; where S b ∈{S1, S2, …, S h}, indicating that each record corresponds to only one deployment architecture;

[0032] Among them, [(T1)D1, D2, …, D d , [(T2)D1, D2, …, D d , …, [(T t )D1, D2, …, D d respectively represent the D1 - D t index data corresponding to the time of T1 - T d .

[0033] In step Step2, the threshold a represents an effect, where 0 < a < 100%;

[0034] The specific method for evaluating the deployment effect of each record is as follows: calculate the deployment effect based on the index data; divide the indexes into best - range indexes and extreme - optimization indexes; the best - range indexes indicate that the indexes have the best effect within a specific range; the extreme - optimization indexes indicate that the higher or lower the index, the better the effect;

[0035] For the best - range indexes, set the upper - limit threshold a1 and the lower - limit threshold a2, where a1 > a2; set the effect at (a1 + a2) / 2 to be 100%, and set the effect at the thresholds a1 and a2 to be a; when the index data V(D i ) > a1 or V(D i ) < a2, set the effect to 0; when the index data a2 < V(D i)When a1, the effect E(D i ) can be expressed as: E(D i ) = 1 - |V(D i ) - (a1 + a2) / 2| / (a1 - (a1 + a2) / 2) × (1 - a);

[0036] Among them, i is a positive integer, i ∈ {1, 2, …, d}; V(D i ) represents the value of the i-th index D i ; E(D i ) represents the effect of the i-th index D i ;

[0037] For extreme optimization type indicators, they are divided into high-end optimization type indicators and low-end optimization type indicators; for high-end optimization type indicators, set the lower limit threshold a 31 , high-end threshold a 32 , a 32 > a 31 , set the effect at the high-end threshold a 32 to be 100%, and set the effect at the lower limit threshold a 31 to be a; when the indicator data V(D j ) < a 31 , the effect is 0; when the indicator data V(D j ) > a 32 , the effect is 100%; when the indicator data a 31 < V(D j ) < a 32 , the effect E(D j ) can be expressed as:

[0038] E(D j ) = a + (V(D j ) - a 31 ) / (a 32 - a 31 ) × (1 - a);

[0039] Among them, j is a positive integer, j ∈ {1, 2, …, d}; V(D j ) represents the value of the j-th index D j ; E(D j ) represents the effect of the j-th index D j ;

[0040] For low-end optimization type indicators, set the upper limit threshold a4, set the effect to be 100% when the indicator data is 0, and set the effect at the threshold a4 to be a; when the indicator data V(D k ) > a4, the effect is set to 0; when the indicator data 0 < V(D k ) < a4, the effect E(D k) can be expressed as: E(D k ) = 1 - V(D k ) / a4×(1 - a);

[0041] where k is a positive integer, k ∈ {1, 2, …, d}; V(D k ) represents the value of the k-th index D k ; E(D k ) represents the effect of the k-th index D k ;

[0042] Calculate the average effect of each index over a period of time:

[0043] E mean (D g ) = (E T1 (D g ) + E T2 (D g ) + … + E Tt (D g )) / t;

[0044] where g is a positive integer, g ∈ {1, 2, …, d}; E T1 (D g ) ~ E Tt (D g ) respectively represent the effects of the g-th index D t at times T1 to T g ; E mean (D g ) represents the average effect of the g-th index D g over t times;

[0045] Calculate the comprehensive effect E(D) of all indexes:

[0046] E(D) = (E mean (D1) + E mean (D2) + … + E mean (D d )) / d, and filter out the records with a deployment effect not less than a.

[0047] In step Step3, for the filtered records, construct a machine learning model according to their requirements and the selected architecture;

[0048] Data processing: For the input feature deployment requirement values [R1, R2, …, R r and the output label deployment architectures [S1, S2, …, S h , perform One-Hot encoding on the categorical features to convert them into binary features, and standardize the numerical features;

[0049] The input feature matrix is represented as:

[0050] ;

[0051] where X represents the input feature matrix, with a size of n×r; n is a positive integer representing the number of records; r is the number of features, and R nr represents the r-th deployment requirement feature value of the n-th record after numericalization and standardization;

[0052] The output label matrix is represented as:

[0053] ;

[0054] where Y represents the output label matrix, with a size of n×1; n is a positive integer representing the number of records; S(1), …, S(n) respectively represent the deployment architecture labels corresponding to the deployment requirements R 1r , …, R nr after numericalization and standardization;

[0055] Use the machine learning algorithm LightGBM to build and train a model for deployment requirements and their corresponding deployment architectures;

[0056] Model building: The model predicts the output label matrix Y through the input feature matrix X, represented as the function f: Y = f(X; θ);

[0057] where f is the functional representation form of the LightGBM model; θ is the set of model parameters;

[0058] The model learns the relationship between the input and output by minimizing the loss function:

[0059] ;

[0060] where L is the loss function, measuring the error between the model's predicted value and the true value; K ∈ {1, 2, …, h}, K is the number of all possible classes; p ∈ {1, 2, …, n}, q ∈ {1, 2, …, K}, and Y pq represents the true class label of the p-th record, which is 1 when it belongs to the q-th class and 0 otherwise; represents the probability that the model predicts the p-th record belongs to the q-th class;

[0061] LightGBM trains multiple decision trees in an iterative manner, and each tree attempts to correct the errors of the previous tree:

[0062] ;

[0063] where, denotes the model prediction result after the e-th iteration; η represents the learning rate, controlling the contribution of each tree to the model prediction; f e (X; θ e ) is the e-th decision tree, which makes predictions based on the input feature matrix X;

[0064] For the new deployment requirement feature vector X′, the model predicts the most suitable deployment architecture label S′: S′ = f(X′; θ).

[0065] In step Step4, the component of the new deployment requirement is obtained by taking the architecture component corresponding to the record with the highest similarity to the new deployment requirement from all the records corresponding to the selected architecture. The specific method is as follows:

[0066] Among the records corresponding to the deployment architecture obtained by classifying the new deployment requirement, the deployment requirements of each record are numerically processed and standardized to obtain the deployment requirement vector of each record; the new deployment requirement is numerically processed and standardized to obtain the new deployment requirement vector; the cosine similarity between the new deployment requirement vector and the deployment requirement vector of each record is calculated, and the architecture component corresponding to the record with the highest cosine similarity value is selected as the component of the new deployment requirement;

[0067] In step Step5, the specific method for predicting the load situation of the server cluster is as follows:

[0068] Regularly collect various indicators of the server cluster and form an indicator log, and use a time series prediction model to predict the trend of each indicator; set the threshold for each indicator, and when there is an indicator prediction value exceeding the corresponding set threshold, send a warning to the management personnel.

[0069] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0070] 1. By introducing a machine learning model, the present invention can accurately predict new deployment requirements based on high-quality deployment architectures selected from historical records;

[0071] 2. Through an automated decision-making process integrating machine learning and similarity matching, the present invention reduces the need for manual intervention, simplifies the operation process of server cluster deployment, and improves the overall operation and maintenance efficiency;

[0072] 3. The present invention makes full use of big data for dynamic deployment effect evaluation and prediction, enabling the system to better adapt to complex and changing environmental requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 is a schematic flow chart of a server cluster deployment monitoring system based on big data according to the present invention;

[0074] Figure 2Schematic diagram of the steps of a method for monitoring the clustered deployment of servers based on big data according to the present invention. Specific implementation mode

[0075] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0076] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical solution, a monitoring system for the clustered deployment of servers based on big data, the system includes a record screening module, an architecture prediction module, a component selection module and a monitoring and warning module;

[0077] The record screening module is used to collect the records of the clustered deployment of servers and screen the records of the clustered deployment of servers through index effects; the architecture prediction module is used to build a model according to the deployment requirements and deployment architectures to classify and predict the corresponding deployment architectures for new deployment requirements; the component selection module is used to select components according to the similarity between the deployment requirements and new deployment requirements under the same architecture; the monitoring and warning module is used to deploy servers in a clustered manner according to the selected architecture and components and perform monitoring and prediction;

[0078] The output end of the record screening module is connected to the input end of the architecture prediction module; the output end of the architecture prediction module is connected to the input end of the component selection module; the output end of the component selection module is connected to the input end of the monitoring and warning module.

[0079] The record screening module includes a record collection unit, an index classification unit, an effect evaluation unit and a record screening unit;

[0080] The record collection unit is used to collect the historical records of the clustered deployment of servers and new deployment requirements; the index classification unit is used to classify the collected index data; the effect evaluation unit is used to evaluate the effect of the clustered deployment according to the index data; the record screening unit is used to set a threshold and screen out the records whose effects are not less than the threshold;

[0081] The output end of the record collection unit is connected to the input end of the index classification unit; the output end of the index classification unit is connected to the input end of the effect evaluation unit; the output end of the effect evaluation unit is connected to the input end of the record screening unit; the output end of the record screening unit is connected to the input end of the architecture prediction module.

[0082] The architecture prediction module includes a model construction unit, a model training unit and a classification prediction unit;

[0083] The model construction unit is used to construct a model according to the deployment requirements and their corresponding deployment architectures; the model training unit is used to train the model; the classification prediction unit is used to classify and predict new deployment requirements according to the trained model;

[0084] The output end of the model construction unit is connected to the input end of the model training unit; the output end of the model training unit is connected to the input end of the classification prediction unit; the output end of the classification prediction unit is connected to the input end of the component selection module.

[0085] The component selection module includes a similarity analysis unit and a component selection unit;

[0086] The similarity analysis unit is used to calculate the cosine similarity between the new deployment requirements and the deployment requirements of each record corresponding to the classification prediction architecture; the component selection unit is used to select the component corresponding to the record with the highest cosine similarity value as the component used for the new deployment requirements deployment;

[0087] The output end of the similarity analysis unit is connected to the input end of the component selection unit; the output end of the component selection unit is connected to the input end of the monitoring and warning module.

[0088] The monitoring and warning module includes a deployment unit, a monitoring unit and a warning unit;

[0089] The deployment unit is used to perform cluster deployment on the server according to the classification prediction architecture and the selected components; the monitoring unit is used to regularly collect the metric data of the server cluster; the prediction unit is used to establish a time series prediction model according to the collected metric data for prediction and set a threshold for warning;

[0090] The output end of the deployment unit is connected to the input end of the monitoring unit; the output end of the monitoring unit is connected to the input end of the prediction unit.

[0091] A method for monitoring the cluster deployment of servers based on big data, the method comprising the following steps:

[0092] Step1. Collect server cluster deployment records from historical deployments and corresponding logs;

[0093] Each of the server cluster deployment records includes: deployment requirements, deployment architecture, architecture components, and metric logs;

[0094] The deployment requirements refer to the business requirements or performance goals that the system or application needs to meet during the server cluster deployment process; expressed as: [system availability, response time];

[0095] The deployment architecture refers to the overall design and structural arrangement for realizing deployment requirements, including the organization mode of the server cluster, the distribution mode of services, and the network topology structure; expressed as: [Load Balancing Architecture, Multi-Region Distributed Architecture];

[0096] The architecture components refer to the specific hardware and software parts that make up the deployment architecture; these components are the basic elements for actually building and running the deployment architecture; expressed as: [Cache Server, Database Cluster];

[0097] The metric log refers to the record of the server cluster system performance data and running status collected over a period of time; the time is expressed as: [T1, T2]; the metrics collected each time are expressed as: [Memory Utilization, Data Throughput];

[0098] A complete record of server cluster deployment is expressed as: [System Availability, Response Time, Load Balancing Architecture / Multi-Region Distributed Architecture, Cache Server, Database Cluster, (T1) Memory Utilization, Data Throughput, (T2) Memory Utilization, Data Throughput];

[0099] Record 1: [System Availability: 90%, Response Time: 200ms, Load Balancing Architecture, Cache Server 1, Database Cluster 1, (T1) Memory Utilization: 60%, Data Throughput 60, (T2) Memory Utilization: 70%, Data Throughput: 70];

[0100] Record 2: [System Availability: 99%, Response Time: 300ms, Multi-Region Distributed Architecture, Cache Server 2, Database Cluster 2, (T1) Memory Utilization: 50%, Data Throughput: 40, (T2) Memory Utilization: 60%, Data Throughput: 50].

[0101] Step 2, set the threshold at 60%, evaluate the deployment effect of each record, and filter out the records with a deployment effect not less than 60%;

[0102] Calculate the deployment effect based on the metric data; classify the metrics into best-range metrics and extreme optimization metrics; the best-range metrics indicate that the metric has the best effect within a specific range; the extreme optimization metrics indicate that the higher or lower the metric, the better the effect;

[0103] For the memory utilization rate, which is an indicator within the optimal range, set the upper threshold at 70% and the lower threshold at 50%; set the effect at 100% when it is at (70% + 50%) / 2 = 60%, and set the effect at 60% when it is at the thresholds of 50% and 70%; when the indicator data V (memory utilization rate) > 70% or V (memory utilization rate) < 50%, set the effect to 0; when the indicator data 50% < V (memory utilization rate) < 70%, the effect E (memory utilization rate) can be expressed as: E (memory utilization rate) = 1 - |V (memory utilization rate) - 60%| / (70% - 60%) × (1 - 60%);

[0104] For the indicators of extreme optimization, divide them into high-end optimization indicators and low-end optimization indicators; for the high-end optimization indicator of data throughput, set the lower threshold at 20 and the upper threshold at 80, set the effect at 100% when it is at the upper threshold of 80, and set the effect at 60% when it is at the lower threshold of 20; when the indicator data V (data throughput) < 20, the effect is 0; when the indicator data V (data throughput) > 80, the effect is 100%; when the indicator data 20 < V (data throughput) < 80, the effect E (data throughput) can be expressed as: E (data throughput) = 60% + (V (data throughput) - 20) / (80 - 20) × (1 - 60%);

[0105] Calculate the average effect of each indicator over a period of time:

[0106] Record 1 memory utilization rate T1: 100%; T2: 60%; E mean = 80%;

[0107] Record 1 data throughput T1: 86.6%; T2: 93.3%; E mean = 89.95%;

[0108] Record 2 memory utilization rate T1: 60%; T2: 100%; E mean = 80%;

[0109] Record 2 data throughput T1: 73.3%; T2: 80%; E mean = 76.65%;

[0110] Calculate the comprehensive effect of all indicators:

[0111] Record 1: E = 84.975%; Record 2: E = 78.325%;

[0112] Filter out the records with a deployment effect not less than 60%. The effects of both Record 1 and Record 2 are greater than 60%.

[0113] Step 3. For the filtered records, construct a machine learning model according to their requirements and the selected architecture, and use the model to predict new requirements;

[0114] In Step 3, for the filtered records, a machine learning model is constructed according to their requirements and the selected architecture.

[0115] Data processing: For the input feature deployment requirements and the output label deployment architecture, the categorical features are converted into binary features by One-Hot encoding, and the numerical features are standardized.

[0116] Use the machine learning algorithm LightGBM to construct and train a model for the deployment requirements and their corresponding deployment architectures.

[0117] Model construction: The model predicts the output label through the input feature matrix; the model learns the relationship between the input and output by minimizing the loss function; LightGBM trains multiple decision trees iteratively, and each tree tries to correct the mistakes of the previous tree.

[0118] For new deployment requirements, the model predicts the most suitable deployment architecture label.

[0119] Step 4: Obtain the architecture components corresponding to the record with the highest similarity to the new deployment requirement from all the records corresponding to the selected architecture as the components of the new deployment requirement.

[0120] In Step 4, the specific method of obtaining the architecture components corresponding to the record with the highest similarity to the new deployment requirement from all the records corresponding to the selected architecture as the components of the new deployment requirement is as follows:

[0121] Among the records corresponding to the deployment architecture obtained by classifying the new deployment requirements, the deployment requirements of each record are numerically processed and standardized to obtain the deployment requirement vector of each record; the new deployment requirement is numerically processed and standardized to obtain the new deployment requirement vector; calculate the cosine similarity between the new deployment requirement vector and the deployment requirement vector of each record, and select the architecture components corresponding to the record with the highest cosine similarity value as the components of the new deployment requirement.

[0122] Step 5: Cluster the servers according to the obtained architecture and architecture components, monitor the load situation of the cluster and make predictions.

[0123] In Step 5, the specific method of predicting the load situation of the server cluster is as follows:

[0124] Regularly collect various indicators of the server cluster and form an indicator log, and use a time series prediction model to predict the trends of each indicator; set the thresholds for each indicator, and send a warning to the management personnel when there is an indicator prediction value exceeding the corresponding set threshold.

[0125] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A monitoring system for server cluster deployment based on big data, characterized in that: The system includes a record screening module, an architecture prediction module, a component selection module, and a monitoring and warning module; The record screening module is used to collect server cluster deployment records and screen server cluster deployment records through metric effects; the architecture prediction module is used to build a model based on deployment requirements and deployment architectures to classify and predict the corresponding deployment architectures for new deployment requirements; The component selection module is used to select components according to the similarity between deployment requirements and new deployment requirements under the same architecture; the monitoring and warning module is used to deploy servers in a cluster according to the selected architecture and components and perform monitoring and prediction; The output end of the record screening module is connected to the input end of the architecture prediction module; the output end of the architecture prediction module is connected to the input end of the component selection module; the output end of the component selection module is connected to the input end of the monitoring and warning module; The record screening module includes a record collection unit, a metric classification unit, an effect evaluation unit, and a record screening unit; The record collection unit is used to collect historical server cluster deployment records and new deployment requirements; the metric classification unit is used to classify the collected metric data; the effect evaluation unit is used to evaluate the cluster deployment effect according to the metric data; The record screening unit is used to set a threshold and screen out records whose effects are not less than the threshold; The output end of the record collection unit is connected to the input end of the metric classification unit; the output end of the metric classification unit is connected to the input end of the effect evaluation unit; the output end of the effect evaluation unit is connected to the input end of the record screening unit; the output end of the record screening unit is connected to the input end of the architecture prediction module; The architecture prediction module includes a model construction unit, a model training unit, and a classification prediction unit; The model construction unit is used to build a model according to deployment requirements and their corresponding deployment architectures; the model training unit is used to train the model; the classification prediction unit is used to classify and predict new deployment requirements according to the trained model; The output end of the model construction unit is connected to the input end of the model training unit; the output end of the model training unit is connected to the input end of the classification prediction unit; the output end of the classification prediction unit is connected to the input end of the component selection module; The component selection module includes a similarity analysis unit and a component selection unit; The similarity analysis unit is used to calculate the cosine similarity between the new deployment requirements and the deployment requirements of each record corresponding to the classified and predicted architecture; the component selection unit is used to select the components corresponding to the record with the highest cosine similarity value as the components used for the new deployment requirements deployment; The output end of the similarity analysis unit is connected to the input end of the component selection unit; the output end of the component selection unit is connected to the input end of the monitoring and warning module.

2. The monitoring system for server cluster deployment based on big data according to claim 1, wherein: The monitoring and warning module includes a deployment unit, a monitoring unit, and a warning unit; The deployment unit is used to cluster-deploy the servers according to the classification prediction architecture and the selected components; the monitoring unit is used to regularly collect the metric data of the server cluster; the prediction unit is used to establish a time series prediction model based on the collected metric data for prediction and set a threshold for early warning; The output end of the deployment unit is connected to the input end of the monitoring unit; the output end of the monitoring unit is connected to the input end of the prediction unit.

3. A monitoring method for server cluster deployment based on big data, which applies the monitoring system for server cluster deployment based on big data described in any one of claims 1-2, and is characterized in that: The method includes the following steps: Step1: Collect the server cluster deployment records from the historical deployments and the corresponding logs; Step2: Set a threshold a, evaluate the deployment effect of each record, and filter out the records whose deployment effect is not less than a; Step3: For the filtered records, build a machine learning model according to their requirements and the selected architecture, and use the model to predict new requirements; Step4: Obtain the architecture components corresponding to the record with the highest similarity to the new deployment requirement among all the records corresponding to the selected architecture as the components of the new deployment requirement; Step5: Cluster-deploy the servers according to the obtained architecture and architecture components, monitor the load situation of the cluster and make predictions.

4. A monitoring method for server cluster deployment based on big data according to claim 3, characterized in that: In step Step1, each of the server cluster deployment records includes: deployment requirements, deployment architecture, architecture components, and metric logs; The deployment requirements refer to the business requirements or performance goals that the system or application needs to meet during the clustered deployment of servers; expressed as: [R1, R2, …, R r ; where r is a positive integer representing the number of deployment requirements, and R r represents the r-th deployment requirement; The deployment architecture refers to the overall design and structural arrangement for realizing deployment requirements, including the organization mode of server clusters, the distribution mode of services, and the network topology structure; expressed as: [S1, S2, …, S h ; where h is a positive integer representing the number of deployment architectures, and S h represents the h-th deployment architecture; The architecture components refer to the specific hardware and software parts that make up the deployment architecture; these components are the basic elements for actually building and running the deployment architecture; expressed as: [M1, M2, …, M m ; where m is a positive integer representing the number of architecture components, and M m represents the m-th architecture component; The said metric log refers to the record of the performance data and running status of the server cluster system collected over a period of time; the time is represented as: [T1, T2, …, T t ; the metrics collected each time are represented as: [D1, D2, …, D d ; where t is a positive integer representing the number of times of collecting metric data, and T t represents the time of collecting metric data for the t-th time; d is a positive integer representing the number of metrics, and D d represents the d-th metric; A complete record of server cluster deployment is represented as: [R1, R2, …, R r , S b , M1, M2, …, M m , [(T1)D1, D2, …, D d , [(T2)D1, D2, …, D d , …, [(T t )D1, D2, …, D d ; where S b ∈{S1, S2, …, S h}, indicating that each record corresponds to only one deployment architecture; Among them, [(T1)D1, D2, …, D d , [(T2)D1, D2, …, D d , …, [(T t )D1, D2, …, D d respectively represent the D1~D t index data corresponding to the moments of T1~T d .

5. The monitoring method for server cluster deployment based on big data according to claim 4, characterized in that: In step Step2, the threshold a represents an effect, where 0 < a < 100%; The specific method for evaluating the deployment effect of each record is as follows: calculate the deployment effect according to the metric data; classify the metrics into best-range metrics and extreme optimization metrics; the best-range metrics indicate that the metric has the best effect within a specific range; the extreme optimization metrics indicate that the higher or lower the metric, the better the effect; For the best range type of indicators, set the upper threshold a1 and the lower threshold a2, where a1 > a2; set the effect at (a1 + a2) / 2 to be 100%, and set the effect at the thresholds a1 and a2 to be a; when the indicator data V(D i ) > a1 or V(D i ) < a2, set the effect to 0; when the indicator data a2 < V(D i ) < a1, the effect E(D i ) is expressed as: E(D i ) = 1 - |V(D i ) - (a1 + a2) / 2| / (a1 - (a1 + a2) / 2) × (1 - a); where \(i\) is a positive integer, \(i\in\{1,2,\ldots,d\}\); \(V(D i )\) represents the value of the \(i\)-th index \(D i \); \(E(D i )\) represents the effect of the \(i\)-th index \(D i \). For extreme optimization - type metrics, they are divided into high - end optimization - type metrics and low - end optimization - type metrics; for high - end optimization - type metrics, a lower - limit threshold \(a\) is set 31 and a high - end threshold \(a\) 32 , where \(a\) 32 \(>\ a\) 31 . The effect at the high - end threshold \(a\) 32 is set to 100%, and the effect at the lower - limit threshold \(a\) 31 is set to \(a\); when the metric data \(V(D\) j ) \(< a\) 31 , the effect is 0; when the metric data \(V(D\) j ) \(> a\) 32 , the effect is 100%; when the metric data \(a\) 31 \(< V(D\) j ) \(< a\) 32 , the effect \(E(D\) j ) is expressed as: E(D j ) = a+(V(D j ) - a 31 ) / (a 32 - a 31 )×(1 - a); where j is a positive integer, j ∈ {1, 2, …, d}; V(D j ) represents the value of the j-th indicator D j ; E(D j ) represents the effect of the j-th indicator D j . For low-end optimization metrics, set the upper limit threshold a4. When the metric data is 0, the effect is 100%. When the threshold is at a4, the effect is a. When the metric data V(D k ) > a4, the effect is set to 0. When the metric data 0 < V(D k ) < a4, the effect E(D k ) is expressed as: E(D k ) = 1 - V(D k ) / a4 × (1 - a); where k is a positive integer, k ∈ {1, 2, …, d}; V(D k ) represents the value of the k-th index D k ; E(D k ) represents the effect of the k-th index D k ; Calculate the average effect of each metric over a period of time: E mean (D g ) = (E T1 (D g ) + E T2 (D g ) + … + E Tt (D g )) / t; where g is a positive integer, g ∈ {1, 2, …, d}; E T1 (D g ) ~ E Tt (D g ) represent the effects of the g-th index D t at the T1 to T g moments respectively; E mean (D g ) represents the average effect of the g-th index D g at t moments; Calculate the comprehensive effect E(D) of all metrics: E(D)=(E mean (D1)+E mean (D2)+…+E mean (D d )) / d, and select the records with a deployment effect not less than a.

6. The method for monitoring the cluster deployment of a server based on big data according to claim 5, characterized in that: In step Step3, for the filtered records, build a machine learning model according to their requirements and the selected architecture; Data processing: For the input feature deployment requirement values [R1, R2, …, R r , and the output label deployment architecture [S1, S2, …, S h , perform One-Hot encoding on the categorical features among them to convert them into binary features, and standardize the numerical features; The input feature matrix is represented as: ; Among them, X represents the input feature matrix, with a size of n×r; n is a positive integer representing the number of records; r is the number of features, and R nr represents the r-th deployment requirement feature value of the n-th record after numericalization and standardization; The output label matrix is represented as: ; Among them, Y represents the output label matrix with a size of n×1; n is a positive integer representing the number of records; S(1), …, S(n) respectively represent the numericalized and standardized deployment architecture labels corresponding to the deployment requirement features R 1r , …, R nr corresponding to the numericalized and standardized deployment architecture labels; Use the machine learning algorithm LightGBM to build and train a model for the deployment requirements and their corresponding deployment architectures; Model construction: The model predicts the output label matrix Y through the input feature matrix X, represented as the function f: Y = f(X; θ); Among them, f is the functional representation form of the LightGBM model; θ is the parameter set of the model; The model learns the relationship between the input and output by minimizing the loss function: ; Among them, L is the loss function, which measures the error between the predicted value of the model and the true value; K ∈ {1, 2, …, h}, where K is the number of all possible classes; p ∈ {1, 2, …, n}, q ∈ {1, 2, …, K}, and Y pq represents the true class label of the p-th record. It is 1 when it belongs to the q-th class, and 0 otherwise; represents the probability that the p-th record predicted by the model belongs to the q-th class; LightGBM trains multiple decision trees iteratively, and each tree tries to correct the mistakes of the previous tree: ; Among them, represents the model prediction result after the e-th iteration; η represents the learning rate, which controls the contribution of each tree to the model prediction; f e (X; θ e ) is the e-th decision tree, which makes predictions based on the input feature matrix X; For the new deployment requirement feature vector X′, the model predicts the most suitable deployment architecture label S′: S′ = f(X′; θ).

7. A method for monitoring the cluster deployment of a server based on big data according to claim 6, characterized in that: In step Step4, the specific method for obtaining the architecture components corresponding to the record with the highest similarity to the new deployment requirement among all the records corresponding to the selected architecture as the components of the new deployment requirement is as follows: In the records corresponding to the deployment architecture obtained by classifying the new deployment requirements, numerical and standardized processing is performed on the deployment requirements of each record to obtain the deployment requirement vector of each record; numerical and standardized processing is performed on the new deployment requirements to obtain the new deployment requirement vector; the cosine similarity between the new deployment requirement vector and the deployment requirement vector of each record is calculated, and the architecture component corresponding to the record with the highest cosine similarity value is selected as the component of the new deployment requirement. In step Step5, the specific method for predicting the load situation of the server cluster is as follows: Regularly collect various indicators of the server cluster and form indicator logs, and use a time series prediction model to predict the trends of each indicator; set the threshold for each indicator, and when there is an indicator prediction value exceeding the corresponding set threshold, an alarm is sent to the management personnel.

Citation Information

Patent Citations

  • Cluster automation monitoring system and method

    CN110287079A

  • Demand scheme recommendation method and device for computing power network

    CN115358554A