Method and device for monitoring cluster, computer storage medium and electronic device
By receiving and predicting monitoring data, the monitoring cluster configuration of the cloud-native application cluster is dynamically adjusted, solving the problems of excessive pressure on the monitoring system and lagging expansion, and achieving efficient monitoring and resource optimization during peak business hours.
Patent Information
- Application Number
- CN202210844967.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2042-07-18
AI Technical Summary
The cloud-native application cluster monitoring system experiences massive data volumes during peak hours, leading to excessive pressure on the monitoring system. Furthermore, existing solutions involve adjusting the monitoring cluster after the business cluster is expanded, resulting in monitoring adaptation delays and risks.
By receiving monitoring data within a preset time period, using a time-series prediction model to predict monitoring data for future time periods, the scheduling configuration information of the monitoring cluster is dynamically adjusted, including the number of monitoring clusters, the indicators to be monitored, and the data collection frequency, and capacity is expanded in advance to cope with business peaks.
By predicting peak and off-peak business times in the future, we can rationally plan monitoring indicators and collection frequencies, reduce the pressure on the monitoring system, avoid monitoring gaps, and lower business risks.
Smart Images

Figure CN115168042B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a management method for a monitoring cluster, a management device for a monitoring cluster, a computer storage medium, and an electronic device. Background Technology
[0002] Cloud-native technologies, such as containers and microservices, have given rise to a new generation of cloud computing systems. However, these new technologies have also brought new challenges. Containerization and microservices have significantly increased the complexity of monitoring cloud-native application clusters.
[0003] In related technologies, monitoring data corresponding to all monitoring metrics is typically collected periodically, and then the operational status of the business cluster is evaluated based on the monitoring data. However, this approach may result in a massive amount of data in the monitoring cluster during peak business hours, thus placing excessive pressure on the monitoring system.
[0004] Therefore, there is an urgent need in this field to develop a new management method and device for monitoring clusters.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure. Summary of the Invention
[0006] The purpose of this disclosure is to provide a management method, management device, computer storage medium, and electronic device for a monitoring cluster, thereby overcoming, at least to some extent, the technical problem of excessive pressure on the monitoring system caused by limitations in related technologies.
[0007] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part by practice of this disclosure.
[0008] According to a first aspect of this disclosure, a management method for a monitoring cluster is provided, the monitoring cluster being used to monitor the operational status of a business cluster. The method includes: the monitoring cluster being used to monitor the operational status of the business cluster, the method including: receiving monitoring data for the business cluster within a preset time period; the monitoring data including monitoring indicator values corresponding to multiple monitoring indicator items; predicting predicted monitoring data values for the business cluster in a future time period based on the monitoring data; and dynamically adjusting the scheduling configuration information of the monitoring cluster in the future time period based on the predicted monitoring data values; wherein the scheduling configuration information includes any one or more of the following: the number of monitoring clusters, the indicators to be monitored, and the data collection frequency.
[0009] In an exemplary embodiment of this disclosure, after receiving monitoring data of the service cluster within a preset time period, the method further includes: cleaning the monitoring data; and performing data aggregation processing on the cleaned monitoring data.
[0010] In an exemplary embodiment of this disclosure, the step of predicting the monitoring data prediction value of the business cluster in a future time period based on the monitoring data includes: inputting the monitoring data into a time-series prediction model, and obtaining the monitoring data prediction value of the business cluster in the future time period based on the output of the time-series prediction model; wherein, the time-series prediction model is used to predict the monitoring data prediction value of the business cluster in a future time period based on the monitoring data.
[0011] In an exemplary embodiment of this disclosure, the step of dynamically adjusting the scheduling configuration information of the monitoring cluster within the future time period based on the predicted value of the monitoring data includes: identifying peak and off-peak business times within the future time period based on the predicted value of the monitoring data; and dynamically adjusting the scheduling configuration information of the monitoring cluster during the peak and off-peak business times.
[0012] In an exemplary embodiment of this disclosure, the future time period includes multiple future moments; the predicted monitoring data value includes predicted monitoring indicator values corresponding to the multiple monitoring indicator items; identifying peak and off-peak business moments within the future time period based on the predicted monitoring data value includes: comparing the predicted monitoring indicator value for each future moment with the indicator threshold for each monitoring indicator item; when the predicted monitoring indicator value is greater than the preset indicator threshold, determining the monitoring indicator item to which the predicted monitoring indicator value belongs as a target monitoring indicator item; in response to the number of target monitoring indicator items being greater than the indicator item threshold, determining the future moment as the peak business moment; in response to the number of target monitoring indicator items not being greater than the indicator item threshold, determining the future moment as the off-peak business moment.
[0013] In an exemplary embodiment of this disclosure, after determining that the future time is the peak business time, the dynamic adjustment of the scheduling configuration information of the monitoring cluster during the peak business time includes: pre-expanding the monitoring cluster within a preset time period before the peak business time to increase the number of monitoring clusters; adjusting the number of monitoring indicators to be monitored by the expanded monitoring cluster during the peak business time so that the number of monitoring indicators is less than the number of multiple monitoring indicators; and adjusting the data acquisition frequency of the expanded monitoring cluster during the peak business time so that the data acquisition frequency is higher than the acquisition frequency of the monitoring data.
[0014] In an exemplary embodiment of this disclosure, after determining that the future time is the off-peak time of the business, the step of dynamically adjusting the scheduling configuration information of the monitoring cluster during the off-peak time of the business includes: adjusting the data acquisition frequency of the monitoring cluster during the off-peak time of the business so that the data acquisition frequency is lower than the acquisition frequency of the monitoring data.
[0015] According to a second aspect of this disclosure, a management device for a monitoring cluster is provided. The monitoring cluster is used to monitor the operating status of a business cluster. The device includes: a data receiving module for receiving monitoring data of the business cluster within a preset time period; the monitoring data includes monitoring indicator values corresponding to multiple monitoring indicator items; a data prediction module for predicting the monitoring data prediction value of the business cluster in a future time period based on the monitoring data; and a scheduling configuration module for dynamically adjusting the scheduling configuration information of the monitoring cluster in the future time period based on the predicted monitoring data value; wherein the scheduling configuration information includes any one or more of the following: the number of monitoring clusters, the indicator items to be monitored, and the data collection frequency.
[0016] According to a third aspect of this disclosure, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the management method of the monitoring cluster described in the first aspect.
[0017] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the management method of the monitoring cluster described in the first aspect by executing the executable instructions.
[0018] As can be seen from the above technical solutions, the monitoring cluster management method, monitoring cluster management device, computer storage medium, and electronic device in the exemplary embodiments of this disclosure have at least the following advantages and positive effects:
[0019] In some embodiments of this disclosure, the technical solutions provide, on the one hand, receiving monitoring data of the business cluster within a preset time period, and predicting the monitoring data prediction value of the business cluster in a future time period based on the monitoring data, thereby enabling advance understanding of the operation status of the business cluster in the future time period. This solves the problem in related technologies where the scheduling configuration information of the monitoring cluster is only changed after changes occur in the business cluster, resulting in some business clusters not being monitored, and reduces the operational risk of the business cluster. On the other hand, based on the predicted monitoring data value, dynamically adjusting the scheduling configuration information of the monitoring cluster in the future time period (the scheduling configuration information includes one or more of the following: the number of monitoring clusters, the indicators to be monitored, and the data collection frequency), it is possible to rationally plan the scheduling configuration information of the monitoring cluster according to the predicted monitoring data value. This avoids the problem in related technologies where the monitoring cluster is scheduled with fixed scheduling configuration information under all circumstances, resulting in low intelligence and excessive pressure on the monitoring system during peak business periods, thus reducing the pressure on the monitoring system.
[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0022] Figure 1 A flowchart illustrating the management method of the monitoring cluster in an embodiment of this disclosure is shown.
[0023] Figure 2 This illustration shows a flowchart of how the scheduling configuration information of the monitoring cluster is dynamically adjusted in a future time period based on the predicted value of the monitoring data in an embodiment of this disclosure.
[0024] Figure 3 This illustration shows a flowchart of identifying peak and off-peak business times in the future based on predicted values from monitoring data in an embodiment of this disclosure.
[0025] Figure 4 This illustration shows a flowchart of how to dynamically adjust the scheduling configuration information of the monitoring cluster during peak business hours in an embodiment of this disclosure;
[0026] Figure 5 This diagram illustrates the overall flow of the management method for the monitoring cluster in an embodiment of this disclosure.
[0027] Figure 6 This diagram illustrates the overall block diagram of the management method for the monitoring cluster in an embodiment of this disclosure;
[0028] Figure 7 This diagram illustrates the structure of a management device for a monitoring cluster in an exemplary embodiment of this disclosure.
[0029] Figure 8 A schematic diagram of the structure of an electronic device in an exemplary embodiment of this disclosure is shown. Detailed Implementation
[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0031] The terms “a,” “an,” “the,” and “the” are used in this specification to indicate the presence of one or more elements / components / etc.; the terms “including” and “having” are used to indicate an open-ended inclusion and to mean that there may be other elements / components / etc. in addition to the listed elements / components / etc.; the terms “first” and “second” are used only as markings and are not a limitation on the number of objects.
[0032] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0033] In cloud-native scenarios, monitoring systems typically design a large number of monitoring metrics to meet observability requirements. Existing solutions generally require periodically collecting all monitoring metrics. However, during peak business periods, this full-scale collection leads to massive metric values, high disk usage, and excessive pressure on the monitoring system.
[0034] Meanwhile, during peak business periods, as the business cluster expands, the monitoring cluster also needs to be expanded to cope with the rapid increase in monitoring metrics. Current solutions typically calculate the required number of monitoring clusters after the business cluster has expanded, and only expand further when the existing number is insufficient. However, this approach suffers from a lag in monitoring cluster adaptation when business clusters increase rapidly. Furthermore, the time required for expansion can lead to monitoring gaps, posing a risk to services with high requirements for network element monitoring.
[0035] In the embodiments of this disclosure, a management method for a monitoring cluster is first provided, which at least to some extent overcomes the shortcomings of excessive pressure on monitoring systems in related technologies.
[0036] Figure 1 The flowchart of the management method of the monitoring cluster in the present disclosure is shown. The execution subject of the management method of the monitoring cluster can be a server that manages the monitoring cluster (hereinafter referred to as the cluster management terminal).
[0037] refer to Figure 1 A method for managing a monitoring cluster according to an embodiment of this disclosure includes the following steps:
[0038] Step S110: Receive monitoring data for the business cluster within a preset time period; the monitoring data includes monitoring indicator values corresponding to multiple monitoring indicator items.
[0039] Step S120: Based on the monitoring data, predict the monitoring data forecast values of the business cluster in the future time period;
[0040] Step S130: Based on the predicted values from the monitoring data, dynamically adjust the scheduling configuration information of the monitoring cluster for future time periods.
[0041] exist Figure 1The technical solution provided in the illustrated embodiment, on the one hand, receives monitoring data of the business cluster within a preset time period, and predicts the monitoring data prediction value of the business cluster in the future time period based on the monitoring data. This allows for advance understanding of the operation status of the business cluster in the future time period, solving the problem in related technologies where the scheduling configuration information of the monitoring cluster is changed only after changes occur in the business cluster, resulting in some business clusters not being monitored, and reducing the operational risk of the business cluster. On the other hand, based on the predicted monitoring data value, the scheduling configuration information of the monitoring cluster in the future time period (the scheduling configuration information includes one or more of the following: the number of monitoring clusters, the indicators to be monitored, and the data collection frequency) is dynamically adjusted. This allows for reasonable planning of the scheduling configuration information of the monitoring cluster based on the predicted monitoring data value, avoiding the problem in related technologies where the monitoring cluster is scheduled with fixed scheduling configuration information under all circumstances, resulting in low intelligence and excessive pressure on the monitoring system during peak business periods, thus reducing the pressure on the monitoring system.
[0042] The following are Figure 1 The specific implementation process of each step in the process will be explained in detail:
[0043] A cluster refers to a group (or several) of independent computers forming a large computer service system using a high-speed communication network. Each cluster node (i.e., each computer in the cluster) is an independent server running its own services. These servers can communicate with each other, collaboratively providing users with applications, system resources, and data, and are managed in a single-system manner. When a user requests a cluster system, the cluster appears to the user as a single, independent server, but in reality, the user is requesting a group of cluster servers.
[0044] The service cluster in this disclosure is used to receive network requests sent by user terminals and to provide corresponding network services to user terminals.
[0045] The monitoring cluster in this disclosure is used to monitor the operating status of the aforementioned business clusters, so as to replace them in a timely manner when the business clusters are abnormal, and ensure the normal operation of related services on the user end.
[0046] Before step S110, the service discovery mechanism can discover services in the business clusters and send the pod information of each business cluster to the cluster management terminal. Then, the cluster management terminal can schedule the monitoring cluster according to the above pod information so that the monitoring cluster collects the monitoring index values of the business cluster within the preset time period at a preset data collection frequency (e.g., once every 10 seconds, which can be set according to the actual situation, and this disclosure does not make any special limitation on this), thereby realizing the monitoring of the business clusters.
[0047] For example, the aforementioned preset time period can be 24 hours, which can be set according to the actual situation. This disclosure does not impose any special limitations on this.
[0048] The aforementioned monitoring data may be monitoring indicator values corresponding to multiple monitoring indicator items. For example, the aforementioned multiple monitoring indicator items may include the CPU (central processing unit) usage, memory usage, disk usage, number of concurrent connections, number of loads, business traffic, etc. of each business cluster. These can be set according to actual conditions, and this disclosure does not impose any special limitations on them.
[0049] After the monitoring cluster collects the monitoring data of the business cluster within the aforementioned preset time period, it can send the monitoring data to the aforementioned cluster management terminal.
[0050] Then, in step S110, monitoring data of the business cluster within a preset time period can be received.
[0051] In this step, the cluster management terminal can receive the aforementioned monitoring data. After receiving the monitoring data, the cluster management terminal can preprocess the monitoring data to remove invalid data and improve data availability.
[0052] For example, the cluster management terminal can first perform data cleaning on the aforementioned monitoring data. Data cleaning is used to discover and correct identifiable errors in the data files, including checking data consistency and handling invalid and missing values. There are four key points to data cleaning: completeness (i.e., detecting whether a single data record contains null values and whether the statistical fields are complete), comprehensiveness (viewing all values in a column, and judging whether the data is comprehensive through maximum, minimum, average, data definitions, etc.), legality (whether the type, content, and size of the values meet the set expectations), and uniqueness (whether the data is recorded repeatedly). The standard for cleaning is to make the data clean and continuous.
[0053] After cleaning the monitoring data, the cluster management terminal can also perform data aggregation processing on some fine-grained monitoring data. For example, taking CPU utilization as a monitoring indicator, assume that the CPU contains multiple shards (e.g., shard 1 and shard 2), and the CPU utilization of the business cluster at time m on shard 1 is 20%, and the CPU utilization at time m on shard 2 is 30%. Therefore, the above two CPU utilizations (20% and 30%) can be aggregated (e.g., the average of the two) to obtain the CPU utilization of the business cluster at time m.
[0054] By aggregating the cleaned monitoring data, the fine-grained fragmented data collected by the monitoring cluster can be aggregated into coarse-grained data that can characterize the overall operating status of the business cluster.
[0055] After preprocessing the above monitoring data, we can proceed to step S120.
[0056] In step S120, the predicted value of the monitoring data of the business cluster in the future time period is obtained based on the monitoring data.
[0057] In this step, the preprocessed monitoring data can be input into a time-series prediction model (e.g., the NeuralProphet model) so that the time-series prediction model can predict the monitoring data forecast values of the business cluster in future periods based on the monitoring data.
[0058] NeuralProphet is a decomposable time series model used for modeling time series data based on neural networks. This library uses PyTorch as its backend. Its components include trend, seasonality, autoregression, special events, future regression terms, and lagged regression terms. This model combines the scalability of neural networks with the interpretability of AR models (autoregressive networks, which are single-layer networks trained to simulate the AR process in time series signals, but on a much larger scale than traditional models), improving its accuracy and scalability.
[0059] Specifically, after inputting the aforementioned monitoring data into the NeuralProphet model, the NeuralProphet model can predict the monitoring data forecast values for the aforementioned business cluster in future time periods based on the following formula 1:
[0060] y(t)=T(t)+S(t)+E(t)+F(t)+A(t)+L(t) Formula 1
[0061] Where y(t) represents the predicted value of the monitoring data of the business cluster at time t;
[0062] T(t) represents the trend at time t, that is, the trend of the time series on a non-periodic surface. For example, the point of change can be used to model the trend.
[0063] S(t) represents the seasonality of time t. For example, Fourier terms can be used for modeling to handle various seasonalities in the data.
[0064] E(t) represents the event and holiday effect at time t. This term is a covariate of the model. For example, it can be modeled using a univariate, single-weight approach.
[0065] F(t) represents the regression effect of a known exogenous variable at time t. This term is also a covariate of the model and can be modeled using a univariate, single-weight approach.
[0066] A(t) represents the autoregressive effect of past observation time t, which can be predicted using AR-Net;
[0067] L(t) represents the regression effect observed after the exogenous variable at time t, and this term can be predicted using a feedforward neural network.
[0068] For example, after inputting the monitoring data of the business cluster within a preset time period into the NeuralProphet model, the NeuralProphet model can predict T(t), S(t), E(t), F(t), A(t) and L(t) in sequence based on the monitoring data, and then predict the monitoring data prediction value y(t) at time t based on the above formula 1.
[0069] For example, taking 48 hours after the preset time period as the future time period, the predicted monitoring data value of the business cluster every 10 seconds within the 48 hours can be obtained.
[0070] After predicting the monitoring data of the business cluster in the future period, step S130 can be entered to dynamically adjust the scheduling configuration information of the monitoring cluster in the future period based on the predicted monitoring data value.
[0071] In this step, the scheduling configuration information mentioned above may include the number of monitoring clusters (i.e., the number of monitoring clusters used to monitor business clusters), the indicators to be monitored (i.e., the monitoring indicators that the monitoring clusters need to collect in the future time period), and the data collection frequency (i.e., how many data samples are collected per unit time).
[0072] refer to Figure 2 , Figure 2 This embodiment of the present disclosure illustrates a flowchart of dynamically adjusting the scheduling configuration information of the monitoring cluster for future time periods based on predicted values from monitoring data, including steps S201-S202:
[0073] In step S201, based on the predicted values from the monitoring data, the peak and off-peak business times for future periods are identified.
[0074] In this step, the aforementioned future time period can include multiple future moments. Taking a data collection frequency of 10 seconds as an example, the interval between two adjacent future moments is 10 seconds. Furthermore, the predicted monitoring data value for each future moment includes multiple predicted monitoring indicator values corresponding to multiple monitoring indicator items.
[0075] refer to Figure 3, Figure 3 This embodiment of the present disclosure illustrates a flowchart for identifying peak and off-peak business hours in the future based on predicted values from monitoring data, including steps S301-S304:
[0076] In step S301, the predicted values of multiple monitoring indicators at each future time are compared with the threshold values of each monitoring indicator item.
[0077] In this step, the predicted value of each monitoring indicator at each future time can be compared with the threshold value of its respective monitoring indicator item to determine the numerical relationship between the predicted value and the threshold value.
[0078] The following explanation uses several monitoring metrics, including CPU utilization, concurrent connections, and service traffic, as examples:
[0079] Assuming the predicted value of the CPU utilization monitoring indicator at time n is 40%, and the threshold value of the CPU utilization indicator is 50%, it can be determined that the predicted value of the CPU utilization monitoring indicator of the business cluster at time n is less than the threshold value.
[0080] Assuming the predicted value of the concurrent connection count monitoring metric at time n is 9, and the threshold value of the concurrent connection count metric is 5, it can be determined that the concurrent connection count of the business cluster at time n is greater than the metric threshold.
[0081] Assuming the predicted value of the monitoring data for business traffic at time n is 1024Mb, and the threshold value for business traffic metrics is 1000Mb, it can be determined that the business traffic of the business cluster at time n is greater than the threshold value.
[0082] After obtaining the above numerical comparison results, we can proceed to step S302. When the predicted value of the monitoring indicator is greater than the indicator threshold, the monitoring indicator item to which the predicted value of the monitoring indicator belongs is determined as the target monitoring indicator item.
[0083] In this step, referring to the relevant explanation of step S301 above, the two monitoring indicators of concurrent connections and service traffic at time n can be determined as the target monitoring indicators, that is, the number of target monitoring indicators is 2.
[0084] In step S303, in response to the number of target monitoring indicators being greater than the indicator threshold, a future time is determined to be a peak business time.
[0085] In this step, taking the threshold of the above indicator item as 1 as an example, it can be determined that the number of the above target monitoring indicator items is greater than the indicator item threshold. Therefore, the future time n can be determined as the peak time of business.
[0086] In step S304, in response to the fact that the number of predicted values of the target monitoring data is not greater than the threshold of the indicator item, the future time is determined to be the off-peak time of business.
[0087] In this step, we will still use the above-mentioned monitoring metrics, including CPU utilization, concurrent requests, and business traffic, as examples for explanation:
[0088] Assuming the predicted value of the CPU utilization monitoring indicator at time n is 40%, and the threshold value of the CPU utilization indicator is 50%, it can be determined that the predicted value of the CPU utilization monitoring indicator of the business cluster at time n is less than the threshold value.
[0089] Assuming the predicted value of the concurrent connection count monitoring metric at time n is 3, and the threshold value of the concurrent connection count metric is 5, it can be determined that the concurrent connection count of the business cluster at time n is less than the threshold value.
[0090] Assuming the predicted value of the monitoring metric for service traffic at time n is 800Mb, and the threshold value for the service traffic metric is 1000Mb, then it can be determined that the service traffic of the service cluster at time n is less than the threshold value.
[0091] Therefore, it can be determined that the predicted values of the monitoring data at time n are all no greater than the threshold of the monitoring indicator item to which they belong. That is, there are 0 predicted values of target monitoring data among the predicted values of the monitoring indicators at time n, and 0 is less than the threshold of the above indicator item 1. Therefore, it can be determined that the future time n is a low-peak time for business.
[0092] After determining the peak and off-peak business hours for the future period, step S202 can be taken to dynamically adjust the scheduling configuration information of the monitoring cluster during peak and off-peak business hours.
[0093] In this step, the following will be combined first. Figure 4 This disclosure explains how to dynamically adjust the scheduling configuration information of the monitoring cluster during peak business hours:
[0094] refer to Figure 4 , Figure 4 This embodiment of the present disclosure illustrates a flowchart of how to dynamically adjust the scheduling configuration information of the monitoring cluster during peak business hours, including steps S401-S403:
[0095] In step S401, the monitoring cluster is expanded in advance within a preset time period before the peak business hours to increase the number of monitoring clusters.
[0096] In this step, the monitoring cluster can be expanded in advance within a preset time period before the arrival of the aforementioned peak business hours. For example, the preset time period can be determined according to the time required for the monitoring cluster expansion process. Taking the time required for the monitoring cluster expansion as 5 seconds as an example, the preset time period can be set to 5 seconds or 6 seconds, which can be set according to the actual situation. This disclosure does not impose any special limitations on this.
[0097] Therefore, this disclosure can solve the problem in related technologies that the monitoring cluster is only expanded when the business cluster is detected to be expanded after the peak business period, resulting in a monitoring blank area in some business clusters, and reduce the risk that the abnormal situation of the business cluster will not be detected due to untimely monitoring.
[0098] In step S402, the number of monitoring indicators to be monitored in the monitoring cluster after the expansion process is adjusted during peak business hours so that the number of monitoring indicators is less than the number of multiple monitoring indicators.
[0099] In this step, during peak business hours, the number of monitoring metrics that the monitoring cluster needs to collect can be reduced. For example, the monitoring metrics that the monitoring cluster needs to collect during peak business hours can be adjusted to: business traffic. This way, the monitoring cluster does not need to collect other metric values that are unrelated to business traffic, reducing resource consumption and reducing the pressure on the monitoring system.
[0100] In step S403, the data acquisition frequency of the monitoring cluster after the expansion process is adjusted during peak business hours so that the data acquisition frequency is higher than the acquisition frequency of monitoring data.
[0101] In this step, after scaling up the monitoring cluster, the data collection frequency of the expanded cluster during peak business hours can be increased. For example, the data collection frequency can be adjusted to once every 5 seconds, making it higher than the collection frequency of the aforementioned monitoring indicators. This allows for closer monitoring of the business cluster's operational status during peak business hours, thereby improving the speed of handling related emergencies.
[0102] The following explains how to dynamically adjust the scheduling configuration information of the monitoring cluster during off-peak hours in this disclosure:
[0103] For example, after determining that a future time will be a low-peak time for business, on the one hand, the monitoring indicators that the monitoring cluster needs to collect can be adjusted to: other monitoring indicators besides the number of concurrent connections. This way, the monitoring cluster does not need to collect unnecessary indicators during the low-peak time for business, reducing the pressure on the monitoring system.
[0104] On the other hand, the data collection frequency of the monitoring cluster can be reduced to be lower than the collection frequency of the aforementioned monitoring indicators. This can prevent excessive monitoring of the monitoring cluster during off-peak hours, thereby reducing the resource consumption and system pressure of the monitoring cluster while ensuring the normal operation of the business cluster within a controllable range.
[0105] Based on the above technical solution, this disclosure has the following technical effects:
[0106] First, by predicting the peak and off-peak times of business clusters in the future, and then rationally planning the monitoring indicators and data collection frequency for different time periods based on the prediction results, it is possible to reduce disk resource consumption and reduce the pressure on the monitoring system.
[0107] Second, it can expand the monitoring cluster in advance based on the prediction results to cope with the surge in business scale, avoid monitoring gaps, and reduce system risks.
[0108] refer to Figure 5 , Figure 5 This diagram illustrates the overall flow of the management method for the monitoring cluster in this embodiment, including steps S501-S507:
[0109] In step S501, the monitoring system collects monitoring data related to the load of the business cluster within a certain time period (with a period of seconds or minutes) and transmits it to the prediction module.
[0110] In step S502, the prediction module performs preprocessing on the monitoring data, such as data cleaning and data aggregation, to improve the effectiveness of the data;
[0111] In step S503, the prediction module uses the NeuralProphet model to predict the peak and off-peak business times in the future period based on the preprocessed monitoring data, and then passes the prediction results to the monitoring task planning module.
[0112] In step S504, the monitoring task planning module dynamically adjusts the scheduling configuration information of the monitoring cluster in future time periods based on the prediction results. For example, during off-peak business periods, the collection frequency is reduced and indicators such as concurrency are not collected, while during peak business periods, the collection frequency is increased and only relevant indicators such as business traffic are collected.
[0113] In step S505, the monitoring task planning module plans the number of monitoring nodes in advance based on the prediction results and expands the monitoring cluster before the peak business hours.
[0114] In step S506, the monitoring task planning module collects information on monitoring nodes and business clusters that need to be monitored, allocates monitoring tasks to the expanded monitoring clusters, and the monitoring tasks include the objects to be monitored and the monitoring indicators to be collected, and sends the monitoring tasks to the monitoring clusters.
[0115] In step S507, the monitoring cluster receives and executes the monitoring task.
[0116] refer to Figure 6 , Figure 6 This diagram illustrates the overall block diagram of the management method for the monitoring cluster in an embodiment of this disclosure:
[0117] The service discovery module probes the business clusters through the service discovery mechanism, and transmits the pod information of each business cluster to the monitoring task planning module after aggregating it.
[0118] The monitoring system contains multiple monitoring nodes, which are used to receive scheduling from the monitoring task planning module, collect monitoring data from the business cluster, transmit the data to the operation and maintenance platform and the time series prediction module, and report their own monitoring node information to the monitoring task planning module.
[0119] The time-series prediction module preprocesses the received monitoring data, and then, based on the NeuralProphet model, predicts the peak and off-peak periods of the business cluster in the future time period, and transmits the prediction results to the monitoring task planning module.
[0120] The monitoring task planning module receives relevant information from the business cluster and the monitoring cluster, plans the monitoring tasks of the monitoring cluster in the future period based on the prediction results transmitted by the prediction module (including the monitoring indicators to be collected and their data collection frequency, etc.), assigns the monitoring tasks to the monitoring cluster, and expands the monitoring cluster in advance before the peak business period.
[0121] It should be noted that this disclosure can be applied to monitoring and maintenance scenarios in cloud-native network element clusters (a network element cluster refers to a cluster of network elements composed of multiple network elements that divide the entire hardware and software system resources of one or more servers into several network elements that process the same application service in order to enhance the server's response and processing capabilities to application services) and edge cloud and other resource-scarce scenarios. This can reduce invalid monitoring data, improve the resource utilization of the monitoring system, avoid monitoring gaps caused by the rapid increase in the scale of the business cluster and the untimely expansion of the monitoring system cluster, reduce business risks, and ensure the accuracy and completeness of cloud-native network element cluster monitoring data.
[0122] This disclosure also provides a management device for monitoring clusters. Figure 7 This diagram illustrates the structure of a management device for a monitoring cluster in an exemplary embodiment of this disclosure; as shown below. Figure 7 As shown, the management device 700 for the monitoring cluster may include a data receiving module 710, a data prediction module 720, and a scheduling configuration module 730. Wherein:
[0123] The data receiving module 710 is used to receive monitoring data of the business cluster within a preset time period; the monitoring data includes monitoring indicator values corresponding to multiple monitoring indicator items.
[0124] The data prediction module 720 is used to predict the monitoring data prediction value of the business cluster in a future time period based on the monitoring data.
[0125] The scheduling configuration module 730 is used to dynamically adjust the scheduling configuration information of the monitoring cluster in the future time period based on the predicted value of the monitoring data; wherein the scheduling configuration information includes any one or more of the following: the number of monitoring clusters, the indicators to be monitored, and the data collection frequency.
[0126] In an exemplary embodiment of this disclosure, after receiving monitoring data of the service cluster within a preset time period, the data receiving module 710 is configured to:
[0127] The monitoring data is cleaned; the cleaned monitoring data is then aggregated.
[0128] In an exemplary embodiment of this disclosure, the data prediction module 720 is configured to:
[0129] The monitoring data is input into the time-series prediction model, and the predicted monitoring data value of the business cluster in the future time period is obtained based on the output of the time-series prediction model; wherein, the time-series prediction model is used to predict the predicted monitoring data value of the business cluster in the future time period based on the monitoring data.
[0130] In an exemplary embodiment of this disclosure, the scheduling configuration module 730 is configured to:
[0131] Based on the predicted values of the monitoring data, identify the peak and off-peak business times within the future time period; and dynamically adjust the scheduling configuration information of the monitoring cluster during the peak and off-peak business times.
[0132] In an exemplary embodiment of this disclosure, the future time period includes multiple future times; the predicted monitoring data values include predicted monitoring indicator values corresponding to the multiple monitoring indicator items; the scheduling configuration module 730 is configured to:
[0133] The predicted value of each monitoring indicator at each future time is compared with the indicator threshold of each monitoring indicator item. When the predicted value of the monitoring indicator is greater than the preset indicator threshold, the monitoring indicator item to which the predicted value belongs is determined as the target monitoring indicator item. In response to the number of target monitoring indicator items being greater than the indicator item threshold, the future time is determined to be the peak time of the business. In response to the number of target monitoring indicator items not being greater than the indicator item threshold, the future time is determined to be the off-peak time of the business.
[0134] In an exemplary embodiment of this disclosure, after determining that the future time is the peak business time, the scheduling configuration module 730 is configured to:
[0135] Within a preset time period before the peak business hours, the monitoring cluster is pre-expanded to increase the number of monitoring clusters; the number of monitoring indicators to be monitored in the expanded monitoring cluster during the peak business hours is adjusted so that the number of monitoring indicators is less than the number of multiple monitoring indicators; the data collection frequency of the expanded monitoring cluster during the peak business hours is adjusted so that the data collection frequency is higher than the collection frequency of the monitoring data.
[0136] In an exemplary embodiment of this disclosure, after determining that the future time is the off-peak time for the service, the scheduling configuration module 730 is configured to:
[0137] Adjust the data collection frequency of the monitoring cluster during off-peak hours so that the data collection frequency is lower than the collection frequency of the monitoring data.
[0138] The specific details of each module in the management device of the aforementioned monitoring cluster have been described in detail in the corresponding management method of the monitoring cluster, so they will not be repeated here.
[0139] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0140] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0141] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0142] This application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device.
[0143] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0144] A computer-readable storage medium can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0145] A computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the above embodiments.
[0146] Furthermore, this disclosure also provides an electronic device capable of implementing the above-described method.
[0147] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0148] The following reference Figure 8 To describe an electronic device 800 according to such an embodiment of the present disclosure. Figure 8 The electronic device 800 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0149] like Figure 8 As shown, the electronic device 800 is manifested in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one processing unit 810, at least one storage unit 820, a bus 830 connecting different system components (including storage unit 820 and processing unit 810), and a display unit 840.
[0150] The storage unit stores program code that can be executed by the processing unit 810, causing the processing unit 810 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 810 can perform actions such as... Figure 1 As shown: Step S110, receive monitoring data for the service cluster within a preset time period; the monitoring data includes monitoring indicator values corresponding to multiple monitoring indicator items; Step S120, predict the monitoring data prediction value of the service cluster in a future time period based on the monitoring data; Step S130, dynamically adjust the scheduling configuration information of the monitoring cluster in the future time period based on the predicted monitoring data value.
[0151] Storage unit 820 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 8201 and / or cache memory 8202, and may further include a read-only memory (ROM) 8203.
[0152] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0153] Bus 830 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0154] Electronic device 800 can also communicate with one or more external devices 900 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 800, and / or with any device that enables electronic device 800 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 850. Furthermore, electronic device 800 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 860. As shown, network adapter 860 communicates with other modules of electronic device 800 via bus 830. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0155] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
Claims
1. A management method of monitoring a cluster, characterized by, The monitoring cluster is used for monitoring the running state of the service cluster, and the method comprises: receiving monitoring data of the service cluster within a preset period; the monitoring data comprises monitoring indicator values corresponding to a plurality of monitoring indicator items; predicting monitoring data prediction values of the service cluster within a future period according to the monitoring data; based on the monitoring data prediction values, dynamically adjusting scheduling configuration information of the monitoring cluster within the future period; wherein the scheduling configuration information comprises any one or more of the following: the number of the monitoring cluster, the to-be-monitored indicator items and the data collection frequency; the dynamic adjustment of the scheduling configuration information of the monitoring cluster within the future period based on the monitoring data prediction values comprises: based on the monitoring data prediction values, identifying service peak hours and service low peak hours within the future period; dynamically adjusting the scheduling configuration information of the monitoring cluster at the service peak hours and the service low peak hours.
2. The method of claim 1, wherein, After receiving the monitoring data of the service cluster within the preset period, the method further comprises: performing data cleaning on the monitoring data; performing data aggregation processing on the data cleaned monitoring data.
3. The method of claim 1, wherein, the prediction of the monitoring data prediction values of the service cluster within the future period according to the monitoring data comprises: inputting the monitoring data into a time series prediction model, and obtaining the monitoring data prediction values of the service cluster within the future period according to the output of the time series prediction model; wherein the time series prediction model is used for predicting the monitoring data prediction values of the service cluster within the future period according to the monitoring data.
4. The method of claim 1, wherein, The future period comprises a plurality of future time points; the monitoring data prediction values comprise monitoring indicator prediction values corresponding to the plurality of monitoring indicator items; the identification of the service peak hours and the service low peak hours within the future period based on the monitoring data prediction values comprises: numerical comparison of each monitoring indicator prediction value of each future time point with an indicator threshold value of each monitoring indicator item; when the monitoring indicator prediction value is greater than the preset indicator threshold value, determining the monitoring indicator item to which the monitoring indicator prediction value belongs as a target monitoring indicator item; in response to the number of target monitoring indicator items being greater than an indicator item threshold value, determining the future time point as the service peak hour; in response to the number of target monitoring indicator items being not greater than the indicator item threshold value, determining the future time point as the service low peak hour.
5. The method of claim 4, wherein, after determining that the future time point is the service peak hour, the dynamic adjustment of the scheduling configuration information of the monitoring cluster at the service peak hour comprises: pre-expanding the monitoring cluster within a preset period before the service peak hour to increase the number of the monitoring cluster; adjusting the to-be-monitored indicator items of the monitoring cluster after the expansion processing at the service peak hour, so that the number of the to-be-monitored indicator items is less than the number of the plurality of monitoring indicator items; adjusting the data collection frequency of the monitoring cluster after the expansion processing at the service peak hour, so that the data collection frequency is higher than the collection frequency of the monitoring data.
6. The method of claim 4, wherein, After determining that the future time is the service low peak time, the dynamic adjustment of the scheduling configuration information of the monitoring cluster at the service low peak time comprises: Adjusting the data collection frequency of the monitoring cluster at the service low peak time, so that the data collection frequency is lower than the monitoring data collection frequency.
7. A management apparatus that monitors a cluster, characterized by comprising: The monitoring cluster is used to monitor the running state of the service cluster, and the method comprises: A data receiving module is configured to receive monitoring data of the service cluster within a preset period; the monitoring data comprises monitoring index values corresponding to a plurality of monitoring index items; A data prediction module is configured to predict monitoring data prediction values of the service cluster within a future period according to the monitoring data; A scheduling configuration module is configured to dynamically adjust scheduling configuration information of the monitoring cluster within the future period based on the monitoring data prediction values; wherein the scheduling configuration information comprises any one or more of the following: the number of the monitoring cluster, the to-be-monitored index item, and the data collection frequency; The scheduling configuration module is configured to identify service peak times and service low peak times within the future period based on the monitoring data prediction values; and dynamically adjust the scheduling configuration information of the monitoring cluster at the service peak times and the service low peak times.
8. A computer storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the management method of the monitoring cluster in any one of claims 1-6.
9. An electronic device, comprising: Comprise: A processor; And A memory for storing executable instructions of the processor; Wherein the processor is configured to execute the executable instructions to perform the management method of the monitoring cluster in any one of claims 1-6.
Citation Information
Patent Citations
Information processing method, device and system
CN114218042A
Intelligent elastic scaling method and system based on CPU and memory occupancy rate
CN114428666A