Cluster node management method, computer program product, equipment and medium
By acquiring historical performance metrics of the cloud platform cluster and using a long short-term memory network model to predict future metrics, and adjusting performance thresholds, the accuracy and real-time issues of cluster node number management in existing technologies are resolved, achieving more reasonable and accurate node management.
Patent Information
- Application Number
- CN202511375141.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, the management of the number of container cluster nodes on cloud platforms relies on manual user intervention and preset thresholds, which makes it difficult to respond to rapidly changing load demands in real time, and the preset thresholds are difficult to accurately reflect the current load situation.
By obtaining historical performance metrics of the target cluster, a basic performance threshold is determined, and a long short-term memory network model is used to predict future performance metrics. The basic performance threshold is then adjusted to obtain the target performance threshold. Finally, the number of cluster nodes is managed by combining the target performance threshold and future performance metrics.
It achieves accuracy and rationality in cluster node number management, breaks through the limitations of static thresholds, and can dynamically adjust according to historical patterns and future trends, thus improving the timeliness and adaptability of management.
Smart Images

Figure CN120881075A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing technology, and more specifically, to a cluster node management method, computer program product, device, and medium. Background Technology
[0002] The development of cloud computing technology has made cloud platforms the infrastructure for storing and processing data. However, with the continuous growth of data volume and computing demands, cloud platform container cluster nodes face challenges in managing their number.
[0003] Currently, the management of the number of container cluster nodes on cloud platforms relies on manual user intervention and preset thresholds, which has significant limitations. First, manual intervention requires substantial manpower and time, making it difficult to respond to rapidly changing load demands in real time. Second, preset thresholds often fail to accurately reflect the current load situation.
[0004] In conclusion, how to accurately manage the number of cluster nodes is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a cluster node management method, which can, to a certain extent, solve the technical problem of how to accurately manage the number of cluster nodes. This invention also provides a computer program product, an electronic device, and a computer-readable storage medium.
[0006] To achieve the above objectives, firstly, a cluster node management method is provided, including: Obtain historical performance metrics for the target cluster; Based on the historical performance metrics, determine the basic performance threshold for managing the number of cluster nodes; Obtain future performance metrics after predicting the performance metrics of the target cluster; Based on the future performance indicators, the basic performance threshold is adjusted to obtain the target performance threshold; The number of cluster nodes is managed according to the target performance threshold and the future performance indicators.
[0007] On the other hand, based on the aforementioned historical performance metrics, the basic performance thresholds for managing the number of cluster nodes are determined, including: The historical performance indicators were analyzed to obtain the historical performance load peak values; The historical performance load peaks are adjusted to obtain the initial performance threshold for managing the number of cluster nodes; The basic performance threshold is determined based on the initial performance threshold.
[0008] On the other hand, the historical performance load peaks are adjusted to obtain the initial performance threshold for managing the number of cluster nodes, including: Determine the target service of the target cluster; Generate a security factor to stabilize the target service; Based on the security factor, the historical performance load peak is adjusted to obtain the initial performance threshold for managing the number of cluster nodes.
[0009] On the other hand, generating a security factor for stabilizing the target service includes: Determine the business type of the target business; If the business type is a core business, a security factor with a value greater than 1 is generated. If the business type is a non-core business, a security factor with a value less than or equal to 1 is generated.
[0010] On the other hand, determining the basic performance threshold based on the initial performance threshold includes: Determine the target service of the target cluster; If the target service is a latency-sensitive service, the initial performance threshold is lowered to obtain the basic performance threshold. If the target service is a non-latency-sensitive service, then the initial performance threshold is used as the basic performance threshold.
[0011] On the other hand, based on the future performance indicators, the basic performance threshold is adjusted to obtain the target performance threshold, including: The future performance indicators are analyzed to determine the load change trend of the performance indicators in the first time period in the future; Based on the load change trend, the basic performance threshold is adjusted to obtain the target performance threshold.
[0012] On the other hand, the basic performance threshold is adjusted according to the load change trend to obtain the target performance threshold, including: Determine the target service of the target cluster; Generate a floating ratio that is positively correlated with the level of the target service; Based on the load change trend, the basic performance threshold is adjusted according to the floating ratio to obtain the target performance threshold.
[0013] On the other hand, based on the load change trend and the floating ratio, the basic performance threshold is adjusted to obtain the target performance threshold, including: In response to the load change trend being upward, the basic performance threshold is adjusted upward based on the floating ratio to obtain the target performance threshold. In response to the load change trend being downward, the base performance threshold is adjusted downward based on the floating ratio to obtain the target performance threshold.
[0014] On the other hand, analyzing the historical performance indicators yields historical performance load peaks, including: The historical performance indicators were analyzed to obtain the historical peak performance load. Adjusting the historical performance load peaks yields the basic performance thresholds for managing the number of cluster nodes, including: The historical peak performance load is adjusted to obtain the basic performance threshold for scaling up cluster nodes.
[0015] On the other hand, the number of cluster nodes is managed according to the target performance threshold and the future performance indicators, including: Obtain the target performance metrics of the target cluster within a preset time period; The target performance indicators are analyzed to obtain multiple performance values of the target cluster to be evaluated within the preset time period; The future performance indicators are analyzed to determine the growth rate of the performance indicators in the second future time period. If the performance value to be evaluated is continuously greater than the target performance threshold and the growth rate is greater than a first set value, then the cluster node is expanded.
[0016] On the other hand, analyzing the historical performance indicators yields historical performance load peaks, including: Analyzing the historical performance indicators yields the historical low peak performance load. Adjusting the historical performance load peaks yields the basic performance thresholds for managing the number of cluster nodes, including: The historical low peak performance load is adjusted to obtain the basic performance threshold for scaling down cluster nodes.
[0017] On the other hand, the number of cluster nodes is managed according to the target performance threshold and the future performance indicators, including: Obtain the target performance metrics of the target cluster within a preset time period; The target performance indicators are analyzed to obtain multiple performance values of the target cluster to be evaluated within the preset time period; The future performance indicators are analyzed to determine the extent of their decrease over the next second time period. If the performance value to be evaluated is consistently lower than the target performance threshold and the decrease is greater than a second set value, then the cluster node is scaled down.
[0018] On the other hand, by analyzing the target performance indicators, multiple performance values of the target cluster to be evaluated are obtained during the preset time period, including: Set the time window value; The target performance indicators are collected according to the time window value to obtain the collected performance indicators; The collected performance metrics are processed to obtain the performance value of the target cluster to be evaluated within the time window. Return to the step of collecting the target performance index according to the time window value, until the collection of the target performance index is completed, and obtain multiple performance values to be evaluated.
[0019] On the other hand, obtaining future performance metrics after predicting the performance metrics of the target cluster includes: A pre-trained long short-term memory network model is determined, which is trained based on performance indicators for different time periods; The future performance metrics are obtained by predicting the performance metrics of the target cluster using the Long Short-Term Memory network model.
[0020] On the other hand, after managing the number of cluster nodes according to the target performance threshold and the future performance indicators, the method further includes: Identify the services to be migrated in the target cluster; Determine the target cluster node to receive the services to be migrated; Migrate a preset percentage of the service traffic in the service to be migrated to the target cluster node; According to the migration rule of increasing the target percentage traffic weight every third time interval, the remaining traffic of the service to be migrated is migrated to the target cluster node until the node load of the target cluster is balanced.
[0021] In a second aspect, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of any of the cluster node management methods described above.
[0022] Thirdly, an electronic device is provided, comprising: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of any of the cluster node management methods described above.
[0023] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the steps of any of the cluster node management methods described above.
[0024] This invention provides a cluster node management method that involves: acquiring historical performance indicators of a target cluster; determining a basic performance threshold for managing the number of cluster nodes based on the historical performance indicators; acquiring future performance indicators obtained by predicting the performance indicators of the target cluster; adjusting the basic performance threshold based on the future performance indicators to obtain a target performance threshold; and managing the number of cluster nodes according to the target performance threshold and the future performance indicators. The beneficial effects of this invention are: determining the basic performance threshold based on the historical performance indicators of the target cluster, and adjusting the basic performance threshold based on the predicted future performance indicators to obtain the target performance threshold, achieves the determination of the target performance threshold based on both historical and future performance indicators. This ensures that the target performance indicators conform to historical patterns while also considering future trends, overcoming the limitations of static thresholds and improving the accuracy and rationality of static thresholds. Furthermore, managing the number of cluster nodes by considering both the target performance threshold and future performance indicators makes changes in the number of cluster nodes more reasonable and accurate. This invention also provides a computer program product, electronic device, and computer-readable storage medium that solves the corresponding technical problems. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0026] Figure 1 A flowchart illustrating a cluster node management method provided in an embodiment of the present invention; Figure 2 A flowchart for adjusting historical performance load peaks; Figure 3 A flowchart for adjusting the basic performance thresholds; Figure 4 A flowchart for business migration; Figure 5 This is a schematic diagram of the cluster node management system. Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention; Figure 7 This is another structural schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] Please see Figure 1 , Figure 1 This is a flowchart of a cluster node management method provided in an embodiment of the present invention.
[0029] An embodiment of the present invention provides a cluster node management method, which may include the following steps: Step S101: Obtain historical performance metrics of the target cluster.
[0030] Step S102: Determine the basic performance threshold for managing the number of cluster nodes based on historical performance metrics.
[0031] In practical applications, cluster node changes are categorized into node expansion and node reduction. Node expansion occurs when the cluster load is too high, while node reduction occurs when the cluster load is too low. Since historical performance metrics reflect the cluster's load, these metrics can be obtained, and a basic performance threshold for managing the number of cluster nodes can be determined based on these metrics. Historical performance metrics can be those from the past 7 days, 10 days, or one month; this invention does not limit the time span of these historical performance metrics.
[0032] It should be noted that the type of cluster performance metrics in this invention can be flexibly determined according to the application scenario. For example, cluster performance metrics can include node-level metrics, container-level metrics, and business-level metrics. Node-level metrics can include hardware resource data such as CPU, memory, and disk I / O collected through Node Exporter; container-level metrics can include resource usage, lifecycle status, and memory working set size monitored by cadvisor; and business-level metrics can include QPS (Queries Per Second), response time (P50 / P90 / P99), and error rate obtained through tracking.
[0033] Step S103: Obtain the future performance metrics obtained after predicting the performance metrics of the target cluster.
[0034] Step S104: Adjust the basic performance threshold based on future performance indicators to obtain the target performance threshold.
[0035] In practical applications, the baseline performance threshold is determined based on historical performance metrics. However, since historical performance metrics are not very timely, the baseline performance threshold is also not very timely and may not match the actual operation of the cluster. To avoid management errors caused by directly applying the baseline performance threshold, we can obtain future performance metrics by predicting the performance metrics of the target cluster. Based on these future performance metrics, the baseline performance threshold is adjusted to obtain the target performance threshold. In this way, since future performance metrics are delayed, adjusting the baseline performance threshold based on future performance metrics is equivalent to combining the untimely historical performance metrics and the delayed future performance metrics to obtain the target performance threshold, thus ensuring the accuracy of the target performance threshold.
[0036] In an exemplary embodiment, the process of obtaining future performance indicators after predicting the performance indicators of the target cluster can be carried out by using a neural network model to predict the performance indicators. For example, a pre-trained Long Short-Term Memory (LSTM) network model can be determined. The LSTM network model is trained based on performance indicators of different time periods. The future performance indicators obtained after the LSTM network model predicts the performance indicators of the target cluster are then obtained.
[0037] Understandably, the application of LSTM models relies heavily on model training. During training, nearly 30 days of historical data (approximately 43,200 metrics) can be exported from Prometheus, an open-source system monitoring and alerting tool. Data cleaning is performed using Python scripts, such as filling missing values with linear interpolation and removing outliers using the 3σ principle. Time features such as hour, day of the week, and holiday status are extracted. Training and testing sets for the LSTM model are then constructed, for example, by dividing the training and testing sets in an 8:2 ratio. The LSTM model is then trained and tested to obtain a well-trained model. This LSTM model can be saved in TensorFlow Serving format, an open-source high-performance model deployment framework. It should be noted that the accuracy of the LSTM model may change over time; therefore, it is advisable to retrain the LSTM model periodically to maintain its accuracy.
[0038] Step S105: Manage the number of cluster nodes according to the target performance threshold and future performance indicators.
[0039] In practical applications, although the number of cluster nodes can be managed solely based on the target performance threshold, i.e., scaling down or expanding the cluster nodes, considering the possibility of accidental triggering of instantaneous peaks, and given that future performance indicators already reflect the development trend of cluster nodes, managing the number of cluster nodes in combination with the target performance threshold and future performance indicators is equivalent to managing cluster nodes in advance, which can enhance the accuracy and adaptability of cluster node management.
[0040] This invention provides a cluster node management method that involves: obtaining historical performance indicators of a target cluster; determining a basic performance threshold for managing the number of cluster nodes based on the historical performance indicators; obtaining future performance indicators obtained by predicting the performance indicators of the target cluster; adjusting the basic performance threshold based on the future performance indicators to obtain a target performance threshold; and managing the number of cluster nodes according to the target performance threshold and the future performance indicators. The beneficial effects of this invention are: determining the basic performance threshold based on the historical performance indicators of the target cluster, and adjusting the basic performance threshold based on the predicted future performance indicators to obtain the target performance threshold, achieves the determination of the target performance threshold based on both historical and future performance indicators. This ensures that the target performance indicators conform to historical patterns while also considering future trends, overcoming the limitations of static thresholds and improving the accuracy and rationality of static thresholds. Furthermore, managing the number of cluster nodes by considering both the target performance threshold and future performance indicators makes changes in the number of cluster nodes more reasonable and accurate.
[0041] Based on the above embodiments, considering that cluster node scaling up or down occurs when the load exceeds a threshold, the threshold can be determined according to the load. Please refer to [link to relevant documentation]. Figure 2 The cluster node management method provided in this embodiment of the invention may include the following steps: Step S201: Obtain historical performance metrics for the target cluster.
[0042] Step S202: Analyze historical performance indicators to obtain historical performance load peak values.
[0043] In other words, when determining the basic performance threshold for managing the number of cluster nodes based on historical performance metrics, the historical performance metrics can be analyzed to obtain historical performance load peaks, so that the performance thresholds can be determined subsequently based on these historical performance load peaks. The historical load peak refers to the maximum value of historical performance metrics (such as CPU utilization, memory utilization, and business QPS) over a past period, which can be a "complete business cycle," such as 7 days or 30 days.
[0044] Step S203: Adjust the historical performance load peak to obtain the initial performance threshold for managing the number of cluster nodes.
[0045] Step S204: Determine the basic performance threshold based on the initial performance threshold.
[0046] In practical applications, historical performance load peaks reflect the load range of cluster nodes. Therefore, by adjusting the historical performance load peaks, an initial performance threshold for managing the number of cluster nodes can be obtained, and then a basic performance threshold can be determined based on the initial performance threshold.
[0047] In the exemplary embodiment, considering that different services have different requirements for cluster load, such as high service stability requiring the cluster load to be not too large to avoid service fluctuations, the target service of the target cluster can be determined during the process of adjusting the historical performance load peak to obtain the initial performance threshold for managing the number of cluster nodes. A safety factor for stabilizing the target service can be generated. For example, the service type of the target service can be determined. If the service type is a core service, such as payment or order, a safety factor with a value greater than 1 is generated, for example, a safety factor selected between 1.05 and 1.1. If the service type is a non-core service, such as log collection, a safety factor with a value less than or equal to 1 is generated, for example, a safety factor selected between 0.95 and 1.0. The safety factor is a redundancy factor set to deal with "sudden load fluctuations", which can avoid "resource exhaustion at peak" caused by "historical peak is the threshold", thus reserving buffer space. The historical performance load peak is adjusted based on the safety factor, for example, by multiplying the safety factor by the historical performance load peak to obtain the initial performance threshold for managing the number of cluster nodes.
[0048] It should be noted that the safety factor in this invention can be adjusted over time. For example, the response time change rate and resource utilization improvement rate after cluster node changes can be calculated periodically every hour. When the feedback of the response time change rate and resource utilization improvement rate is good, the safety factor can be increased; when the feedback of the response time change rate and resource utilization improvement rate is poor, the safety factor can be decreased. The response time change rate can be obtained by (P99 after expansion - P99 before expansion) / P99 before expansion, and the resource utilization improvement rate can be obtained by (average utilization rate after expansion - average utilization rate before expansion) / average utilization rate before expansion.
[0049] In an exemplary embodiment, considering that different services have different response time requirements, and that node scaling up or down affects not only the number of nodes in the cluster but also the cluster resource configuration, which in turn affects the service response time, node scaling up or down will affect service response time. Therefore, to ensure that services meet response time requirements, the target service of the target cluster can be identified during the process of determining the basic performance threshold based on the initial performance threshold. If the target service is a latency-sensitive service (e.g., its tag indicates it is latency-sensitive), the initial performance threshold is lowered to obtain the basic performance threshold, for example, it can be set to 70%-80% of the initial performance threshold. If the target service is a non-latency-sensitive service, the initial performance threshold is used as the basic performance threshold. In this way, since the basic performance threshold for latency-sensitive services is obtained by lowering the initial performance threshold, the basic performance threshold for latency-sensitive services will be lower. This allows for node number management when the cluster load is not high, enabling advance management of the node number and reserving time for cluster resource changes, thus avoiding impacting the response time of latency-sensitive services.
[0050] Step S205: Obtain the future performance metrics obtained after predicting the performance metrics of the target cluster.
[0051] Step S206: Adjust the basic performance threshold based on future performance indicators to obtain the target performance threshold.
[0052] Step S207: Manage the number of cluster nodes according to the target performance threshold and future performance indicators.
[0053] As can be seen from the implementation process, in the process of determining the basic performance threshold for cluster node quantity management based on historical performance indicators, the present invention analyzes the historical performance indicators to obtain the historical performance load peak, and can adjust the historical performance load peak as needed to obtain the initial performance threshold for cluster node quantity management. Then, the basic performance threshold is determined based on the initial performance threshold. This achieves the goal of determining the basic performance threshold according to historical performance indicators, while avoiding directly using the historical performance load peak as the basic performance threshold. This improves the flexibility and adaptability of the basic performance threshold and enables more accurate cluster node quantity management.
[0054] Based on the above embodiments, please refer to Figure 3 The cluster node management method provided in this embodiment of the invention may include the following steps: Step S301: Obtain historical performance metrics for the target cluster.
[0055] Step S302: Analyze historical performance indicators to obtain historical performance load peak values.
[0056] Step S303: Adjust the historical performance load peak to obtain the initial performance threshold for managing the number of cluster nodes.
[0057] Step S304: Determine the basic performance threshold based on the initial performance threshold.
[0058] Step S305: Obtain the future performance metrics obtained after predicting the performance metrics of the target cluster.
[0059] Step S306: Analyze future performance indicators to determine the load change trend of performance indicators in the first time period in the future.
[0060] Step S307: Adjust the basic performance threshold according to the load change trend to obtain the target performance threshold.
[0061] In practical applications, when adjusting the basic performance threshold based on future performance metrics to obtain the target performance threshold, the future performance metrics reflect the future trend of performance metrics. This trend can be used to guide node scaling up or down. Therefore, the future performance metrics can be analyzed to determine the load change trend of the performance metrics within a first time period in the future. The first time period can be flexibly determined according to the application scenario, such as 30 minutes, 1 hour, etc. Then, the basic performance threshold is adjusted according to the load change trend to obtain the target performance threshold.
[0062] In an exemplary embodiment, changes in the number of nodes can affect business operations. Since different businesses have different operational requirements, the target business of the target cluster can be determined during the process of adjusting the basic performance threshold according to the load change trend to obtain the target performance threshold. A floating ratio positively correlated with the level of the target business is generated. This level can be an importance level, etc. For example, the more core the business, the higher the floating ratio. In order to avoid the floating ratio being too large, the range of the floating ratio can be limited, such as the floating ratio being 10%-15%. The basic performance threshold is adjusted based on the floating ratio according to the load change trend to obtain the target performance threshold.
[0063] In specific application scenarios, when adjusting the basic performance threshold according to the load change trend to obtain the target performance threshold, if the load change trend is upward, the basic performance threshold is increased based on the floating ratio to obtain the target performance threshold, so as to avoid untimely expansion affecting business operations; if the load change trend is downward, the basic performance threshold is decreased based on the floating ratio to obtain the target performance threshold, so as to avoid excessive expansion affecting business operations.
[0064] For ease of understanding, let's assume the load trend is "upward," such as a predicted "continuous increase" in load, like a 15% increase in QPS over the next hour. Then, the target performance threshold = base performance threshold × (1 + fluctuation percentage). Specifically, if the base performance threshold is 82.5%, and it increases by 10%, the target performance threshold is 82.5% × 1.1 = 90.75%. Conversely, let's assume the load trend is "downward," such as a 5% decrease in CPU utilization over the next hour. Then, the target performance threshold = base performance threshold × (1 - fluctuation percentage). Specifically, if the base threshold is 82.5%, and it decreases by 10%, the target performance threshold is 82.5% × 0.9 = 74.25%.
[0065] Step S308: Manage the number of cluster nodes according to the target performance threshold and future performance indicators.
[0066] In an exemplary embodiment, if it is necessary to expand the number of nodes, in the process of analyzing historical performance indicators to obtain the historical performance load peak value, the historical performance indicators can be analyzed to obtain the historical performance load high peak value; correspondingly, in the process of adjusting the historical performance load peak value to obtain the basic performance threshold for managing the number of cluster nodes, the historical performance load high peak value can be adjusted to obtain the basic performance threshold for expanding the number of cluster nodes. Subsequently, during the process of managing the number of cluster nodes according to the target performance threshold and future performance indicators, the target performance indicators of the target cluster within a preset time period can be obtained. The preset time period can be a continuous period determined based on the current moment. The target performance indicators are analyzed to obtain multiple performance values to be evaluated for the target cluster within the preset time period, avoiding the interference of instantaneous fluctuations in node number management caused by a single performance value to be evaluated. The future performance indicators are analyzed to determine the growth rate of the performance indicators in the second future time period. The growth rate is calculated as (predicted load in the second future time period - current load) / current load × 100%. The second time period can be flexibly determined according to the application scenario, such as 30 minutes, 1 hour, etc. In response to the performance value to be evaluated continuously exceeding the target performance threshold and the growth rate exceeding a first set value, which can be flexibly determined according to the application scenario, such as 20%, the cluster nodes are expanded.
[0067] For ease of understanding, assuming the target performance threshold is 94.875%, after monitoring the cluster, if the average CPU utilization over 5 minutes is 95.2% and the minimum is 94.9%, then it is determined that "the load continuously exceeds the target performance threshold". At the same time, assuming the current CPU utilization is 95.2%, and the predicted CPU utilization after 1 hour is 115%, then the increase is (115-95.2) / 95.2×100%≈20.8%, which exceeds 20%. Therefore, it can be determined that the performance value to be evaluated is continuously greater than the target performance threshold, and the increase is greater than the first set value. Then, the cluster can be expanded to include more nodes.
[0068] In an exemplary embodiment, if it is necessary to scale down the number of nodes, during the process of analyzing historical performance indicators to obtain the historical performance load peak, the historical performance indicators can be analyzed to obtain the historical performance load low peak. Correspondingly, during the process of adjusting the historical performance load peak to obtain the basic performance threshold for managing the number of cluster nodes, the historical performance load low peak can be adjusted to obtain the basic performance threshold for scaling down the number of cluster nodes. Subsequently, during the process of managing the number of cluster nodes according to the target performance threshold and future performance indicators, the target performance indicators of the target cluster within a preset time period can be obtained; the target performance indicators can be analyzed to obtain multiple performance values to be evaluated for the target cluster within the preset time period; the future performance indicators can be analyzed to determine the reduction rate of the performance indicators in a second future time period, where the reduction rate is (current load - predicted load in the second future time period) / current load × 100%; in response to the performance values to be evaluated being continuously less than the target performance threshold and the reduction rate being greater than a second set value, the cluster nodes are scaled down.
[0069] In specific application scenarios, during the analysis of target performance indicators to obtain multiple performance values to be evaluated for the target cluster within a preset time period, a time window value can be set, such as 5-10 minutes. The time window value can be determined based on the business response sensitivity; for example, 5 minutes for real-time transaction business and 10 minutes for batch processing business. Target performance indicators are collected according to the time window value to obtain the collected performance indicators. The collected performance indicators are processed to obtain the performance values to be evaluated for the target cluster within the time window value, such as using the average and / or minimum values of the collected performance indicators as the performance values to be evaluated for the target cluster within the time window value. The process of collecting target performance indicators according to the time window value is then repeated until the collection of target performance indicators is completed and multiple performance values to be evaluated are obtained.
[0070] As can be seen from the implementation process, in the process of adjusting the basic performance threshold based on future performance indicators to obtain the target performance threshold, the present invention determines the load change trend of the performance indicators in the first time period in the future, and then adjusts the basic performance threshold according to the load change trend to obtain the target performance threshold. In this way, the target performance threshold is consistent with the load change trend of the performance indicators, which can improve the matching degree between the target performance threshold and the future performance indicators, thereby enabling more accurate advance management of the number of cluster nodes.
[0071] Based on the above embodiments, changing the number of cluster nodes also involves resource reallocation. Please refer to [link / reference needed]. Figure 4 The cluster node management method provided in this embodiment of the invention may include the following steps: Step S401: Obtain historical performance metrics for the target cluster.
[0072] Step S402: Determine the basic performance threshold for managing the number of cluster nodes based on historical performance metrics.
[0073] Step S403: Obtain the future performance metrics obtained after predicting the performance metrics of the target cluster.
[0074] Step S404: Based on future performance indicators, adjust the basic performance threshold to obtain the target performance threshold.
[0075] Step S405: Manage the number of cluster nodes according to the target performance threshold and future performance indicators.
[0076] Step S406: Identify the services to be migrated in the target cluster.
[0077] Step S407: Determine the target cluster node to receive the services to be migrated.
[0078] In practical applications, after a cluster node change, it is necessary to migrate services from the original node to another node. Therefore, after managing the number of cluster nodes according to the target performance threshold and future performance indicators, it is necessary to determine the services to be migrated in the target cluster and the target cluster node to receive the services to be migrated, so as to migrate the services to be migrated according to the target cluster node.
[0079] Step S408: Migrate a preset percentage of the service traffic in the service to be migrated to the target cluster node.
[0080] Step S409: According to the migration rule of increasing the target percentage traffic weight every third time interval, migrate the remaining traffic of the service to be migrated to the target cluster node until the node load of the target cluster is balanced.
[0081] In practical applications, during the migration of services to be migrated, a predetermined percentage of the traffic from the services to be migrated can be migrated to the target cluster nodes first, for example, 10% of the traffic from the services to be migrated can be migrated to the target cluster nodes. Then, according to the migration rule of increasing the target percentage traffic weight every third time interval, for example, increasing by 15% every 5 minutes, the remaining traffic from the services to be migrated can be migrated to the target cluster nodes until the load of the nodes in the target cluster is balanced. In this way, the services to be migrated are equivalent to being migrated in a gradient of 10% → 25% → 40% → ... → balanced.
[0082] As can be seen from the implementation process, in the process of business migration, the present invention determines the business to be migrated in the target cluster, determines the target cluster node to receive the business to be migrated, migrates a preset percentage of the business traffic in the business to be migrated to the target cluster node, and migrates the remaining traffic of the business to be migrated to the target cluster node according to the migration rule of increasing the target percentage traffic weight every third time interval, until the node load of the target cluster is balanced. This is equivalent to migrating the business to be migrated according to the weight gradient, realizing smooth load migration and avoiding service jitter during the migration process.
[0083] To facilitate understanding of the present invention, let's assume that we need to scale up the number of nodes in a Kubernetes (k8s) cluster. Kubernetes is an open-source application used to manage containerized applications across multiple machines in a cloud platform. Kubernetes aims to make deploying containerized applications simple and powerful. Kubernetes provides a mechanism for application deployment, planning, updating, and maintenance. In a Kubernetes (k8s) cluster, nodes are the foundation for workload operation, divided into control plane nodes and worker nodes. Their functions are clearly defined, working together to support the operation and management of the cluster. In Kubernetes (k8s), scaling refers to the process of dynamically adjusting cluster resources according to business needs. This is mainly divided into two categories: horizontal scaling (increasing the number of Pod replicas to improve application processing capacity) and vertical scaling (increasing the resource requests (CPU, memory) of individual Pods to improve performance). A Pod is the smallest deployable and manageable unit in Kubernetes. Furthermore, node scaling, i.e., increasing the number of compute nodes, is also involved.
[0084] According to the present invention, the system used is as follows: Figure 5 As shown, the system adopts a four-layer hierarchical architecture to achieve full-process automation and functional decoupling. Each layer forms a closed loop through standardized communication and toolchain integration.
[0085] The data acquisition layer acts as the "sensory nerve" of the system, responsible for collecting indicators across all dimensions. After all data is standardized, it is transmitted to downstream modules through the Kafka message queue to provide a basis for decision-making. Specifically, a Node Exporter is deployed on all Kubernetes nodes via the DaemonSet controller, configured with parameters such as `--collector.systemd` to collect system-level metrics, and port 9100 is exposed for Prometheus to crawl. The DaemonSet is used to run specific Pod replicas on each node of the cluster. A cadvisor is deployed as a container metrics collector, integrated into the kubelet or deployed independently, monitoring core metrics such as container CPU utilization (`container_cpu_usage_seconds_total`) and memory working set (`container_memory_working_set_bytes`). Business metrics are exposed through application instrumentation, such as through Spring Boot Actuator or Prometheus Client. A custom Exporter is configured to collect data such as QPS (`http_requests_total`) and response time (`http_request_duration_seconds_bucket`), and unified access is provided to the Prometheus monitoring system. A Kafka cluster (with 3 redundant nodes, etc.) is deployed, and a `metrics-stream` topic is created for metric data forwarding, with message retention set to 72 hours.
[0086] The decision control layer assumes the "brain decision-making" function and includes two core components: an LSTM prediction model built on TensorFlow, which uses more than 7 days of historical data exported from Prometheus, integrates time features (hours, working days), historical load and business indicators to predict the load trend in the next 15-30 minutes; and a rule engine that generates decisions based on target performance thresholds, such as triggering expansion when the load exceeds the threshold and the predicted growth exceeds 20%.
[0087] The execution layer acts as the "execution hands and feet," connecting the cloud platform and cluster tools. It creates new nodes by calling the cloud platform API and completes the initialization of components such as Docker and kubelet. Docker is a containerization platform, and kubelet is a node agent in the Kubernetes cluster. It executes the kubeadm join command to connect to the Kubernetes cluster. kubeadm join is used to add new nodes to the existing Kubernetes cluster. At the same time, it achieves smooth traffic migration through Service weight control. Service is the core abstraction in Kubernetes used to expose applications and achieve service discovery and load balancing. The new node initially carries 10% of the traffic, which increases by 15% every 5 minutes until the load is balanced.
[0088] The feedback optimization layer is responsible for "effect calibration". By evaluating indicators such as the rate of change in response time after scaling, the improvement in resource utilization, and decision accuracy, it dynamically optimizes system parameters. For example, it retrains the LSTM model weekly, adjusts the safety factor monthly, automatically identifies and improves inefficient scaling modes, and forms a closed loop of continuous iteration.
[0089] The node expansion of a Kubernetes cluster can include the following processes: Obtain historical performance metrics for the Kubernetes cluster; Historical performance indicators were analyzed to obtain historical peak performance loads. Identify the target services for the target cluster, and determine the service types of the target services; If the business type is a core business, a security coefficient with a value greater than 1 is generated; if the business type is a non-core business, a security coefficient with a value less than or equal to 1 is generated. Based on the safety factor, the historical peak performance load is adjusted to obtain the initial performance threshold for managing the number of cluster nodes; Determine the target business of the target cluster; If the target service is a latency-sensitive service, the initial performance threshold is lowered to obtain the base performance threshold; if the target service is a non-latency-sensitive service, the initial performance threshold is used as the base performance threshold. A pre-trained long short-term memory network model is determined, which is trained based on performance metrics at different time periods; The future performance metrics are obtained by predicting the performance metrics of the target cluster using a long short-term memory network model. Analyze future performance metrics to determine the load change trend of performance metrics in the next hour; Determine the target business of the target cluster; generate a floating ratio that is positively correlated with the level of the target business; If the load change trend is upward, the base performance threshold is adjusted upward based on the floating ratio to obtain the target performance threshold; if the load change trend is downward, the base performance threshold is adjusted downward based on the floating ratio to obtain the target performance threshold. Obtain the target performance metrics of the target cluster over 30 minutes; Set a time window value of 5 minutes; collect target performance metrics according to the time window value to obtain the collected performance metrics; process the collected performance metrics to obtain the performance values of the target cluster to be evaluated within the time window value; return to execute the step of collecting target performance metrics according to the time window value until the collection of target performance metrics is completed and multiple performance values to be evaluated are obtained. Analyze future performance metrics to determine the growth rate of performance metrics in the next hour; If the performance value to be evaluated continues to exceed the target performance threshold and the increase is greater than 20%, the cluster nodes will be expanded. Identify the services to be migrated in the target cluster; Determine the target cluster node to receive the services to be migrated; Migrate 10% of the traffic from the services to be migrated to the target cluster node; Following the migration rule of increasing traffic weight by 15% every 5 minutes, the remaining traffic of the service to be migrated is migrated to the target cluster node until the target cluster node load is balanced.
[0090] To facilitate understanding of the effectiveness of this invention, a scenario of exceeding the load threshold is simulated using an API call rule engine to manually trigger expansion, verifying the entire process time of node creation, initialization, and cluster access, such as a target time of less than or equal to 3 minutes; automatic response testing is performed, for example, using the K6 tool to simulate a traffic surge from 100 QPS to 1000 QPS, checking whether the system triggers expansion within 5 minutes and whether the traffic allocation of the new nodes conforms to the gradient rules; stress testing is performed, simulating 5000 QPS traffic in a 10-node cluster, verifying that the prediction model accuracy is ≤15% and the expansion response time is ≤2 minutes; security verification is performed, checking RBAC permission configuration (the controller only has create / update permissions for node resources), API key storage (encrypted), and audit log integrity (including the time of node creation / deletion operations and operator information). After testing, the solution demonstrated that this invention improved cluster resource utilization by over 15%, significantly reducing cloud platform operating costs. It can respond to sudden traffic surges within 3 minutes, ensuring the stability of core services such as payment interfaces during peak traffic through P99 response time monitoring and priority scaling mechanisms, reducing the error rate to below 0.1%. It can replace over 70% of manual scaling operations, freeing the operations team from repetitive tasks and allowing them to focus on strategy optimization. The canary deployment strategy (non-production environment verification → 10% → 30% → 100% production traffic) reduces system deployment risks. Furthermore, it is compatible with mainstream cloud platforms and the container ecosystem, making it widely applicable to scenarios with drastic traffic fluctuations, such as e-commerce promotions, financial transactions, and online education. It provides a standardized elastic scaling solution for cloud-native architectures, offering excellent scalability and adaptability.
[0091] In other words, this invention dynamically adjusts the threshold for node expansion based on historical data and business needs, thereby automatically judging and executing node expansion operations. It can not only automatically adjust resource configuration under different business loads, but also respond to changes in the cluster environment in real time, significantly improving resource utilization efficiency. Ultimately, it achieves "zero-intervention elastic scaling" of cloud platform container clusters, enabling a dynamic balance between resource supply and business needs, greatly improving the economy and reliability of the cloud platform. Furthermore, it realizes a closed-loop process from indicator collection and intelligent decision-making to automatic execution, maximizing resource utilization efficiency while ensuring service stability.
[0092] This invention also provides an electronic device and a computer-readable storage medium, both of which have the corresponding effects of the cluster node management method provided in the embodiments of this invention. Please refer to... Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0093] An electronic device provided by an embodiment of the present invention includes a memory 201 and a processor 202. The memory 201 stores a computer program, and the processor 202 executes the computer program to implement the steps of the cluster node management method described in any of the above embodiments.
[0094] Please see Figure 7 Another electronic device provided in this embodiment of the invention may further include: an input port 203 connected to the processor 202 for transmitting commands input from the outside to the processor 202; a display unit 204 connected to the processor 202 for displaying the processing results of the processor 202 to the outside; and a communication module 205 connected to the processor 202 for enabling communication between the electronic device and the outside. The display unit 204 may be a display panel, a laser scanner, or the like; the communication method used by the communication module 205 includes, but is not limited to, Mobile High-Definition Link (MHL), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connectivity: Wireless Fidelity (WiFi), Bluetooth communication technology, Bluetooth Low Energy communication technology, and communication technology based on IEEE 802.11s.
[0095] The present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the cluster node management method described in any of the above embodiments.
[0096] The computer-readable storage media involved in this invention include random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs (compact disc read-only memory), or any other form of storage media known in the art.
[0097] The present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the cluster node management method described in any of the above embodiments.
[0098] For descriptions of relevant parts of the computer program product, electronic device, and computer-readable storage medium provided in this embodiment of the invention, please refer to the detailed description of the corresponding parts in the cluster node management method provided in this embodiment of the invention, which will not be repeated here. Furthermore, parts of the technical solutions provided in this embodiment of the invention that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.
[0099] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0100] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A cluster node management method, characterized in that, include: Obtain historical performance metrics for the target cluster; Based on the historical performance metrics, determine the basic performance threshold for managing the number of cluster nodes; Obtain future performance metrics after predicting the performance metrics of the target cluster; Based on the future performance indicators, the basic performance threshold is adjusted to obtain the target performance threshold; The number of cluster nodes is managed according to the target performance threshold and the future performance indicators.
2. The cluster node management method according to claim 1, characterized in that, Based on the historical performance metrics, determine the basic performance thresholds for managing the number of cluster nodes, including: The historical performance indicators were analyzed to obtain the historical performance load peak values; The historical performance load peaks are adjusted to obtain the initial performance threshold for managing the number of cluster nodes; The basic performance threshold is determined based on the initial performance threshold.
3. The cluster node management method according to claim 2, characterized in that, Adjusting the historical performance load peaks yields an initial performance threshold for managing the number of cluster nodes, including: Determine the target service of the target cluster; Generate a security factor to stabilize the target service; Based on the security factor, the historical performance load peak is adjusted to obtain the initial performance threshold for managing the number of cluster nodes.
4. The cluster node management method according to claim 3, characterized in that, Generating a security factor for stabilizing the target service includes: Determine the business type of the target business; If the business type is a core business, a security factor with a value greater than 1 is generated. If the business type is a non-core business, a security factor with a value less than or equal to 1 is generated.
5. The cluster node management method according to claim 2, characterized in that, Determining the basic performance threshold based on the initial performance threshold includes: Determine the target service of the target cluster; If the target service is a latency-sensitive service, the initial performance threshold is lowered to obtain the basic performance threshold. If the target service is a non-latency-sensitive service, then the initial performance threshold is used as the basic performance threshold.
6. The cluster node management method according to any one of claims 2 to 5, characterized in that, Based on the future performance indicators, the basic performance threshold is adjusted to obtain the target performance threshold, including: The future performance indicators are analyzed to determine the load change trend of the performance indicators in the first time period in the future; Based on the load change trend, the basic performance threshold is adjusted to obtain the target performance threshold.
7. The cluster node management method according to claim 6, characterized in that, Based on the load change trend, the basic performance threshold is adjusted to obtain the target performance threshold, including: Determine the target service of the target cluster; Generate a floating ratio that is positively correlated with the level of the target service; Based on the load change trend, the basic performance threshold is adjusted according to the floating ratio to obtain the target performance threshold.
8. The cluster node management method according to claim 7, characterized in that, Based on the load change trend and the floating ratio, the basic performance threshold is adjusted to obtain the target performance threshold, including: In response to the load change trend being upward, the basic performance threshold is adjusted upward based on the floating ratio to obtain the target performance threshold. In response to the load change trend being downward, the base performance threshold is adjusted downward based on the floating ratio to obtain the target performance threshold.
9. The cluster node management method according to claim 6, characterized in that, Analyzing the historical performance indicators yields historical performance load peaks, including: The historical performance indicators were analyzed to obtain the historical peak performance load. Adjusting the historical performance load peaks yields the basic performance thresholds for managing the number of cluster nodes, including: The historical peak performance load is adjusted to obtain the basic performance threshold for scaling up cluster nodes.
10. The cluster node management method according to claim 9, characterized in that, The number of cluster nodes is managed according to the target performance threshold and the future performance indicators, including: Obtain the target performance metrics of the target cluster within a preset time period; The target performance indicators are analyzed to obtain multiple performance values of the target cluster to be evaluated within the preset time period; The future performance indicators are analyzed to determine the growth rate of the performance indicators in the second future time period. If the performance value to be evaluated is continuously greater than the target performance threshold and the growth rate is greater than a first set value, then the cluster node is expanded.
11. The cluster node management method according to claim 6, characterized in that, Analyzing the historical performance indicators yields historical performance load peaks, including: Analyzing the historical performance indicators yields the historical low peak performance load. Adjusting the historical performance load peaks yields the basic performance thresholds for managing the number of cluster nodes, including: The historical low peak performance load is adjusted to obtain the basic performance threshold for scaling down cluster nodes.
12. The cluster node management method according to claim 11, characterized in that, The number of cluster nodes is managed according to the target performance threshold and the future performance indicators, including: Obtain the target performance metrics of the target cluster within a preset time period; The target performance indicators are analyzed to obtain multiple performance values of the target cluster to be evaluated within the preset time period; The future performance indicators are analyzed to determine the extent of their decrease over the next second time period. If the performance value to be evaluated is consistently lower than the target performance threshold and the decrease is greater than a second set value, then the cluster node is scaled down.
13. The cluster node management method according to claim 12, characterized in that, Analyzing the target performance metrics yields multiple performance values of the target cluster to be evaluated within the preset time period, including: Set the time window value; The target performance indicators are collected according to the time window value to obtain the collected performance indicators; The collected performance metrics are processed to obtain the performance value of the target cluster to be evaluated within the time window. Return to the step of collecting the target performance index according to the time window value, until the collection of the target performance index is completed, and obtain multiple performance values to be evaluated.
14. The cluster node management method according to claim 1, characterized in that, Obtaining future performance metrics after predicting the performance metrics of the target cluster includes: A pre-trained long short-term memory network model is determined, which is trained based on performance indicators for different time periods; The future performance metrics are obtained by predicting the performance metrics of the target cluster using the Long Short-Term Memory network model.
15. The cluster node management method according to claim 1, characterized in that, After managing the number of cluster nodes according to the target performance threshold and the future performance indicators, the method further includes: Identify the services to be migrated in the target cluster; Determine the target cluster node to receive the services to be migrated; Migrate a preset percentage of the service traffic in the service to be migrated to the target cluster node; According to the migration rule of increasing the target percentage traffic weight every third time interval, the remaining traffic of the service to be migrated is migrated to the target cluster node until the node load of the target cluster is balanced.
16. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the cluster node management method as described in any one of claims 1 to 15.
17. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the cluster node management method as described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the cluster node management method as described in any one of claims 1 to 15.
Citation Information
Patent Citations
Elastic scaling method based on Kubernetes cluster
CN114637650A
Method and device for supporting elastic scalability of computing power resources of intelligent computing center
CN120560858A