High-availability task level capacity expansion and contraction method for cluster in restricted environment

By introducing a dual-mode scaling mechanism and ant colony optimization algorithm in the K3s cluster, the load imbalance and scheduling lag of K3s in resource-constrained environments is solved, and intelligent scheduling is realized for multi-dimensional resource perception, ensuring the high availability and resource utilization of the cluster.

CN120371523APending Publication Date: 2025-07-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510486965.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The horizontal scaling mechanism of K3s has problems such as insufficient single task consideration, incomplete resource utilization and scheduling lag in resource-constrained environments, which affects the high availability of the cluster.

Method used

The dual-mode dynamic adjustment mechanism and intelligent pre-scheduling algorithm are adopted to optimize the number and distribution of task replicas through the horizontal expansion and capacity mode of single-task and global tasks, combined with the GRU algorithm and improved ant colony optimization algorithm.

Benefits of technology

Dynamic optimization of cluster load balancing and high task availability in resource-constrained environments is realized, which reduces computing overhead, improves the timeliness and accuracy of scaling, and ensures high availability and resource utilization of clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371523A_ABST
    Figure CN120371523A_ABST
Patent Text Reader

Abstract

The invention relates to the field of distributed cluster high availability, and particularly discloses a task level capacity expansion and contraction method oriented to cluster high availability in a limited environment. According to the method, optimization is achieved through a dual-mode dynamic adjustment mechanism and an intelligent pre-scheduling algorithm aiming at the problems of single-task capacity expansion and contraction limitation, global load unbalance, scheduling lag and the like of a K3s cluster in a resource-constrained environment. According to the technical scheme, high availability is guaranteed preferentially in a single-task mode, and copy adjustment is carried out by integrating node resource prediction in a global mode; constructing a lightweight resource prediction model based on a GRU algorithm, and performing UINT quantization processing on a CPU / memory to reduce calculation overhead; an improved ant colony optimization algorithm is used for pre-scheduling scoring, and multi-dimensional resource load balancing and task distribution optimization are achieved by dynamically adjusting pheromone weights and a heuristic matrix. According to the method, the fault disaster tolerance capability of the cluster in the limited environment is remarkably improved, and the number of copies is dynamically maintained in the optimal interval for guaranteeing high availability and resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high availability of distributed clusters, and particularly to a method for horizontally scaling tasks for high availability of clusters in restricted environments. Background Art

[0002] K3s is a lightweight version of Kubernetes designed to provide a simplified Kubernetes cluster solution for resource-constrained environments such as edge computing, Internet of Things devices, developer machines, etc. It is highly compatible with the original Kubernetes (K8s), but many non-core components and features are removed, making it more efficient in resource utilization and deployment. In terms of high availability, K3s provides replicas and distributed deployment. Users can use Kubernetes resources such as Deployment and StatefulSet to ensure that replicas of container applications are distributed across multiple nodes, avoiding application unavailability when a node or container instance fails. However, the default horizontal scaling mechanism of K3s only scales individual Pods and does not balance the high availability of Pods and the load balance of the cluster by considering the global resource situation and reasonably estimating the number of task replicas. In addition, the default horizontal scaling mechanism of K3s only considers CPU and memory, and is applicable to limited tasks.

[0003] In summary, in restricted environments, the horizontal scaling mechanism of K3s has the following disadvantages:

[0004] 1. The default horizontal scaling mechanism of K3s mainly starts from the high availability of the task itself. That is, when the task consumes too much CPU and memory resources due to factors such as the request volume, K3s will trigger horizontal expansion of the task, increasing the number of tasks to achieve the purpose of traffic diversion. When the resource consumption is below the threshold, horizontal scaling of the task is performed to save cluster resource overhead. In restricted environments, the main purpose of task replica redundancy is to prevent task failure after a cluster node fails, but excessive replica redundancy will cause the overloaded resource-limited cluster.

[0005] 2. The default horizontal scaling mechanism of K3s only targets individual tasks and cannot globally consider the impact of adjusting the number of task replicas on the cluster load and task high availability based on the resource situation of the entire cluster.

[0006] 3. The default horizontal scaling mechanism of K3s is based on CPU and memory resources, and sets thresholds for resource occupancy to perform scaling adjustments. The resource consideration is not comprehensive enough, which results in tasks with special resource requirements not being able to be scheduled to healthy nodes that meet the requirements.

[0007] 4. The default horizontal scaling mechanism of K3s is based on the real-time monitoring of resources by Metric-Server, and there will be a certain lag problem in the horizontal scaling decision-making and execution, that is, the resources change greatly after the decision-making and before the execution, thus affecting the actual horizontal scaling effect.

[0008] The disadvantages of the existing K3s horizontal scaling mechanism mentioned above will affect the high availability of the cluster in a restricted environment. Therefore, the present invention proposes a task horizontal scaling method for high availability of the cluster in a restricted environment. Summary of the Invention

[0009] The technical problem solved by the present invention is: in the resource-constrained K3s cluster environment, by predicting resource usage and dynamically optimizing the scheduling strategy, overcome the defects of the traditional horizontal scaling mechanism that only focuses on single tasks, single resource metrics and lagged scheduling, and achieve intelligent scaling that takes into account global load balancing and task high availability. Ensure that under the multi-dimensional resource constraints such as CPU / memory / disk / network I0, through quantitative prediction and improved ant colony algorithm pre-scheduling, dynamically optimize the number and distribution of replicas, and solve the balance problem of redundant resource occupation and node failure disaster tolerance.

[0010] In order to achieve the above object, the present invention adopts the following technical means:

[0011] The present invention provides a task horizontal scaling method for high availability of the cluster in a restricted environment, including the following steps:

[0012] Step S1: Select the task horizontal scaling mode, and adjust the replicas according to the replica adjustment strategy of single-task horizontal scaling or global task horizontal scaling;

[0013] Step S2: Generate a lightweight resource prediction model based on the historical resource usage data of the nodes, and predict the resource usage of the nodes according to this model;

[0014] Step S3: According to the predicted node resources, task resource occupancy information and scaling mode, use the improved ant colony optimization algorithm for task pre-scheduling, and score based on high availability and load balancing;

[0015] Step S4: Determine the optimal number of replicas and pre-scheduling plan according to the scoring results, and execute the actual scaling and scheduling of the task replicas.

[0016] In the above solution, step S1 includes the following steps:

[0017] Step S11: Select the single-task horizontal scaling mode or the global task horizontal scaling mode according to the cluster running environment, where the single-task mode gives priority to ensuring the high availability of the task, and the global mode takes into account both the cluster load balancing and the task high availability;

[0018] Step S111: In the single-task horizontal scaling mode, set the replica threshold interval based on the number of nodes, perform scheduling feasibility scoring on multiple candidate replica number schemes according to the pre-scheduling algorithm, and select the highest-scoring scheme that meets the node distribution redundancy requirements;

[0019] Step S112: In the global-task horizontal scaling mode, adjust the replicas according to the threshold specified by the number of nodes and the current number of replicas of the task. For all current tasks, adopt a triple replica adjustment strategy: keep the current number of replicas, increase one replica, or decrease one replica. Determine the optimal replica combination scheme that minimizes the variance of the cluster resource utilization through combined pre-scheduling calculation.

[0020] In the above solution, step S2 includes the following steps:

[0021] Step S21: Collect the CPU, memory, and hard disk resource usage data of all nodes in the cluster at a preset time interval and store it in the time series database;

[0022] Step S22: Build a prediction model based on the GRU neural network and perform 8-bit / 16-bit fixed-point quantization processing on the CPU and memory historical data: quantize the CPU occupancy rate from FP32 floating-point type to UINT8 integer type, and quantize the memory occupancy from FP32 floating-point type to UINT16 integer type;

[0023] Step S23: Use the quantized historical data to train the GRU model to generate a lightweight resource prediction model. When predicting, perform the same quantization conversion on the input data and then input it into the model to output the CPU, memory, and hard disk resource prediction values of the nodes.

[0024] In the above solution, step S3 includes the following steps:

[0025] Step S31: Obtain the predicted values of the node CPU, memory, and hard disk output by the resource prediction model, and combine the task CPU / memory occupancy data monitored by metrics-server and the disk requirements and task types labeled by the task tags;

[0026] Step S32: Use the weighted fusion algorithm to calculate the task resource requirements. The formula is:

[0027]

[0028] Among them, CPU max , Mem max represent the upper bounds of CPU and memory set by the system, CPU min , Mem min represent the lower bounds of CPU and memory set by the system, CPU now , Mem now represent the current CPU and memory, ω and is an adjustable parameter, where

[0029] Step S33: Perform task pre-scheduling under the current task copy based on the ant colony optimization algorithm, and score according to the task scaling mode.

[0030] In the above solution, step S33 includes the following steps:

[0031] Step S331: Initialize the number of ants, the number of iterations, the pheromone matrix, the pheromone weight, the information emission coefficient, the heuristic information matrix, and the heuristic information weight parameter;

[0032] Step S332: The ants complete path selection and score based on high availability and load. If the scheduling fails, the score is set to 0, indicating that no scheduling path is found. The total score S e Based on the replica redundancy score C i and the overall cluster load score S l Perform weighted scoring, and the scoring formula is as follows:

[0033]

[0034] where, S e represents the total score, and the total score is based on the replica redundancy score C i and the overall cluster load score S l Perform weighted scoring, μ and are adjustable parameters, S l = 0 indicates that the scheduling cannot be completed under the current replica conditions, and the total score S e is set to 0;

[0035] r represents the number of replicas, k is a positive constant proportional to the score growth rate, h is the threshold upper limit, h / 2 is the critical point where the replica number growth slows down, and C is a constant;

[0036] D represents the decay parameter, ∑ represents the mean square deviation of resources, a, b, c, d are adjustable parameters, and IO j represents the network IO situation of the node carrying task i, and IO max is the IO threshold defined by the algorithm;

[0037] Step S333: After a single ant colony iteration, record the highest score, the pre-scheduling result of the highest score, and update the pheromone. Use the priority queue method to record the selection of the top few ants with higher scores and enhance the pheromone according to these high-quality ants. The global pheromone evaporation formula and the pheromone enhancement formula are as follows:

[0038] τ ij = (1 - ρ)τij

[0039]

[0040] Among them, τ ij represents the value of each point in the pheromone matrix, ω is an adjustable parameter, and S p represents the path score of high-quality ants, Q represents the pheromone enhancement parameter, and Q min represents the minimum information enhancement parameter, and Q max represents the maximum information enhancement parameter, r represents the current iteration number, and T represents the total number of iterations;

[0041] Step S334: Perform the ant colony search iteration according to the initialized number of iterations for steps S332 and S333 until the number of times the optimal solution change threshold stabilizes within a suitable range reaches the threshold or the iteration is completed, and finally output the optimal solution to obtain the optimal scheduling plan and score under the current copy number.

[0042] In the above solution, step S331 includes the following steps:

[0043] Step S3311: Initialize the number of ants. The formula is as follows:

[0044]

[0045] Among them, N represents the number of ants for calculation, t represents the number of tasks, n represents the number of nodes, θ and represent adjustment parameters, l represents the load of the node where the pre-scheduling algorithm is located, and this load is mainly determined by the usage rates of the node's CPU and memory. The calculation of l is as follows:

[0046]

[0047] Among them, CPU idle and Memory idle represent the idle CPU and idle memory, and CPU all and Memory all represent the total amount of the node's CPU and memory;

[0048] Dynamically calculate the number of ants according to the node load:

[0049]

[0050] Step S3312: Initialize the number of iterations. The number of iterations is set between 100 and 2000 times. For task pre-scheduling, analyze and judge the number of iterations according to the total amount of tasks and nodes. The calculation of the number of iterations is as follows:

[0051] T p = a·n + b·t

[0052]

[0053] Among them, T p represents the number of iterations calculated during initialization, n represents the number of nodes, t represents the number of tasks, a and b are adjustment parameters, and T represents the number of iterations during actual initialization;

[0054] S3313: Initialize the pheromone weight α and the pheromone evaporation coefficient ρ:

[0055]

[0056] Among them, T i represents the number of iterations after the i-th iteration, T represents the total number of iterations, and δ and σ represent adjustable parameters;

[0057] Step S3314: Initialize the heuristic information matrix η ij , and the nodes of the heuristic information matrix are composed of the score S ij as follows:

[0058] S ij = u·S1 ij + v·S2 ij

[0059] Among them, S1 ij is the load score of the hard disk, S2 ij is the IO score of the network IO type task, and u + v = 1;

[0060] Based on the principle of strengthening exploration in the early stage and accelerating convergence in the later stage, initialize the heuristic information weight β, and the initialization formula is as follows:

[0061]

[0062] In the above solution, step 4 includes the following steps:

[0063] Step S41: Adjust the number of task replicas according to the S1 replica strategy;

[0064] Step S42: After adjusting the number of replicas, use the node resources predicted by S2 for the task pre-scheduling algorithm in S3 and record the highest score and the scheduling scheme with the highest score;

[0065] Step S43: Repeat the process of step S41 and step S42 until the replica strategy of S1 is completed;

[0066] Step S44: Perform actual horizontal scaling of tasks according to the number of replicas and the pre-scheduling scheme with the highest score, and complete the actual scheduling or deletion of task replicas based on the node affinity of K3s.

[0067] Since the present invention adopts the above technical means, it has the following beneficial effects:

[0068] 1. When the cluster has a high load due to limited resources or node downtime, or when the cluster load is low, task horizontal scaling will be triggered to keep the number of task replicas in the cluster within a range that can ensure both the load balance of the cluster and the high availability of the business function.

[0069] 2. Different task horizontal scaling strategies are set for single tasks and global tasks. The global task horizontal scaling strategy pays more attention to balancing the load of the cluster and the availability of the business function, while the single task horizontal scaling mainly considers the high availability of the task.

[0070] 3. The method implements a pre-scheduling algorithm based on CPU, memory, disk resources, etc. This algorithm can not only ensure the load balance of CPU, memory, and disk, but also consider special tasks of network IO type, enabling different types of tasks to be scheduled to ideal nodes.

[0071] 4. The method uses a resource prediction algorithm based on the GRU algorithm to predict the node resources in the cluster, and reduces the resource occupancy of the prediction algorithm for CPU and memory through data quantization, solving the problem of lag in the execution of task horizontal scaling under the premise of being relatively lightweight. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a block diagram of a task horizontal scaling method for cluster high availability in a restricted environment provided for the implementation of the present invention.

[0073] Figure 2 It is a schematic structural diagram of a task horizontal scaling method for cluster high availability in a restricted environment provided for the implementation of the present invention.

[0074] Figure 3 It is a flowchart of a task horizontal scaling method for cluster high availability in a restricted environment provided for the implementation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0075] The following will give a detailed description of the embodiments of the present invention. Although the present invention will be described and explained in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments only. On the contrary, any modifications or equivalent replacements made to the present invention should be covered within the scope of the claims of the present invention.

[0076] In addition, in order to better illustrate the present invention, numerous specific details are given in the following detailed description. Those skilled in the art will understand that the present invention can be implemented without these specific details.

[0077] Based on the above problems, the present invention proposes a method for task horizontal scaling in a cluster with high availability in a restricted environment, including the following steps:

[0078] S1: Select the task horizontal scaling mode and adjust the replicas according to the corresponding replica adjustment strategy, which is divided into single-task horizontal scaling and global task horizontal scaling.

[0079] S2: Generate a lightweight resource prediction model based on the historical usage of nodes, and predict the resource usage of nodes according to this model.

[0080] S3: Perform task pre-scheduling and scoring according to the predicted node resource situation, the average resource occupancy of tasks, and the replica adjustment strategy in the task horizontal scaling mode.

[0081] S4: Complete task replica adjustment and pre-scheduling according to the replica adjustment strategy, record the best replica and pre-scheduling plan, and perform task scheduling according to this plan.

[0082] In the above technical solution, S1 is specifically described as:

[0083] S11: The task horizontal scaling mode is mainly divided into two types: single-task horizontal scaling and global task horizontal scaling. Among them, the single-task horizontal scaling has limited impact on the overall load of the cluster, so the high availability of the task is mainly considered. While the global task scaling mainly considers the overall load situation of the cluster and the high availability of the task itself in a balanced manner.

[0084] S111: The single-task horizontal scaling mainly stipulates a threshold according to the number of nodes and scores the replica adjustment attempts according to the pre-scheduling algorithm in S2, and selects the best plan to determine the number of replicas and the scheduling plan.

[0085] S112: The global task horizontal scaling mainly adjusts the replicas according to the threshold stipulated by the number of nodes and the current number of replicas of the task. In the case of limited resources, to ensure that the replica adjustment is completed quickly, the global task horizontal scaling makes combination attempts in three cases: the current number of task replicas remains unchanged, increases by one, and decreases by one, and scores according to the pre-scheduling algorithm in S2, selects the best plan, and determines the number of replicas and the scheduling plan.

[0086] In the above technical solution, S2 is specifically described as:

[0087] S21: Collect the CPU, memory, and hard disk occupancy of all nodes in the cluster at regular intervals and store them in InfluxDB.

[0088] S22: According to the historical resource situation of the node, the node historical resource information is passed into the GRU prediction algorithm to form a resource prediction algorithm model. Among them, aiming at the numerical range and precision requirements of memory and CPU in a restricted environment, quantization processing is performed on memory and CPU to reduce the resource consumption of the resource prediction algorithm. Since the numerical range of the hard disk is large, quantization processing is not performed.

[0089] S221: Perform quantization processing on CPU resources, quantizing from FP32 to UINT8, and perform quantization processing on memory resources, quantizing from FP32 to UINT16.

[0090] S23: Perform resource prediction on CPU, memory, and hard disk according to the resource prediction algorithm model. During prediction, CPU and memory are quantized according to the steps of S221 to reduce resource occupancy during prediction.

[0091] In the above technical solution, the specific description of S3 is as follows:

[0092] S31: Obtain the future usage of CPU, memory, and hard disk of the node according to the resource prediction algorithm, obtain the CPU and memory occupancy information of the task according to metrics-server, and obtain the disk occupancy and task type information of the task according to the label.

[0093] S32: Process the CPU and memory occupancy information of the task. The formula is as follows:

[0094]

[0095] Among them, CPU max , Mem max represent the upper bounds of CPU and memory set by the system, CPU min , Mem min represent the lower bounds of CPU and memory set by the system, CPU now , Mem now represent the current CPU and memory, ω and are adjustable parameters. Among them when ω is greater than , the average value scheme restricted by the system is mainly adopted for CPU and memory. when it is greater than ω, the values of CPU and memory are mainly taken according to the current CPU and memory usage. Generally, ω is taken as 0.5.

[0096] S33: Perform task pre-scheduling under the current task replica based on the ant colony optimization algorithm and score according to the task scaling mode.

[0097] S331: Initialize relevant parameters such as the number of ants, the number of iterations, the pheromone matrix, the pheromone weight, the information evaporation coefficient, the heuristic information matrix, and the heuristic information weight.

[0098] S3311: Initialize the number of ants. The formula is as follows:

[0099]

[0100] where N p represents the number of ants for calculation, t represents the number of tasks, n represents the number of nodes, θ and μ are adjustment parameters for conveniently adjusting the number of ants for calculation. l represents the load of the node where the pre-scheduling algorithm is located, and this load is mainly determined by the usage rates of the node's CPU and memory. The calculation of l is as follows:

[0101]

[0102] where, CPU idle and Memory idle represent the idle CPU and idle memory, and CPU all and Memory all represent the total amount of the node's CPU and memory.

[0103] N p Theoretically, as the number of ants for calculation, N, can increase infinitely, but the actual number of ants needs to be maintained within a reasonable range to ensure the execution efficiency of the algorithm. The formula for the actual number of ants is as follows:

[0104]

[0105] S3312: Initialize the number of iterations. Generally, it is common to set the number of iterations between 100 and 2000. For task pre-scheduling, the number of iterations can be analyzed and judged based on the total amounts of tasks and nodes. The calculation of the number of iterations is as follows:

[0106] T p = a·n + b·t

[0107]

[0108] where, T p represents the number of iterations calculated during initialization, n represents the number of nodes, t represents the number of tasks, a and b are adjustment parameters, and T represents the actual number of iterations during initialization.

[0109] S3313: Initialize the pheromone weight α and the pheromone evaporation coefficient ρ. At the beginning of the iteration, the pheromone weight should be reduced and the pheromone evaporation rate should be increased to enhance the random search ability. In the later stage of the iteration, the pheromone weight needs to be increased and the pheromone evaporation rate should be reduced to accelerate convergence to the optimal solution. The initialization formula is as follows:

[0110]

[0111] Among them, \(T_i\) i represents the number of iterations after the \(i\)-th iteration, \(T\) represents the total number of iterations. The initialization method of adding pheromone based on logarithm can significantly improve the pheromone concentration in the middle stage, facilitating rapid convergence in the later stage. \(\delta\) and \(\sigma\) represent adjustable parameters.

[0112] S3314: Initialize the heuristic information matrix \(\eta\) ij , and the nodes of the heuristic information matrix are composed of scores \(S\) ij which includes the load scores \(S_1\) of CPU, memory, and hard disk ij and the IO score \(S_2\) of network IO type tasks ij in two parts. The initialization formula is as follows:

[0113] \(S\) ij = \(u\cdot S_1\) ij + \(v\cdot S_2\) ij

[0114]

[0115] Among them, CPU j , MEM j , DISK j represent the free amounts of CPU, memory, and hard disk of node \(j\), and \(CPU^i\) i , \(MEM^i\) i , \(DISK^i\) i represent the demands of CPU, memory, and hard disk of task \(i\). \(u\), \(v\), \(a\), \(b\), \(c\) are adjustable parameters, \(u + v = 1\). If it is a network IO type task, then the proportion of \(v\) is higher, otherwise \(v = 0\). \(a + b + c = 1\), and according to the task type, the parameter proportions are different. If it is a computing task, then the proportion of \(a\) is larger; for large memory tasks, the proportion of \(b\) is larger; for hard disk storage tasks, the proportion of \(c\) is larger. \(IO_{th}\) max represents the upper threshold of network IO per second, and \(IO_j\) j represents the current network IO per second of node \(j\).

[0116] Based on the principle of strengthening exploration in the early stage and accelerating convergence in the later stage, initialize the heuristic information weight \(\beta\), and the initialization formula is as follows:

[0117]

[0118] S332: The ant completes path selection and scores based on high availability and load. If the scheduling fails, it is set to 0 points, indicating that no scheduling path is found. The total score \(S\) e is mainly weighted and scored based on the replica redundancy score \(C\) i and the overall cluster load score \(S\) l , and the scoring formula is as follows:

[0119]

[0120]

[0121] Among them, S e represents the total score, and the total score is based on the replica redundancy score C i and the overall cluster load score S l for weighted scoring. μ and are adjustable parameters, for the replica scaling of single tasks, the weight μ of the number of replicas is larger. For the global replica scaling of multi-tasks, more attention should be paid to the overall cluster load. Therefore is larger. In addition, when S l = 0, it indicates that scheduling cannot be completed under the current replica conditions, and the total score S e is set to 0.

[0122] r represents the number of replicas, k is a positive constant, proportional to the score growth rate, h is set as the threshold upper limit, h / 2 is the critical point where the growth of the number of replicas slows down, and C is a constant used to adjust the score offset. Under this formula, when the number of replicas is below h / 2, the score grows rapidly, reflecting the high-availability effect brought by replica redundancy. It reaches 70 points at h / 2 and then grows slowly, that is, after the replicas grow to a certain extent, the improvement of availability is no longer as large as that of low-task replicas.

[0123] D represents the attenuation parameter, used to control the attenuation error during CPU, memory, and hard disk calculations. When the mean square error is too high, a higher score can be obtained by reducing D. ∑ represents the mean square error of resources, used as a standard to measure its load balancing. a, b, c, and d are adjustable parameters, and the score weights can be adjusted according to the task type. IO j represents the network IO situation of the node carrying task i, and IO max is the IO threshold defined by the algorithm.

[0124] S333: After a single ant colony iteration, record the pre-scheduling results of the highest score and the highest score and update the pheromone. In this paper, the priority queue method is adopted to record the selection of the first few ants with higher scores and enhance the pheromone based on these high-quality ants. The global pheromone evaporation formula and the pheromone enhancement formula are as follows:

[0125] τ ij = (1 - ρ)τ ij

[0126]

[0127] Among them, τ ij represents the value of each point in the pheromone matrix. ω is an adjustable parameter used to adjust the score Sp Perform mapping to map pheromone enhancement to an appropriate range. S p represents the path score of high-quality ants, Q represents the pheromone enhancement parameter, Q min represents the minimum information enhancement parameter, Q max represents the maximum information enhancement parameter, r represents the current iteration number, and T represents the total number of iterations.

[0128] S334: Perform the ant colony search iteration according to the initialized number of iterations for steps S332 and S333 until the number of times the optimal solution change threshold stabilizes within an appropriate range reaches the threshold or the iteration is completed, and finally output the optimal solution to obtain the optimal scheduling plan and score under the current replica number.

[0129] In the above technical solution, S4 is specifically described as:

[0130] S41: Adjust the number of task replicas according to the S1 replica strategy

[0131] S42: After adjusting the number of replicas, use the node resources predicted by S2 to perform the task pre-scheduling algorithm in S3 and record the highest score and the scheduling plan with the highest score.

[0132] S43: Repeat the processes of S41 and S42 until the replica strategy of S1 is completed.

[0133] S44: Perform actual horizontal scaling of tasks according to the replica number and pre-scheduling plan with the highest score, and complete the actual scheduling or deletion of task replicas based on the node affinity of K3s.

[0134] Experimental example:

[0135] The present invention will be further described below in conjunction with specific embodiments.

[0136] Set up a cluster consisting of 10 nodes. The cluster is built based on K3s. To ensure high availability in a restricted environment, set 3 nodes as Server nodes and the rest as Agent nodes, and set up an embedded etcd cluster to ensure high availability of storage. The node configurations are shown in the following table:

[0137]

[0138] Implement two schemes: perform single-task horizontal scaling for a single task and horizontal scaling for all tasks under the same namespace. The scaling steps are as follows:

[0139] S1: Select the task horizontal scaling mode and perform replica adjustment according to the corresponding replica adjustment strategy, which is divided into single-task horizontal scaling and global task horizontal scaling. The task replica threshold is taken Nodes represents the total number of nodes.

[0140] S2: Generate a lightweight resource prediction model based on the historical usage of nodes, and predict the resource usage of nodes according to this model.

[0141] S3: Perform task pre-scheduling and scoring according to the predicted node resource situation, the average resource occupancy of tasks, and the replica adjustment strategy in the task horizontal scaling mode.

[0142] S4: Complete task replica adjustment and pre-scheduling according to the replica adjustment strategy, record the best replica and pre-scheduling plan, and perform task scheduling according to this plan.

[0143] S1 is specifically described as:

[0144] S11: The task horizontal scaling mode is mainly divided into two types: single-task horizontal scaling and global task horizontal scaling. Among them, the single-task horizontal scaling has limited impact on the overall load of the cluster, so the high availability of tasks is mainly considered. While the global task scaling mainly balances the overall load of the cluster and the high availability of the tasks themselves.

[0145] S111: Single-task horizontal scaling mainly stipulates thresholds according to the number of nodes and scores the replica adjustment attempts according to the pre-scheduling algorithm in S2, and selects the best plan to determine the number of replicas and the scheduling plan.

[0146] S112: Global task horizontal scaling mainly adjusts replicas according to the thresholds stipulated by the number of nodes and the current number of replicas of the task. In the case of limited resources, to ensure the faster completion of replica adjustment, the global task horizontal scaling makes combined attempts in three cases: the current number of task replicas remains unchanged, increases by one, and decreases by one, and scores according to the pre-scheduling algorithm in S2, selects the best plan, and determines the number of replicas and the scheduling plan.

[0147] S2 is specifically described as:

[0148] S21: Collect the CPU, memory, and hard disk occupancy of all nodes in the cluster at regular intervals and store them in influxDB.

[0149] S22: According to the historical resource situation of nodes, input the node historical resource information into the GRU prediction algorithm to form a resource prediction algorithm model. Among them, for the numerical range and precision requirements of memory and CPU in a restricted environment, memory and CPU are quantized to reduce the resource consumption of the resource prediction algorithm, while the hard disk has a large numerical range and is not quantized.

[0150] S221: Quantize the CPU resources from FP32 to UINT8 and quantize the memory resources from FP32 to UINT16.

[0151] S23: Perform resource prediction for CPU, memory, and hard disk according to the resource prediction algorithm model. When predicting, quantize the CPU and memory according to the steps in S221 to reduce resource occupancy during prediction.

[0152] The specific description of S3 is as follows:

[0153] S31: Obtain the future usage of CPU, memory, and hard disk of the node according to the resource prediction algorithm, obtain the CPU and memory occupancy information of the task according to the metrics-server, and obtain the disk occupancy and task type information of the task according to the labels.

[0154] S32: Process the CPU and memory occupancy information of the task.

[0155] S33: Perform task pre-scheduling under the current task replica based on the ant colony optimization algorithm and score according to the task scaling mode.

[0156] S331: Initialize relevant parameters such as the number of ants, the number of iterations, the pheromone matrix, the pheromone weight, the pheromone evaporation coefficient, the heuristic information matrix, and the heuristic information weight.

[0157] S3311: Initialize the number of ants.

[0158] S3312: Initialize the number of iterations. Generally speaking, it is more common to set the number of iterations between 100 and 2000 times. For task pre-scheduling, the number of iterations can be analyzed and judged according to the total amount of tasks and nodes.

[0159] S3313: Initialize the pheromone weight α and the pheromone evaporation coefficient ρ. In the initial stage of iteration, the pheromone weight should be reduced and the pheromone evaporation rate should be increased to enhance the random search ability. In the later stage of iteration, the pheromone weight needs to be increased and the pheromone evaporation rate needs to be reduced to accelerate convergence to the optimal solution.

[0160] S3314: Initialize the heuristic information matrix η ij 。

[0161] S332: The ants complete path selection and score based on high availability and load. If the scheduling fails, it is set to 0 points, indicating that no scheduling path is found.

[0162] S333: After a single ant colony iteration, record the highest score, the pre-scheduling result of the highest score, and update the pheromone. In this paper, the priority queue method is adopted to record the selection of the top few ants with higher scores and enhance the pheromone according to these high-quality ants.

[0163] S334: Conduct the ant colony search iteration according to the initialized number of iterations for steps S332 and S333 until the number of times the optimal solution change threshold stabilizes within a suitable range reaches the threshold or the iteration is completed, and finally output the optimal solution to obtain the optimal scheduling plan and score under the current replica count.

[0164] S4 is specifically described as follows:

[0165] S41: Adjust the number of task replicas according to the S1 replica strategy

[0166] S42: After adjusting the number of replicas, use the node resources predicted by S2 to perform the task pre-scheduling algorithm in S3 and record the highest score and the scheduling plan with the highest score.

[0167] S43: Repeat the processes of S41 and S42 until the execution of the S1 replica strategy is completed.

[0168] S44: Perform actual horizontal scaling of the tasks according to the replica count and pre-scheduling plan with the highest score, and complete the actual scheduling or deletion of task replicas based on the node affinity of K3s.

[0169] Test Description

[0170] The core of the present invention lies in introducing the horizontal scaling of global tasks and using an improved ant colony optimization algorithm to schedule task replicas. Briefly, under the replica count specified by the task horizontal scaling strategy, calculate whether there is a possibility of replica scheduling based on task resources and node resources, as well as the optimal scheduling plan for the current replicas. Adjust the number of replicas through the horizontal scaling strategy and perform pre-scheduling calculations until the attempt is completed to obtain the optimal replica solution and scheduling solution. This verification test mainly tests the replica scaling of single tasks and global tasks to verify whether it meets the high-availability requirements of business functions and the load balancing requirements of the cluster.

[0171] Perform replica scaling on single tasks through the above steps, and select different task sizes for scaling. The task resource occupancy mainly includes CPU, memory, and disk. The test results of single-task replica scaling for different requirements are as follows:

[0172] Task CPU / % Memory / MB Disk / MB Network I / O Type Original Replica Number Current Replica Number task1 0.1 200 532 No 1 7 task2 0.2 1500 1421 No 1 6 task3 0.15 1150 893 Yes 1 4

[0173] It can be seen that the horizontal scaling of single tasks generally ensures the high availability of tasks mainly according to the number of nodes, because the scaled replicas of single tasks are distributed on different nodes and the resource demand nodes can be satisfied.

[0174] For horizontal scaling of all tasks under the same namespace, 10 different tasks are tested and placed in the same namespace. The average CPU occupancy is 0.2, the memory occupancy is 750MB, and the disk occupancy is 600MB. Perform global task horizontal scaling according to the above steps, and finally calculate the load balancing degree of the cluster according to the following formula

[0175]

[0176] where L i represents the resource load of node i, and μ represents the average load of the cluster. The resource load mainly includes CPU load, memory load, and hard disk load. The number of replicas and the load balancing degree of the cluster can meet the requirements of high availability of business functions and the requirements of cluster load balancing, and there is a certain margin of resources in the cluster.

[0177] In summary, the present invention has the following characteristics:

[0178] The method for horizontal scaling of tasks for high availability of clusters in restricted environments provided by the present invention has the following significant advantages:

[0179] 1. Dynamically balance cluster load and high availability

[0180] By introducing a dual-mode scaling mechanism for single tasks and global tasks, intelligent decision-making is achieved in resource-constrained scenarios. The single-task mode gives priority to ensuring the high availability of critical services and ensuring service continuity in case of failures; the global mode dynamically adjusts the replica distribution through multi-dimensional resource evaluation, and realizes load balancing optimization across tasks under node resource constraints, effectively avoiding cluster performance bottlenecks caused by local resource overload.

[0181] 2. Intelligent scheduling with multi-dimensional resource awareness

[0182] Break through the traditional decision-making mode based only on CPU / memory thresholds, innovatively integrate key indicators such as disk I / O and network throughput, and construct a multi-dimensional resource evaluation system. Through the dynamic weight adjustment mechanism of the ant colony optimization algorithm, the optimal scheduling strategy can be automatically adapted according to the task type (computation-intensive, storage-sensitive, network IO-intensive, etc.), significantly improving the node adaptability of heterogeneous tasks.

[0183] 3. Forward-looking resource scheduling decision-making ability

[0184] Adopt a lightweight GRU prediction model to perform time-series prediction on node resources, and combine quantization compression technology to reduce the model calculation overhead. Through the pre-scheduling mechanism, the scaling decision is advanced, effectively overcoming the lag defect of traditional real-time monitoring, enabling the system to complete the optimal scheduling plan in advance based on the predicted resource status, and ensuring the timeliness and accuracy of the scaling operation.

[0185] 4. Elastic and Scalable Optimization Algorithm Architecture

[0186] Design an adaptive dynamic adjustment mechanism for ant colony algorithm parameters to intelligently balance exploration and convergence efficiency during the iteration process. By introducing an ant quantity control strategy that senses node load and a phased pheromone update rule, it significantly reduces the consumption of computing resources while ensuring scheduling quality, and is particularly suitable for low-computing-power environments such as edge devices.

[0187] 5. System-Level Fault Tolerance and Resource Protection Mechanism

[0188] Establish a feasibility verification mechanism for scaling operations through a pre-scheduling scoring system. Pre-evaluate the remaining node resources and the rationality of distribution before replica adjustment to effectively avoid resource fragmentation problems caused by invalid scaling operations. Combine with the native affinity strategy of K3s to achieve seamless scheduling migration, ensuring that business continuity is not affected by the scaling process.

[0189] Through the above innovative mechanisms, this method achieves the Pareto optimality of cluster resource utilization and task availability in a restricted environment, providing a highly robust container orchestration solution for scenarios such as edge computing and the Internet of Things.

Claims

1. A task horizontal scaling method for high availability of clusters in restricted environments, characterized in that, It includes the following steps: Step S1: Select the task-level scaling mode, and adjust the replicas according to the replica adjustment strategy for single-task-level scaling or global-task-level scaling; Step S2: Generate a lightweight resource prediction model based on the historical resource usage data of the nodes, and predict the resource usage of the nodes according to this model; Step S3: According to the predicted node resources, task resource occupancy information, and scaling mode, use an improved ant colony optimization algorithm for task pre-scheduling, and score based on high availability and load balancing; Step S4: Determine the optimal number of replicas and pre-scheduling plan according to the scoring results, and execute the actual scaling and scheduling of task replicas.

2. The method according to claim 1, wherein: Step S1 includes the following steps: Step S11: Select the single-task-level scaling mode or the global-task-level scaling mode according to the cluster running environment, where the single-task mode gives priority to ensuring the high availability of tasks, and the global mode takes into account both cluster load balancing and task high availability; Step S111: In the single-task-level scaling mode, set the replica threshold interval based on the number of nodes, score the scheduling feasibility of multiple candidate replica number schemes according to the pre-scheduling algorithm, and select the highest-score scheme that meets the node distribution redundancy requirements; Step S112: In the global-task-level scaling mode, adjust the replicas according to the threshold specified by the number of nodes and the current number of replicas of the task. For all current tasks, adopt a triple replica adjustment strategy: keep the current number of replicas, add one replica, or reduce one replica, and determine the optimal replica combination scheme that minimizes the variance of cluster resource utilization through combinatorial pre-scheduling calculation.

3. The method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: Collect the CPU, memory, and hard disk resource usage data of all nodes in the cluster at a preset time interval and store it in the time series database; Step S22: Build a prediction model based on the GRU neural network, and perform 8-bit / 16-bit fixed-point quantization processing on the CPU and memory historical data: quantize the CPU occupancy rate from FP32 floating-point type to UINT8 integer type, and the memory occupancy from FP32 floating-point type to UINT16 integer type; Step S23: Use the quantized historical data to train the GRU model to generate a lightweight resource prediction model. When predicting, perform the same quantization conversion on the input data and then input it into the model to output the predicted values of the CPU, memory, and hard disk resources of the nodes.

4. The method according to claim 1, wherein: Step S3 includes the following steps: Step S31: Obtain the predicted values of the node CPU, memory, and hard disk output by the resource prediction model, and combine the task CPU / memory occupancy data monitored by metrics-server in real time and the disk requirements and task types labeled by the task tags; Step S32: Use the weighted fusion algorithm to calculate the task resource requirements. The formula is: Among them, CPU max , Mem max represent the upper bounds of CPU and memory set by the system, CPU min , Mem min represent the lower bounds of CPU and memory set by the system, CPU now , Mem now represent the current CPU and memory, ω and are adjustable parameters, where Step S33: Perform task pre-scheduling for the current task replicas based on the ant colony optimization algorithm, and score according to the task scaling mode.

5. The method according to claim 4, wherein: Step S33 includes the following steps: Step S331: Initialize the number of ants, number of iterations, pheromone matrix, pheromone weight, information evaporation coefficient, heuristic information matrix, and heuristic information weight parameters; Step S332: The ant completes path selection and scores based on high availability and load. If the scheduling fails, it is set to 0 points, indicating that no scheduling path is found, and the total score is S e Based on the replica redundancy score C i and the overall cluster load score S l Perform weighted scoring, and the scoring formula is as follows: Among them, S e represents the total score, and the total score is based on the replica redundancy score C i and the overall cluster load score S l for weighted scoring, where μ and are adjustable parameters, When S l = 0, it indicates that scheduling cannot be completed under the current replica conditions, and the total score S e is set to 0; r represents the number of replicas, k is a positive constant proportional to the score growth rate, h is the upper threshold, h / 2 is the critical point where the growth rate of the number of replicas slows down, and C is a constant; D represents the attenuation parameter, ∑ represents the mean square deviation of resources, and a, b, c, and d are adjustable parameters. IO j represents the network IO situation of the node carrying task i. IO max is the IO threshold defined for the algorithm; Step S333: After a single ant colony iteration, record the highest score, the pre-scheduling result of the highest score, and update the pheromone. Using the priority queue method, record the choices of the top several ants with higher scores and enhance the pheromone based on these high-quality ants. The global pheromone evaporation formula and the pheromone enhancement formula are as follows: τ ij = (1 - ρ)τ ij Among them, τ ij represents the value of each point in the pheromone matrix, ω is an adjustable parameter, S p represents the path score of high-quality ants, Q represents the pheromone enhancement parameter, Q min represents the minimum information enhancement parameter, Q max represents the maximum information enhancement parameter, r represents the current iteration number, and T represents the total number of iterations; Step S334: Perform ant colony search iterations according to the initialized number of iterations for steps S332 and S333 until the number of times the optimal solution change threshold stabilizes within a suitable range reaches the threshold or the iteration is completed. Finally, output the optimal solution to obtain the optimal scheduling plan and score under the current number of replicas.

6. The method according to claim 5, wherein: Step S331 includes the following steps: Step S3311: Initialize the number of ants, and the formula is as follows: where N p represents the number of calculations of the ants, t represents the number of tasks, n represents the number of nodes, θ and μ represent adjustment parameters, l represents the load of the node where the pre-scheduling algorithm is located, and this load is mainly determined by the usage rates of the node's CPU and memory. The calculation of l is shown as follows: Among them, CPU idle and Memory idle represent idle CPU and idle memory, and CPU all and Memory all represent the total amount of node CPU and memory; Dynamically calculate the number of ants according to the node load: Step S3312: Initialize the number of iterations. The number of iterations is set between 100 and 2000 times. For task pre-scheduling, analyze and judge the number of iterations based on the total number of tasks and nodes. The calculation of the number of iterations is as follows: Tp = a·n + b·t Among them, T p represents the number of iterations calculated during initialization, n represents the number of nodes, t represents the number of tasks, a and b are adjustment parameters, and T represents the actual number of iterations during initialization; S3313: Initialize the pheromone weight α and the pheromone evaporation coefficient ρ: Among them, T i represents the number of iterations after the i-th iteration, T represents the total number of iterations, and δ and σ represent adjustable parameters; Step S3314: Initialize the heuristic information matrix η ij , and the nodes of the heuristic information matrix are composed of scores S ij as follows: S ij = u·S1 ij + v·S2 ij Among them, S1 ij is the load score of the hard disk, S2 ij is the IO score of the network IO type task, and u + v = 1; Based on the principle of strengthening exploration in the early stage and accelerating convergence in the later stage, initialize the heuristic information weight β, and the initialization formula is as follows:

7. The method according to claim 1, characterized in that: Step 4 includes the following steps: Step S41: Adjust the number of task replicas according to the S1 replica strategy; Step S42: After adjusting the number of replicas, use the node resources predicted by S2 for the task pre-scheduling algorithm in S3 and record the highest score and the scheduling plan with the highest score; Step S43: Repeat the processes of Step S41 and Step S42 until the S1 replica strategy is executed; Step S44: Perform actual horizontal scaling of the tasks according to the number of replicas and the pre-scheduling plan with the highest score, and complete the actual scheduling or deletion of the task replicas based on the node affinity of K3s.