A Jenkins and K8s-based task execution resource scaling method, device and medium
Patent Information
- Application Number
- CN202610369432.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-03-25
AI Technical Summary
这种架构存在明显的局限性:首先,由于所有子任务混排在单一队列中,系统无法根据子任务的不同资源需求、优先级或类型进行差异化调度,导致资源分配不精准,高资源消耗任务可能阻塞低资源需求任务的执行;其次,执行节点通常为静态预配的虚拟机,其数量与规格固定,无法根据任务队列长度或负载变化动态调整,当任务量突增时容易因资源不足造成队列堆积、任务执行延迟,而在任务空闲时又导致资源闲置浪费;此外,基于顺序执行的模式难以充分利用多核与分布式计算能力,系统整体吞吐量与响应效率受限
[0015] This invention presents a task execution resource scaling method based on Jenkins and Kubernetes. By pre-setting independent sub-task queues for different types of pipeline tasks and associating them with dedicated execution instance groups, it achieves refined task scheduling and precise resource allocation, effectively solving the problems of mixed task scheduling and high-load tasks blocking low-load tasks in traditional single-queue architectures. Furthermore, by monitoring queue length and introducing a scaling trigger mechanism based on duration judgment, it drives Kubernetes to dynamically adjust the number of execution instances, achieving a high degree of elasticity and intelligence in resource supply. This overcomes the contradiction of resource idleness and queue accumulation coexisting in the static resource pre-provisioning mode. At the same time, based on the rapid elasticity of containerized instances and the parallel processing capability of multiple queues, it significantly improves the overall throughput and execution efficiency of the system, thereby comprehensively improving the response speed, resource utilization, and operational stability of continuous integration/continuous deployment processes.
Smart Images

Figure CN122152392B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of task publishing technology, and in particular to a method, device and medium for scaling up and down task execution resources based on Jenkins and K8s. Background Technology
[0002] Traditional Jenkins continuous integration / continuous deployment systems typically employ a single-queue task scheduling architecture. Specifically, Jenkins places all subtasks generated during the build process into a centralized task queue, with one or more fixed-configuration virtual machines (or physical machines) acting as execution nodes, sequentially retrieving and executing tasks from this queue. This architecture has significant limitations: First, because all subtasks are mixed in a single queue, the system cannot differentiate scheduling based on the different resource requirements, priorities, or types of subtasks, leading to inaccurate resource allocation, where high-resource-consuming tasks may block the execution of low-resource-requirement tasks. Second, execution nodes are usually statically pre-configured virtual machines with fixed numbers and specifications, unable to dynamically adjust according to task queue length or load changes. When the task volume surges, insufficient resources can easily cause queue congestion and task execution delays, while idle tasks result in wasted resources. Furthermore, the sequential execution model struggles to fully utilize multi-core and distributed computing capabilities, limiting overall system throughput and response efficiency. Therefore, existing solutions still require improvement in terms of task scheduling flexibility, resource utilization elasticity, and overall system execution efficiency. Summary of the Invention
[0003] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:
[0004] According to a first aspect of this application, a method for scaling up and down task execution resources based on Jenkins and Kubernetes is provided, the method comprising the following steps:
[0005] S100, obtain each preset subtask queue; wherein, each subtask queue is used to store subtask configuration information of a pipeline task type, and each subtask queue corresponds to several execution instances for executing that type of subtask.
[0006] S200 responds to any subtask QR corresponding to the Jenkins build-to-deploy task after it has been completed and stores the configuration information corresponding to the QR into the subtask queue corresponding to the QR.
[0007] S300: For any subtask queue RA, if there is an idle execution instance among the several execution instances corresponding to RA, the configuration information in RA is sent to the idle execution instance, so that the idle execution instance executes the subtask corresponding to the received configuration information.
[0008] S400, if the number of configuration information in RA is greater than or equal to the first preset queue length threshold, then obtain the first duration for which the number of configuration information in RA is greater than or equal to the preset queue length threshold.
[0009] S500, if the first duration is greater than or equal to the first preset duration, then the execution instance corresponding to RA is expanded;
[0010] S600, if the number of configuration information in RA is less than or equal to the second preset queue length threshold, then obtain the second duration for which the number of configuration information in RA is less than or equal to the second preset queue length threshold;
[0011] S700, if the second duration is greater than or equal to the second preset duration, then the execution instance corresponding to RA is scaled down.
[0012] According to another aspect of this application, a non-transitory computer-readable storage medium is also provided, wherein at least one instruction or at least one program is stored in the storage medium, and the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described method for scaling up and down task execution resources based on Jenkins and K8s.
[0013] According to another aspect of this application, an electronic device is also provided, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0014] The present invention has at least the following beneficial effects:
[0015] This invention presents a task execution resource scaling method based on Jenkins and Kubernetes. By pre-setting independent sub-task queues for different types of pipeline tasks and associating them with dedicated execution instance groups, it achieves refined task scheduling and precise resource allocation, effectively solving the problems of mixed task scheduling and high-load tasks blocking low-load tasks in traditional single-queue architectures. Furthermore, by monitoring queue length and introducing a scaling trigger mechanism based on duration judgment, it drives Kubernetes to dynamically adjust the number of execution instances, achieving a high degree of elasticity and intelligence in resource supply. This overcomes the contradiction of resource idleness and queue accumulation coexisting in the static resource pre-provisioning mode. At the same time, based on the rapid elasticity of containerized instances and the parallel processing capability of multiple queues, it significantly improves the overall throughput and execution efficiency of the system, thereby comprehensively improving the response speed, resource utilization, and operational stability of continuous integration / continuous deployment processes. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating a method for scaling up and down task execution resources based on Jenkins and Kubernetes, provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] It should be noted that, based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Furthermore, this device and / or practice the method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0020] Example 1:
[0021] The following will refer to Figure 1 The flowchart shown illustrates a method for scaling up and down task execution resources based on Jenkins and Kubernetes, introducing such a method.
[0022] This method for scaling up and down task execution resources based on Jenkins and Kubernetes may include the following steps:
[0023] S100, obtain each preset subtask queue; wherein, each subtask queue is used to store subtask configuration information of a pipeline task type, and each subtask queue corresponds to several execution instances for executing subtasks of that type.
[0024] Furthermore, the pipeline task types include: production release, canary release, test release, and continuous integration build.
[0025] Predefine different "pipeline task types" in the system configuration, such as "Java CI build," "Python test release," "frontend production release," and "Android canary release." Create a logical "sub-task queue" for each type (which can be implemented using message middleware such as RabbitMQ, Kafka, or Redis List), and associate it with a Kubernetes Deployment or StatefulSet resource template. This template defines the image of the execution instance Pod (including the Jenkins Agent and corresponding language / toolchain), resource requests (CPU / memory), etc., required to run this type of task.
[0026] When the system starts, it creates a corresponding Kubernetes workload for each subtask queue according to the configuration and sets the initial number of replicas (e.g., keeping a minimum of 1 instance to maintain connectivity). After these Pods start, they are automatically registered with JenkinsMaster, becoming the queue's dedicated, label-matched build agent.
[0027] This step achieves physical isolation and specialization of resources. By dividing queues and dedicated instance groups according to task type, it fundamentally avoids dependency conflicts between different types of tasks on the execution environment (such as the different environments required for Java Maven builds and Python script execution), and lays the architectural foundation for subsequent fine-grained scaling up and down for specific task load characteristics.
[0028] S200 responds to any subtask QR corresponding to the Jenkins build-to-deployment task, storing the configuration information corresponding to the QR into the subtask queue corresponding to the QR.
[0029] When triggering a task in a Jenkins Pipeline or via the Jenkins API, you must explicitly specify the "pipeline task type" of the task in the task parameters or Pipeline script.
[0030] When Jenkins completes the stage splitting of a release task and generates specific subtasks (such as "QR" representing "compile backend service"), a central scheduler component (which can be a custom Jenkins plugin or a standalone service) reads the type identifier of the subtask. Subsequently, the scheduler serializes the configuration information of the subtask (such as code repository address, branch, build script path, etc.) into a message according to the preset mapping relationship and publishes (Pushes) it to its corresponding message queue.
[0031] This step achieves logical divide-and-conquer and ordering of the task flow. By decoupling mixed task flows by type and distributing them to independent queues, the "head-of-queue blocking" problem of a single queue is avoided. For example, a time-consuming Android packaging task will not block subsequent fast front-end deployment tasks, improving the overall parallelism and response speed of the task flow.
[0032] S300: For any subtask queue RA, if there is an idle execution instance among the several execution instances corresponding to RA, the configuration information in RA is sent to the idle execution instance, so that the idle execution instance executes the subtask corresponding to the received configuration information.
[0033] Each execution instance Pod runs a lightweight queue consumer service. This service continuously listens to its specific message queue.
[0034] When the consumer service detects a new message (subtask configuration information) in the queue, it first checks whether it is currently executing a task (i.e., whether it is in an "idle" state). If it is idle, it retrieves (Pops) the message from the queue and deserializes it to obtain the task configuration.
[0035] Based on the obtained configuration information, the consumer service registers with the Jenkins Master through the Jenkins Agent's communication channels (such as JNLP or WebSocket) and reports that it is ready. It then waits for the Jenkins Master to issue specific build instructions according to the global scheduling policy, thereby starting and executing the subtask. During task execution, the instance is marked as "busy" in the Jenkins Master's node list and no longer receives new task assignments; after execution, it is marked back as "idle" and continues to listen to the queue, waiting for the next scheduling.
[0036] This step achieves seamless integration with Jenkins' native scheduling mechanism: Jenkins Master acts as the central scheduler, uniformly allocating tasks based on the real-time status and tags of each execution instance. This ensures both the consistency and reliability of the scheduling strategy, while also enabling on-demand allocation and efficient utilization of resources. Each instance focuses on its strengths in specific task types through a tagging mechanism, reducing environment switching overhead and thus improving the execution efficiency of individual tasks.
[0037] S400, if the number of configuration information in RA is greater than or equal to the first preset queue length threshold, then obtain the first duration for which the number of configuration information in RA is greater than or equal to the preset queue length threshold.
[0038] Deploy a monitoring component (such as Prometheus) to periodically capture the current length (queue_length) of each message queue through exposed interfaces. If queue_length is greater than or equal to a first preset queue length threshold, then start timing (recorded as T_overload_start). Only when the above conditions are continuously and uninterruptedly met, until the cumulative duration reaches the "first preset duration", will a specific "expansion event" or "shrinkage event" signal be generated.
[0039] In this embodiment, the first preset duration can be set to a fixed value, for example, the first preset duration is 30 seconds.
[0040] Furthermore, the first preset queue length threshold can be determined through the following steps:
[0041] S410: Obtain all subtasks completed by RA within the past statistical period T, and calculate the corresponding average execution time t_task.
[0042] Within a statistical period T (e.g., the past 24 hours or week), monitoring systems (such as Prometheus) need to record the actual execution time of each completed subtask in the subtask queue RA. This is typically achieved by timestamping the start and end of the task and reporting metrics via the Jenkins Agent or task executor.
[0043] After period T ends, the execution times of all collected tasks are summed, and then divided by the total number of tasks to obtain the average execution time t_task. The formula is: t_task = Σ(execution time of each task) / total number of tasks.
[0044] This step establishes the threshold based on a realistic and measurable baseline of workload performance. `t_task` reflects the "standard processing speed" of this type of task in the current environment, avoiding misjudgments caused by setting thresholds based on theories or outdated data, and making scaling decisions more aligned with actual processing capabilities.
[0045] S420 obtains the average cold start time t_startup required to start a corresponding execution instance for RA from the K8s cluster monitoring data.
[0046] Record the time elapsed from when the system issues the command to create the execution instance Pod corresponding to the RA queue (such as kubectl create or Deployment expand replica) to when the Pod's state becomes Ready (i.e., the internal Jenkins Agent has completed startup and successfully registered with the Master) through K8s event monitoring, Pod state change logs, or dedicated monitoring agents (such as kube-state-metrics).
[0047] Within period T, the startup time of all newly started Pods of this type is sampled and the average value is calculated to obtain t_startup.
[0048] This step quantifies the delayed costs of resource provisioning. By incorporating the infrastructure response time (t_startup) into the decision-making model, it is recognized that expansion is not instantaneous, thus requiring the threshold to reserve sufficient task buffers for this "resource delivery window".
[0049] S430, based on t_task and t_startup, determine the base threshold N_base=ceil(t_startup / t_task); where ceil() is the rounding up function.
[0050] Assuming t_task = 2 minutes and t_startup = 5 minutes, then N_base = ceil(5 / 2) = ceil(2.5) = 3. The physical meaning is: starting a new instance takes 5 minutes, while an existing instance can process one task every 2 minutes. Therefore, a "buffer" of at least 3 tasks is needed to ensure that the backlog of tasks does not exhaust the processing capacity of the existing instance before the new instance is ready, preventing the queue from growing uncontrollably.
[0051] This step establishes a preventative buffering model based on resource delivery delays. N_base defines the minimum safe queue length required to prevent the queue from growing indefinitely during resource expansion. It ensures that the expansion trigger point always occurs before resources are truly exhausted, achieving proactive resource scheduling.
[0052] S440, get the average number of tasks N_blocked in RA due to waiting for the execution instance to start within T.
[0053] Within period T, a "blockage" occurs whenever a task in the queue is in a "ready" state but cannot be consumed immediately because all execution instances are busy. More specifically, the growth of the queue length during instance startup can be monitored.
[0054] Calculate the peak increment of queue length or the cumulative number of waiting tasks caused by each instance startup (usually accompanied by a scaling operation), and then average these values over period T to obtain N_block.
[0055] This step captures the "pain index" of historical scaling behavior. N_block reflects the actual task latency experienced by the system when dealing with load growth under the existing strategy. It is a key feedback signal used to determine whether the current buffer (N_base) is sufficient.
[0056] S450, based on N_base and N_block, determines the cold start compensation factor α = 1 + (N_block / N_base).
[0057] The ratio N_block / N_base measures the severity of historical blocking. If the ratio is 0, N_base is perfectly adequate, α=1, and the threshold requires no compensation. If the ratio is 0.5, there has been significant historical blocking, α=1.5, and the threshold will be increased by 50% to provide more buffering.
[0058] This step enables adaptive adjustments based on historical performance feedback. It transforms the system's "lessons learned" (N_block) into adjustments (α) for future strategies. If historical data shows frequent task waiting during expansion, the trigger threshold is automatically increased, making future expansion decisions earlier and more proactive, thereby continuously optimizing system responsiveness.
[0059] S460, based on α and N_base, determine the first preset queue length threshold L_threshold = α × N_base.
[0060] This step generates a dynamic, intelligent, and learning-capable expansion trigger standard. The final threshold L_threshold is no longer a statically configured fixed number, but a composite metric that integrates task processing capacity (t_task), infrastructure latency (t_startup), and historical service quality feedback (N_block). It can automatically evolve with system behavior and environmental changes, achieving a dynamic balance between "avoiding premature expansion that wastes resources" and "preventing late expansion that leads to queue backlog," significantly improving the accuracy of resource elasticity and the overall stability of the system.
[0061] S500, if the first duration is greater than or equal to the first preset duration, then the execution instance corresponding to RA is expanded.
[0062] Furthermore, the execution instance in S100 can be an all-in-one execution instance. In this step, if the first duration is greater than or equal to the first preset duration, step Q500 in Embodiment 2 can be executed first.
[0063] Furthermore, the first preset duration is determined through the following steps:
[0064] S510, based on the historical sequence of queue length of RA in the most recent M statistical windows, determine the average rate of change V_avg and instantaneous volatility σ of the queue length corresponding to RA.
[0065] Extract the time series of the queue length of the subtask queue RA from a monitoring system (such as Prometheus) over the most recent M consecutive statistical windows (e.g., M=10 windows, each window is 5 minutes). Assuming that each window records an average length value, the resulting sequence is L1, L2, ..., LM.
[0066] Calculate the average rate of change V_avg: Calculate the slope of the linear regression for the sequence, or simply calculate the average of the length differences between adjacent windows. V_avg = (LM - L1) / (M × window duration). A positive value indicates the queue is growing, and a negative value indicates it is decreasing.
[0067] Calculate instantaneous volatility σ: Calculate the deviation of each data point in the sequence from the sequence mean (or a moving average), and then calculate the standard deviation of these deviations. σ reflects the degree of fluctuation in the queue length around its trend line.
[0068] This step establishes a quantitative understanding of the queue's dynamic behavior. V_avg reveals the macro-level trend of the load (whether it's stable, growing slowly, or growing rapidly), while σ reveals the micro-level instability of the load. These two metrics provide crucial data input for subsequent duration decisions, enabling the system to "sense" the load's stress patterns and noise levels.
[0069] S511, based on V_avg, determine the initial reference duration T_base=β / (|V_avg|+ε); where β is the preset system response coefficient and ε is the preset constant.
[0070] β is a preset "system response coefficient" that can be adjusted according to the service's tolerance for latency (e.g., β = 300 seconds, representing the system's baseline response scale). ε is a very small positive number (e.g., 0.001) used to prevent the denominator from being zero when V_avg is close to zero.
[0071] T_base is inversely proportional to the rate of change of the queue. If the queue grows rapidly (|V_avg| is large), such as adding 1 task per second, then T_base will become very small (e.g., 300 / 1 = 300 seconds, or 5 minutes, is already quite long; in reality, β needs to be adjusted), meaning the system needs to make expansion decisions faster and cannot wait too long. If the queue remains almost unchanged (|V_avg| is close to 0), T_base will become very large, and the system can be very "patient" because a brief exceedance of the threshold is likely noise.
[0072] This step is the core of "trend-driven decision-making," ensuring that the system reacts quickly to clear and strong load growth signals and remains cautious when faced with ambiguous or weak signals, avoiding the problems of sluggish response or oversensitivity caused by a fixed-duration "one-size-fits-all" approach.
[0073] S512, based on σ, determine the fluctuation suppression factor γ = 1 + (σ / σ_ref); where σ_ref is the reference benchmark value for queue length fluctuation.
[0074] σ_ref is a reference value for queue volatility set based on historical data or experience, representing a "normal" level of volatility. γ increases linearly with volatility σ. When queue jitter is severe (large σ), γ > 1, for example, when σ = 2 × σ_ref, γ = 3. This means that a high-volatility environment will "suppress" or "prolong" the final decision-making time. Because high-frequency noise can easily cause momentary over-thresholds, the γ factor, by increasing the duration, requires the over-threshold state to last longer before triggering an action, thereby filtering out brief spikes.
[0075] This step effectively prevents "jitter expansion" (frequent expansion and contraction) caused by random fluctuations in queue length, significantly improving system stability and decision reliability, and reducing unnecessary resource operation overhead.
[0076] S513, combined with the ratio of remaining schedulable resources available for RA in the current K8s cluster R_available, determine the resource pressure coefficient η=max(0.5,R_available); max() is the preset maximum value function.
[0077] Query the scheduler status of the Kubernetes cluster or use monitoring data from interfaces such as `kubectl describe node` to calculate the total remaining allocable CPU and memory (i.e., "schedulable resources") on all nodes. Divide this by the total schedulable resources of the cluster to obtain the global `R_available`. A more refined approach is to only calculate the resources on nodes that can run RA type instances (considering node selectors, taints, etc.).
[0078] η reflects the resource availability of the cluster and has a safety lower bound. When cluster resources are plentiful (high R_available, such as 0.8), η is also high (0.8). When resources are scarce (low R_available, such as 0.2), η is clamped at 0.5. In subsequent formulas, η is in the denominator.
[0079] This step achieves resource-aware risk control. When resources are scarce (η takes a smaller value of 0.5), the final T_threshold is increased (because the denominator is small), making the system more conservative and cautious in deciding whether to consume the remaining resources for expansion. This requires a longer confirmation time, preventing reckless expansion decisions during resource bottleneck periods, which could exacerbate cluster pressure or even lead to scheduling failures. This embodies the intelligence of "acting within one's means."
[0080] S514, determine the first preset duration T_threshold=(T_base×γ) / η based on T_base, γ and η.
[0081] Ultimately, the duration is driven by three forces: Trend-driven force (T_base): The stronger the trend, the shorter the duration (faster response is required). Volatility-suppressing force (γ): The greater the noise, the longer the duration (greater stability is required). Resource pressure (η): The tighter the resources, the longer the duration (careful decision-making is required). The formula (T_base×γ) / η perfectly integrates these three sometimes contradictory demands.
[0082] T_threshold is no longer a static configuration, but an intelligent parameter calculated in real time that balances "response speed," "decision stability," and "resource security." It gives the expansion triggering mechanism the experience-based judgment of a seasoned driver.
[0083] S515, constrain T_threshold so that T_threshold is within the feasible operating range [T_min, T_max]; T_min is the first preset minimum duration, and T_max is the first preset maximum duration.
[0084] Preset an absolute minimum value T_min (e.g., 30 seconds) and a maximum value T_max (e.g., 30 minutes). Regardless of the calculated values, the final effective duration will be limited to the range [T_min, T_max]. T_min prevents the system from reacting too quickly (avoiding responses to extremely short-term disturbances), and T_max prevents the system from reacting too slowly (ensuring a response limit even in the worst-case scenario).
[0085] The above steps, through a layered, multi-dimensional dynamic calculation model, transform the fixed "first preset duration" into an intelligent system state function. First, S510 quantifies queue behavior (trends and fluctuations), establishing a perceptual foundation for decision-making. Then, a trend-driven formula establishes the core principle of rapid response to clear growth. Next, a fluctuation suppression factor cleverly filters noise interference, enhancing decision stability. Furthermore, an innovative resource pressure coefficient is introduced, giving scaling decisions a "resource cost awareness," automatically switching to a conservative mode when cluster resources are strained. Finally, a comprehensive calculation and application of safety boundaries are performed. This entire mechanism works collaboratively, enabling the system to achieve a delicate, adaptive balance between "acting quickly in the face of real load growth" and "avoiding erroneous actions due to noise or resource constraints." This not only greatly improves the accuracy and reliability of scaling decisions and effectively eliminates resource jitter, but also optimizes overall resource utilization efficiency and system operating costs, making it a core intelligent controller for building highly elastic and stable CI / CD infrastructure.
[0086] Furthermore, the execution instances corresponding to RA are expanded, including the following steps:
[0087] S520, based on the pipeline task type LA corresponding to RA, obtains the resource weight coefficient W and the number of parallel tasks per instance P of LA from the preset configuration table.
[0088] The system maintains a task type configuration table (which can be stored in a ConfigMap, database, or configuration center), predefining two key parameters for each pipeline task type (such as "Java CI Build" and "Front-end Production Deployment").
[0089] Resource Weight Coefficient (W): A dimensionless value representing the relative amount of resources required to start a Pod of this type. Calculation method: Using the resource request (CPU + memory) of a baseline task type (such as a simple script task) as the benchmark (W=1), the W value for other types = the resource request of that type / the baseline resource request. For example, a Java build Pod requiring 2 cores and 4GB of memory might have a W value of 2; while a Python test Pod requiring only 0.5 cores and 1GB of memory might have a W value of 0.5.
[0090] Single-instance parallel task count (P): An integer representing the maximum number of this type of subtask that a single execution instance Pod can process simultaneously. This value is typically determined by the numExecutors configuration of the Jenkins Agent within the Pod. For example, an Agent configured with 2 executors has P=2.
[0091] Parameter query: When it is necessary to expand the queue RA, query this configuration table according to the task type LA associated with RA to obtain the corresponding W and P values.
[0092] This step enables templated and differentiated management of task resource configuration. Through predefined W and P, the system can accurately recognize the resource requirements (cost) and processing capabilities (efficiency) of different types of tasks, providing standardized input for subsequent precise capacity calculations and avoiding a "one-size-fits-all" resource allocation strategy.
[0093] S521, based on the length of RA L_current and the first preset queue length threshold L_threshold, determine the number of tasks to be resolved based on queue pressure ΔQ=L_current-L_threshold.
[0094] Query the API or monitoring metrics of message queue services (such as RabbitMQ and Kafka) in real time to obtain the current length L_current of the subtask queue RA (i.e., the number of messages waiting in the queue). Retrieve the first preset queue length threshold L_threshold calculated for queue RA from the dynamic threshold calculation module or configuration storage. ΔQ represents the number of backlogged tasks that exceed the safety buffer and urgently need to be processed.
[0095] This step concretizes and quantifies the expansion requirements, transforming the qualitative judgment of "the queue is too long" into a quantitative target of "ΔQ additional tasks need to be processed," providing clear input for subsequent resource calculations.
[0096] S522, based on ΔQ and P, determine the theoretical number of new instances N_need required to satisfy the current queue pressure: N_need = ceil(ΔQ / P); ceil() is the floor function.
[0097] Assuming ΔQ=7 (there are 7 backlog tasks) and P=2 (each instance can handle 2 tasks at the same time), then N_need=ceil(7 / 2)=ceil(3.5)=4. This means that to process these 7 backlog tasks at the same time, ideally 4 new execution instances are needed (3 instances each handle 2 tasks, and 1 instance handles 1 task).
[0098] By considering the parallel processing capability P of a single instance, the calculated N_need is more in line with the actual concurrent execution scenario, avoiding the underestimation of the number of instances caused by the sequential processing thinking (P=1), thus designing the expansion scale more efficiently to quickly resolve the queue.
[0099] S523, query the total amount of idle resources currently available for LA in the K8s cluster, and calculate the maximum number of instantaneous expansion instances N_max allowed under the current resource conditions, in combination with the resource weight coefficient W.
[0100] Call the Kubernetes API or use a cluster monitoring system (such as Prometheus combined with kube-state-metrics) to obtain the allocatable and allocated resources (Requests) of all worker nodes. Calculate the total amount of remaining schedulable resources in the cluster, typically by calculating CPU (cores) and memory (GB).
[0101] Divide the remaining CPU and memory resources by the Pod resource request specification for task type LA to obtain the maximum number of instances that can be created based on CPU and memory constraints, respectively. Take the smaller of the two values as the theoretical maximum value under the resource constraints.
[0102] Since the W coefficient has been normalized, the above calculation is equivalent to: N_max based on a certain resource = (total remaining amount of that resource in the cluster / baseline resource request) / W. Ultimately, N_max takes the minimum value calculated using CPU and memory.
[0103] Further consideration can be given to scheduling rules such as node selector, taints and tolerations, and Pod anti-affinity, and N_max can be further modified through kube-scheduler-simulator or empirical rules.
[0104] This step ensures safe boundary control for capacity expansion operations. It is crucial for preventing uncontrolled expansion and guaranteeing overall cluster stability. It ensures that expansion operations are conducted within the hard constraints of the cluster's real-time resource capabilities, preventing uncontrolled expansion of a single queue from leading to node resource exhaustion, system scheduling failures, or impacts on other critical services. This reflects a holistic resource management approach that prioritizes prudent resource allocation.
[0105] S524, based on N_need and N_max, determine the number of expanded instances N_scale=min(N_need,N_max), and call the K8s cluster interface to create N_scale new execution instances for RA; min() is the preset minimum value function.
[0106] Call the Kubernetes API (e.g., via the Kubernetes Go client or the kubectl command-line tool) to update the replica count of the Deployment or StatefulSet corresponding to the queue RA. The update operation is: New replica count = Current replica count + N_scale.
[0107] When the Kubernetes control plane receives a change in the number of replicas, it schedules and creates N_scale new Pods on the corresponding nodes. After these Pods start, their internal Jenkins Agents automatically connect to the JenkinsMaster based on preset tags and begin listening to the corresponding queues, becoming new execution instances.
[0108] This step completes the closed loop from intelligent decision-making to secure implementation. The `min()` function ensures that the scaling operation is both proactive (striving to meet demand) and robust (never exceeding resource limits). Automated API calls enable rapid, unattended elastic scaling, instantly translating calculated strategies into actual resource supply, significantly shortening the cycle from load perception to resource provision, and effectively improving the system's agility and automation level.
[0109] The aforementioned scaling method constructs a three-layer funnel-shaped decision-making and execution model. First, it establishes quantitative standards for resources and capabilities through task type parameterization. Second, it transforms the vague "queue length" into the precise "N instances needed" through a demand quantification funnel. Then, it imposes constraints on scaling requirements from a cluster-wide perspective through a resource constraint funnel, calculating a safe upper limit. Finally, it completes the transformation from decision to resources through a safe execution layer. This entire process achieves precise planning (based on W and P), clear objectives (based on ΔQ), clear boundaries (based on N_max), and automated execution. This ensures that every scaling operation is justified, measured, and effective, maximizing resource utilization to cope with peak loads while steadfastly safeguarding the overall stability and health of the cluster—a core operational guarantee for the long-term reliable operation of the elastic system.
[0110] S600, if the number of configuration information in RA is less than or equal to the second preset queue length threshold, then obtain the second duration for which the number of configuration information in RA is less than or equal to the second preset queue length threshold.
[0111] Deploy a monitoring component (such as Prometheus) to periodically capture the current length (queue_length) of each message queue through the exposed interface. If queue_length ≤ the second preset queue length threshold, start timing (recorded as T_underload_start).
[0112] Furthermore, the second preset queue length threshold is determined through the following steps:
[0113] S610, obtain the average resource utilization rate of the execution instance cluster corresponding to RA in the current statistical period.
[0114] Use a monitoring system (such as Prometheus Node Exporter, cAdvisor, and kube-state-metrics) to collect the CPU and memory usage of all execution instance Pods corresponding to the RA queue within the target period. Simultaneously, collect the resource requests declared by these Pods as the denominator.
[0115] For each Pod, calculate its CPU utilization (actual usage / requests) and memory utilization (actual usage / requests). To avoid the impact of instantaneous peaks, the average value over a period of time (such as the most recent 15 minutes) is usually taken.
[0116] Calculate the arithmetic mean of the average CPU utilization and average memory utilization of all Pods belonging to the RA queue to obtain a value representing the overall CPU and memory load. To obtain a comprehensive metric, a weighted average can be applied (e.g., CPU weight 0.7, memory weight 0.3) to obtain the final average resource utilization (U_avg). The value of U_avg ranges from 0 to 1.
[0117] U_avg directly reflects the degree to which the computing resources currently allocated to this queue are being "effectively used." This is the most direct and core indicator for judging whether resources are idle or need to be reclaimed, providing the primary data-driven basis for scaling-down decisions.
[0118] S611, Determine the resource release conservatism coefficient based on the average resource utilization rate; wherein, the lower the average resource utilization rate, the larger the resource release conservatism coefficient.
[0119] Define a function that maps U_avg to a resource release conservatism coefficient (λ). The core mapping rule is: the lower U_avg (the more idle the resource), the larger the value of λ should be to promote more aggressive scaling down. The function can be: λ = 1 + (1 - U_avg). For example, when U_avg = 0.2 (utilization 20%), λ = 1.8; when U_avg = 0.8 (utilization 80%), λ = 1.2.
[0120] This coefficient, acting as an amplification factor in subsequent calculations, dynamically adjusts the system's "thrifty" tendency. The more idle the resources, the larger this coefficient becomes, resulting in a higher calculated scaling-down threshold (making it easier to trigger scaling-down), thus releasing idle resources back to the cluster resource pool more quickly and improving overall resource turnover.
[0121] S612, obtain the typical minimum queue length of RA during the historical steady state period.
[0122] Historical steady-state periods refer to a timeframe (e.g., the past week) during which the RA queue has not undergone any expansion operations and its length has remained relatively stable. These periods can be identified by analyzing expansion / shrinkage event logs and queue length monitoring data.
[0123] Extract the length sequence of RA queues during these steady-state periods from the monitoring database. Perform statistical analysis on this sequence (e.g., take the 5th percentile or the minimum value after removing outliers) to obtain the typical minimum queue length (L_min_typical). This value represents the number of "background" tasks (e.g., resident low-priority tasks or tasks about to be scheduled) that may still exist in the queue when the system is calm and resource supply and demand are balanced.
[0124] L_min_typical reflects the system's "natural water level" under no external pressure. Using it as a benchmark for calculation can prevent the shrinkage threshold from being set too low, avoiding accidental triggering of shrinkage when the queue length returns to a normal calm state (but is slightly above zero), thus ensuring the system's basic stability during low-load periods and retaining the ability to cope with minor fluctuations.
[0125] S613, determine the basic shrinkage threshold based on the typical minimum queue length and the resource release conservative coefficient.
[0126] Multiply the typical minimum queue length (L_min_typical) by the resource release conservatism coefficient (λ) to obtain the base shrinkage threshold (L_base). That is: L_base = L_min_typical × λ.
[0127] L_base combines the system's historical calm water level (L_min_typical) and the current resource idleness level (λ). When resources are idle (λ is large), L_base will be significantly higher than the historical minimum length, meaning the system can "tolerate" longer queues before triggering shrinkage, exhibiting more aggressive behavior. When resources are fully utilized (λ is small), L_base approaches the historical minimum length, and the behavior becomes more conservative.
[0128] This step generates a preliminary decision line that integrates historical patterns and real-time status. L_base is a dynamic, context-aware threshold, no longer a static value. This allows scaling down decisions to both respect the system's inherent load patterns and respond sensitively to changes in resource utilization, achieving an initial balance between the goals of "maintaining stability" and "reclaiming resources."
[0129] S614. Determine the cost efficiency factor based on the average task digestion rate corresponding to RA and the keep-alive maintenance cost index of the execution instance.
[0130] Calculate the average task completion rate (V): In the most recent statistical period, divide the total number of tasks successfully completed by the RA queue by the total duration of that period to obtain the average number of tasks completed per second (or per minute). V = Total number of tasks / Duration of period.
[0131] Quantitative Keep-Alive Cost Metric (C): This is a comprehensive metric designed to quantify the "cost" required to keep an execution instance running. It can use the instance's resource consumption cost, for example: C = (Pod CPU requests × CPU unit price) + (Pod memory requests × memory unit price). More complex models can incorporate the instance's own management overhead, the impact of node fragmentation, etc.
[0132] The cost efficiency factor (μ) is calculated as μ = V / C, which represents the efficiency per unit cost. A higher μ value means that each execution instance in the queue is more cost-effective, processing tasks faster while consuming relatively fewer resources. To facilitate subsequent calculations, μ is usually normalized, for example, by dividing it by a reference value (μ_ref), resulting in μ_normalized = μ / μ_ref.
[0133] This step introduces an economic benefit dimension, enabling intelligent cost optimization. It links technical decisions to business costs. Through cost efficiency factors, the system can identify and "protect" instance groups that are highly efficient (V large) or have low resource costs (C small). Inefficient queues (μ small) will have their scaling-down threshold lowered (in the next step), making them easier to scale down, thus freeing up resources for more efficient task types and optimizing the cluster's "return on investment" from a global perspective.
[0134] S615, combine the basic shrinkage threshold with the cost efficiency factor to obtain a preliminary queue length threshold.
[0135] Combine the base shrinkage threshold (L_base) with the cost efficiency factor (μ_normalized). A straightforward approach is to multiply them: L_preliminary = L_base × μ_normalized.
[0136] If μ_normalized > 1 (efficiency higher than the baseline), then L_preliminary > L_base, increasing the shrinkage threshold. This means that the system will be more "lenient" towards this high-efficiency queue, allowing it to maintain a relatively long queue (i.e., retain more instances) to fully utilize its efficient processing capabilities.
[0137] If μ_normalized < 1 (efficiency is lower than the baseline), then L_preliminary < L_base, lowering the shrinkage threshold. This means that for this inefficient queue, the system will be more "strict," considering shrinkage when the queue is shorter to reduce its resource consumption.
[0138] This step is crucial for global resource optimization. It ensures that when deciding "whose capacity to reduce," we consider not only whose resources are idle, but also whose "efficiency" in using those resources is high. This guides the system to allocate limited resources to high-efficiency tasks, indirectly improving the throughput and cost-effectiveness of the entire cluster, and achieving refined, value-driven operations and maintenance.
[0139] S616, the maximum value between the initial queue length threshold and the preset absolute safety lower limit is determined as the second preset queue length threshold.
[0140] Set an absolute safety lower limit (L_safe): This is a static value preset based on system reliability and business continuity requirements. For example, it can be set to 1 to ensure that scaling down is only triggered when at least one task is in the queue at any time; or it can be set to the number of instances × 1 to ensure that each reserved instance has at least one pending task buffer.
[0141] The above steps construct a multi-layered, multi-objective optimized dynamic scaling-down decision engine. It begins with direct perception of resource utilization, driving the system to actively reclaim resources when they are idle; then it anchors to historical steady-state baselines, ensuring decisions do not deviate from the system's inherent patterns; next, it introduces cost-efficiency analysis, linking resource allocation to economic benefits, intelligently differentiating the "value density" of resource use, and prioritizing high-efficiency loads; finally, it uses a safety lower limit as a safety net to firmly safeguard stability. This entire mechanism upgrades the "scaling down" judgment from a simple comparison based on a single length threshold to a complex optimization process that comprehensively considers real-time load, historical patterns, processing efficiency, and economic costs. This not only greatly improves the rationality, fairness, and global optimality of scaling-down decisions, achieving the core goal of cost reduction and efficiency improvement, but also ensures operational security and system resilience through multiple safeguard mechanisms. It is a key intelligent component for building efficient, economical, stable, and reliable cloud-native elastic systems.
[0142] S700, if the second duration is greater than or equal to the second preset duration, then the execution instance corresponding to RA is scaled down.
[0143] Furthermore, the second preset duration ranges from 3 minutes to 6 minutes, for example, the second preset duration is 4 minutes.
[0144] Furthermore, the execution instance corresponding to RA is scaled down, including the following steps:
[0145] S710, for several execution instances corresponding to RA, identify at least one target execution instance that is in an idle state and has been idle for more than a third preset duration.
[0146] Each execution instance Pod runs a health reporting service that periodically (e.g., every 30 seconds) sends a heartbeat signal to the central controller (or Jenkins Master), while also reporting its current status (busy / idle) and the currently executing task ID (empty if idle).
[0147] The central controller records the timestamp of the last time each instance transitioned from a busy state to an idle state. The difference between the current time and this timestamp represents the duration of the instance's continuous idle time.
[0148] Query all execution instances belonging to queue RA, and filter out a list of instances whose status is idle and whose continuous idle time is greater than the third preset time. This list is the candidate set of "target execution instances".
[0149] To further optimize the selection, the candidate set can be sorted from longest to shortest idle time, or from oldest to newest instance creation time (assuming older instances may have run longer and have a higher degree of resource fragmentation).
[0150] S720, determine the number of instances to be released from the target instances.
[0151] Based on the current load of queue RA (queue length, task arrival rate) and the processing capacity of the execution instances, calculate the theoretical number of instances that should be reduced.
[0152] Ensure that the remaining number of instances after scaling down still maintains a certain processing buffer. The formula can be: N_reduce=max(0, total number of current instances-ceil(L_current / P)-N_keep), where L_current is the current queue length, P is the number of parallel tasks per instance, and N_keep is the number of additional buffer instances to be retained (e.g., 1-2).
[0153] The calculated theoretical reduction capacity N_reduce is compared with the number of target instance candidates N_idle_candidate identified in step S710. The actual number of instances to be released, N_scale_down, should be the smaller of the two: N_scale_down = min(N_reduce, N_idle_candidate).
[0154] Regardless, after scaling down, it should be ensured that the queue RA retains at least minReplicas (the preset minimum number of replicas, such as 1) execution instances. Therefore, N_scale_down must ultimately satisfy: N_scale_down ≤ total number of current instances - minReplicas.
[0155] This step balances two dimensions: "how much should be scaled down" (based on the load model) and "how much can be scaled down" (based on actual idle instances). It avoids excessive scaling down (exceeding the actual number of idle instances) due to an aggressive model, while also preventing excessive scaling down at once (exceeding load requirements) due to too many idle instances, thus achieving safe and gradual resource release. At the same time, the minimum replica count guarantee ensures basic service availability.
[0156] S730, call the K8s cluster interface to delete the execution instances corresponding to the number of instances to be released, and update the mapping relationship between RA and execution instances.
[0157] Calling Kubernetes APIs (such as kubectl delete pod) <pod-name>Alternatively, you can delete the selected target Pod instance via a Kubernetes client library. The key point is that Kubernetes sends a SIGTERM signal to the Pod, triggering graceful termination.
[0158] During the graceful termination period (30 seconds by default), the queue consumer service within the Pod should: stop pulling new tasks from the message queue; continue completing currently executing tasks (if any). If the task execution time might exceed the graceful termination period, the consumer service should, upon receiving a termination signal, save the task progress (e.g., update the task status to "in interruption") and republish the task configuration information back to the message queue (or notify the task to be rescheduled through other mechanisms).
[0159] Remove the deleted instance information from the service registry of the central controller or the node list of the Jenkins Master. Update the association record between the queue RA and the execution instance to ensure that subsequent tasks are not scheduled to instances that no longer exist.
[0160] Monitor the results of Pod deletion operations to confirm successful deletion. If a Pod deletion fails (e.g., stuck in the Terminating state), an alert should be triggered and logged for manual intervention.
[0161] Through the graceful termination mechanism of Kubernetes and the cooperation of consumer services, the scaling-down operation ensures that no tasks are lost or abnormally interrupted. The mapping relationships after deletion are updated in real time, maintaining the consistency of the system view and preventing invalid scheduling. The entire operation is automated and monitorable, releasing idle resources while maximizing business continuity and system reliability.
[0162] Example 2:
[0163] In the first embodiment above, scaling up and down operations may be performed frequently, which will generate additional computing power consumption. To avoid this problem, the execution instance corresponding to each subtask queue can be set as an all-purpose execution instance, specifically including the following steps:
[0164] Q100 retrieves each preset subtask queue; each subtask queue stores subtask configuration information for a pipeline task type, and each subtask queue corresponds to several universal execution instances, which can execute subtasks of each pipeline task type.
[0165] Furthermore, the pipeline task types include: production release, canary release, test release, and continuous integration build.
[0166] The execution instance in S100 can be a full-featured execution instance.
[0167] In the system configuration, separate logical queues are created for different pipeline task types (such as "Java build", "front-end deployment", "Android packaging" etc.), which can be implemented using message middleware (such as RabbitMQ, Kafka) or database tables.
[0168] Deploy one or more "universal execution instance" Pods in a Kubernetes cluster. The container images of these Pods contain the common build environment required for all task types (such as JDK, Node.js, Python, Docker, etc.) and install a common JenkinsAgent. They are tagged with a common Jenkins label (such as label=universal-agent).
[0169] The all-in-one instances are deployed as Kubernetes Deployments, with a reasonable initial number of replicas (e.g., 2-3) to form a shared resource pool. These instances automatically register with the Jenkins Master after startup, awaiting task assignment.
[0170] This step establishes a shared resource pool across task types, breaking down the barriers of static resource allocation by task type in traditional architectures. This lays the foundation for subsequent intelligent scheduling and resource reuse, and improves the utilization rate of basic resources.
[0171] Q200 responds to any subtask QR corresponding to the Jenkins build-to-deployment task, storing the configuration information corresponding to QR into the subtask queue corresponding to QR.
[0172] In the Jenkins Pipeline definition, explicitly specify the "Pipeline Task Type" label for each subtask. This information can be stored in the Pipeline script or added via a Jenkins plugin.
[0173] When Jenkins completes the stage split of a release task and generates a specific subtask (QR), a central scheduling service (or an enhanced Jenkins plugin) reads the task type tag of the QR, then serializes the complete configuration information of the subtask (including code repository, build parameters, environment variables, etc.) into JSON format and publishes it to the corresponding subtask queue.
[0174] This step enables intelligent classification and orderly management of task flows, ensuring that different types of tasks enter dedicated processing channels and providing clear input signals for subsequent differentiated and flexible strategies.
[0175] Q300: If there are idle execution instances among the execution instances corresponding to RA, the configuration information in RA will be sent to the idle execution instances, so that the idle execution instances can execute the subtasks corresponding to the received configuration information.
[0176] The specific implementation of this step is as follows:
[0177] Each execution instance Pod runs a generic task consumer service. This service continuously listens to one or more message queues associated with it (for a general instance, it listens to all queues; for a dedicated instance, it only listens to queues of its corresponding type). When a new message is detected, the consumer service retrieves the subtask configuration information from the queue.
[0178] The consumer service parses the task configuration and identifies the pipeline type labels required for the task. Subsequently, the consumer service establishes a connection with the Jenkins Master through the Jenkins Agent's communication channels (such as JNLP or WebSocket), dynamically adjusts its own Agent labels to match the labels required by the task to be executed, and reports to the Master that it is in an idle state.
[0179] The Jenkins Master acts as a central scheduler, maintaining the real-time status (tags, load, idle / busy) of all registered agents. When a task needs to be executed, the Master proactively selects a matching agent based on the required tags and the agent's idle status, and issues specific build instructions to it via the Agent protocol.
[0180] After receiving the instruction, the selected Agent starts and executes the subtask. During task execution, the instance is marked as "busy" in the JenkinsMaster's node list, and the consumer service pauses pulling new messages from the message queue. After execution, the Agent reports "idle" to the Master, and the consumer service resumes listening to the message queue, preparing to process the next task.
[0181] This step achieves seamless integration with Jenkins' native scheduling mechanism. The solution employs a combination of event-driven and Master-driven proactive scheduling: the consumer service retrieves task configurations from the message queue and dynamically adjusts Agent labels to prepare for Master scheduling; the Jenkins Master acts as a unified central scheduler, proactively allocating tasks to the most suitable idle Agent based on the global state. This design leverages the advantages of Jenkins' mature scheduler, avoiding the duplication of complex scheduling logic on the Kubernetes side, while ensuring the consistency and reliability of the scheduling strategy. Simultaneously, by transforming Kubernetes-level queue load information into label signals understandable by Jenkins through the consumer service, the Master is guided to accurately allocate tasks, achieving elastic resource orientation based on queue load. This significantly improves overall resource utilization and task execution efficiency while ensuring scheduling reliability.
[0182] Q400: If the number of configuration information in RA is greater than or equal to the first preset queue length threshold, then check whether there are idle execution instances in the execution instances corresponding to other subtask queues besides RA.
[0183] Furthermore, the first preset queue length threshold is determined through the following steps:
[0184] Q410: Obtain the peak queue length of RA for each time segment in the past N complete periods, forming a peak sequence L_peak(t).
[0185] Monitoring systems (such as Prometheus) continuously collect the length data of each subtask queue, using fixed time segments (such as every 15 minutes) as windows, and record the peak (maximum) queue length within each window rather than the average, in order to capture the worst-case scenario.
[0186] Sequence Construction: Select N complete and comparable periods from the past (e.g., N=4, period is 1 week), extract the peak values of the same time segment within each period (e.g., every Monday morning 9:00-9:15), arrange them in chronological order to form the peak sequence L_peak(t). t represents the segment index arranged in chronological order.
[0187] This step captures the true peak stress experienced by the system during the same historical period. By analyzing peak values rather than average values, the resulting model is better equipped to handle worst-case scenarios, providing the system with sufficient safety margin.
[0188] Q411, perform time series decomposition on L_peak(t) to extract its seasonal component C(t) and trend component T(t).
[0189] Apply a time series decomposition algorithm (such as STL, X-12-ARIMA, or Prophet) to L_peak(t). This algorithm can decompose the sequence into three main components:
[0190] Seasonal component C(t): reflects the regular fluctuations that recur within a fixed period (such as daily or weekly).
[0191] Trend component T(t): reflects the long-term upward or downward trend of the sequence.
[0192] Residual component R(t): represents random noise that cannot be explained by periodicity and trend (not extracted in this step).
[0193] After the algorithm runs, it outputs the seasonal component value C(t) and the trend component value T(t) for each time point t.
[0194] This step deconstructs the complex patterns of queue load. It breaks down seemingly chaotic peak data into interpretable, regular components (seasonality and trends), laying the mathematical foundation for subsequent accurate predictions.
[0195] Q412. Based on the phase of the current time point t_now in the cycle, determine the corresponding seasonal component value C(t_now), and use the trend component to predict the current trend value T(t_now), and determine the expected value of the base load L_base=C(t_now)+T(t_now).
[0196] Calculate the specific position (phase) of the current time point t_now within a complete cycle (such as one week). For example, Tuesday at 10:15 AM is the Xth time segment after the start of the cycle.
[0197] Obtain the seasonal value: Based on the phase of t_now, find the value C(t_now) of the corresponding phase from the seasonal component C(t) obtained by decomposition.
[0198] Predicting trend value: Using a model (such as linear regression or moving average) based on the trend component T(t), extrapolate and predict the trend value T(t_now) at the current time point t_now.
[0199] Calculate the expected value: Add the two together: L_base = C(t_now) + T(t_now). L_base represents the typical peak load expected at the current point in time based on historical patterns.
[0200] This step enables intelligent load prediction based on historical patterns. It allows the system to "anticipate" upcoming regular load peaks (such as the weekly build peak on Mondays), thus allowing for advance resource planning.
[0201] Q413, obtain the historical average execution time t_exec_avg of the subtask of the corresponding type of RA, the average scheduling startup time t_schedule_avg of the omnipotent execution instance, and the average startup time t_startup_avg of the dedicated execution instance.
[0202] Average execution time (t_exec_avg): The total execution time of RA type subtasks in the most recent statistical period is calculated from the task execution log, and then divided by the number of tasks to obtain the average execution time.
[0203] Average scheduling startup time (t_schedule_avg): Calculated from system event logs, this is the average time historically taken to successfully schedule and bind an idle omnipotent instance to a queue. It includes steps such as tag update, Agent reconfiguration, and queue listening taking effect.
[0204] Average startup time (t_startup_avg): The average time taken from creating a new dedicated execution instance Pod in history until its status is Ready and its internal agent is successfully registered, based on K8s events and Pod status logs.
[0205] This step quantifies the time cost of different elastic response paths. It clarifies the preparation time required for the two core response strategies, "scheduling existing idle instances" and "creating new dedicated instances," providing key input for calculating the necessary buffer size.
[0206] Q414. Based on t_exec_avg, t_schedule_avg, and t_startup_avg, determine the queue length compensation amount ΔL = max(t_schedule_avg, t_startup_avg) / t_exec_avg; max() is the function to find the maximum value.
[0207] The physical meaning of the compensation amount ΔL is: in the worst case (the response time is the longer of the two paths), the number of tasks that an existing execution instance can handle during the time it takes for the system to decide to take action and for new processing capacity to become available. Taking the maximum value (max) ensures that the buffer is sufficient to cover any possible response latency.
[0208] Example: If t_schedule_avg = 45 seconds, t_startup_avg = 90 seconds, and t_exec_avg = 30 seconds, then ΔL = max(45, 90) / 30 = 90 / 30 = 3. This means that in order to cover a resource delivery delay of up to 90 seconds, the queue needs to buffer an additional 3 tasks.
[0209] This step transforms resource delivery latency into measurable task buffering demand. This transformation ensures that the expansion trigger point is set early enough that when new resources become available, the backlog of tasks remains manageable, preventing uncontrolled queue growth due to response latency.
[0210] Q415. Based on L_base and ΔL, determine the initial threshold L_init = L_base + ΔL.
[0211] L_init is the initial expansion trigger line. It requires that the queue length simultaneously exceed both the predicted normal load and the buffer needed to handle latency before triggering subsequent scheduling or expansion checks.
[0212] This step led to the formation of preliminary decision-making criteria that comprehensively consider both routine load and contingency buffers. These criteria respect the system's historical operating patterns while providing a safety net for resource supply delays, ensuring that capacity expansion decisions are both predictable and robust.
[0213] Q416. Obtain the real-time availability ratio P_avail of the omnipotent execution instances in the current K8s cluster, and calculate the elastically corrected threshold L_adj=L_init×[1+τ×(1-P_avail)]; where τ is the preset elasticity coefficient, 0<τ<1.
[0214] The system queries the status of all omnipotent execution instances in real time, calculates the proportion of instances currently in an "idle" state to the total number of omnipotent instances, and obtains P_avail. τ is a preset elasticity coefficient (e.g., 0.3) that controls the adjustment range.
[0215] This formula implements reverse adjustment. When P_avail is high (the shared resource pool is ample), (1-P_avail) is small, and L_adj is slightly greater than or equal to L_init. The system tends to slightly increase the threshold because low-cost scheduling resources are sufficient, allowing for more "leisure" scheduling and avoiding premature scheduling, prioritizing the consumption of idle resources.
[0216] When P_avail is low (shared resource pool is strained), (1-P_avail) is large, and L_adj is significantly larger than L_init. The system will significantly increase the threshold. This is because the success rate of triggering "schedule idle instances" is low at this time, and it is more likely to directly proceed to "expand dedicated instances". Since the latter is more costly, the system needs to be more "cautious" and only execute it when the queue backlog is more severe and longer-lasting, which is equivalent to raising the trigger threshold for dedicated expansion.
[0217] This step is the core of the solution's innovation. It makes the expansion trigger threshold a dynamic lever that can automatically adjust the system's preference between "fast scheduling" and "robust expansion" strategies based on the abundance of low-cost resources (universal instances), achieving a fine balance between cost and response speed.
[0218] Q417, min(L_adj,L_max) is determined as the first preset queue length threshold; L_max is the preset static system upper limit value; min() is the function to find the minimum value.
[0219] Set a static upper limit for the queue length L_max based on the message queue's maximum capacity, system service level objectives (SLO), or operational experience.
[0220] This step sets the final safety boundary. No matter how high the threshold calculated by the preceding intelligent algorithm, this absolute upper limit, L_max, prevents the trigger line from being set too high, avoiding system sluggish response to queue backlog due to model bias or extreme cases, and ensuring the bottom line of service quality.
[0221] The above steps construct a four-stage, multi-factor fusion dynamic threshold decision-making model. First, time-series analysis and prediction enable the system to "foresee the future," allowing the threshold to adapt to periodic load changes. Second, by quantifying latency costs, the physical latency of resource supply is transformed into explicit task buffer requirements, injecting a "safety buffer" into the threshold. Third, through resource status-aware elastic correction, the real-time status of the shared resource pool is innovatively used as a regulating valve, enabling the threshold to dynamically balance "rapid scheduling" and "more costly dedicated expansion" strategies, achieving optimal cost-effectiveness in resource utilization. Finally, system upper limit constraints provide ultimate security. This entire mechanism works collaboratively, producing not a static configuration number, but an intelligently predictive, dynamically adaptable, economically optimal, and secure elastic trigger signal. This significantly improves the accuracy, economy, and reliability of the entire elastic scaling system, enabling it to handle complex real-world production environment loads with ease. It is a core intelligent decision-making component for building a high-standard cloud-native CI / CD platform.
[0222] Q500: If there are idle execution instances in the execution instances corresponding to other subtask queues besides RA, then several idle execution instances will be scheduled to RA; otherwise, the number of dedicated execution instances for RA will be expanded; the dedicated execution instances can execute the subtasks of the corresponding subtask queues.
[0223] If step Q400 finds available idle instances, the scheduler selects several instances according to a preset strategy (such as least usage priority, nearest geographical location priority, etc.) and "temporarily schedules" them to the RA queue in the following way:
[0224] Tag reconfiguration: Modify the Jenkins Agent tags for these instances to add exclusive tags for the RA queue.
[0225] Queue Binding: Update the configuration of the consumer services within these instances so that they listen to the RA queue simultaneously.
[0226] Dedicated expansion trigger: If no available idle instances are available, a dedicated expansion process is triggered.
[0227] Template Acquisition: Based on the task type of RA, the corresponding K8sDeployment template is retrieved from the predefined dedicated instance template library.
[0228] Scaling up: Calculate the required number of dedicated instances based on metrics such as the backlog of the RA queue and the average processing time of tasks.
[0229] Instance creation: The Kubernetes API is called to increase the number of replicas for the corresponding Deployment and create new dedicated Pod instances. The images for these Pods are specifically optimized for Ragnarok Online (RA) task types.
[0230] Resource registration: After the newly created dedicated instance starts, it is registered with the Jenkins Master with the RA exclusive tag, and its consumer service is configured to listen only to the RA queue.
[0231] This step prioritizes the allocation of idle resources, enabling rapid and low-cost emergency response; dedicated capacity expansion is only performed when necessary, avoiding over-provisioning of resources. This tiered strategy achieves an optimal balance between response speed and resource cost.
[0232] Furthermore, the expansion of dedicated execution instances for RA includes the following steps:
[0233] Q510. Based on the current backlog of configuration information L_current in RA and the first preset queue length threshold L_threshold, determine the amount of tasks to be resolved ΔL=L_current-L_threshold.
[0234] The current message count L_current of the subtask queue RA is queried in real time by calling a message queue service (such as RabbitMQ's HTTP API or monitoring interface). The first preset queue length threshold L_threshold, calculated in real time for the RA queue, is obtained from the dynamic threshold calculation module or cache. The result ΔL represents the backlog of tasks that exceeds the system's expected buffering range and must be processed using additional dedicated resources.
[0235] This step transforms the expansion decision from a qualitative perception of "the queue is too long" into a quantitative objective of "needing to process ΔL additional tasks," providing clear and objective input for subsequent refined capacity calculations and avoiding blind expansion.
[0236] Q511, obtain the historical average execution time t_exec_avg of the task type corresponding to RA and the preset target queue clearing time T_target, and determine the parallel capability increment P_need=ΔL×t_exec_avg / T_target.
[0237] Obtain the historical average execution time t_exec_avg of the corresponding task type for RA in the most recent statistical period from the monitoring database or task log analysis system (e.g., each task executes for an average of 120 seconds).
[0238] T_target is a preset Service Level Target (SLO) parameter, set by operations personnel based on business priority and user experience requirements. For example, for core service queues, T_target can be set to 300 seconds (5 minutes), requiring backlogged tasks to be cleared within 5 minutes; for non-core tasks, T_target can be set to 1800 seconds (30 minutes).
[0239] ΔL×t_exec_avg represents the total workload time (in seconds) required to process all backlogged tasks. Dividing this total workload by the target duration T_target gives the additional "processing rate" (in tasks / second) the system needs to provide to complete the work within the target time. More intuitively, P_need is a parallel processing capability coefficient. For example, if P_need = 2.5, it means that an additional processing rate equivalent to 2.5 times the baseline is required.
[0240] This step is the key to the solution's innovation. It transforms scaling decisions from a passive, reactive process (scaling up when the queue gets too long) into a proactive, goal-oriented process (how much scaling is needed to clear the queue within X minutes). This aligns resource allocation with clear business objectives, enabling predictable and measurable elastic scaling.
[0241] Q512, obtain the standard task processing capacity P_unit of the dedicated execution instance, and determine the theoretical number of instances that need to be expanded N_need=ceil(P_need / P_unit); ceil() is the preset rounding up function.
[0242] P_unit represents the standard task processing capacity of a dedicated execution instance. This is a predefined baseline value, typically related to the baseline of t_exec_avg. In the simplest model, P_unit can be defined as 1, indicating that an instance can process "1 unit" of work per unit time (its dimension is consistent with P_need). A more precise approach is to calibrate P_unit based on instance specifications (CPU / memory) and actual performance benchmarks.
[0243] Example: Assuming P_need=2.5 and P_unit=1, then N_need=ceil(2.5)=3. This means that 3 new dedicated execution instances are needed to provide a 2.5x increase in processing power (because instance capabilities are discrete).
[0244] By introducing the standardized parameter P_unit, the capabilities of instances of different specifications or performance can be uniformly measured and calculated, achieving an accurate mapping from business objectives (processing rate) to infrastructure resources (number of instances).
[0245] Q513: Query the node resource status in the K8s cluster that can be used to deploy dedicated execution instances for the corresponding task type of RA. Based on node affinity and anti-affinity rules, and the real-time remaining schedulable resources of each node, determine the maximum number of instances that can be safely deployed, N_max.
[0246] Call the K8s API (such as kubectl top nodes combined with kubectl describe nodes) or through the cluster monitoring system to obtain the allocatable and allocated resources (total requests) of all worker nodes, and calculate the real-time remaining schedulable resources (number of CPU cores and GB of memory) of each node.
[0247] Node affinity: Check if nodeSelector or nodeAffinity is defined in the Pod template of the RA-specific instance to ensure that only nodes that meet the tag requirements are considered.
[0248] Pod Anti-Affinity: If the Pod template defines podAntiAffinity (for example, to prevent two Pods of the same application from being scheduled to the same node), scheduling needs to be simulated to ensure that new instances are distributed to different nodes.
[0249] Resource fragmentation and binning: Taking into account the remaining resources of each node, the algorithm uses best-fit or worst-fit to simulate and calculate the maximum number of Pods that can be scheduled under the current resource fragmentation conditions.
[0250] Determine N_max: Taking into account all the above constraints, calculate the maximum number of dedicated execution instances that can be safely deployed for the RA queue under the current cluster state, without causing any node overload and satisfying all scheduling rules.
[0251] Beneficial effects: It achieves global resource security verification and boundary control for expansion operations. This step is core to ensuring the overall stability of the cluster. It ensures that expansion operations do not "grow wildly," and will not exhaust node resources, cause node pressure, or violate high availability rules (such as anti-affinity) in order to meet the needs of a single queue. This reflects the system design philosophy of examining local resilience from a global perspective.
[0252] Q514. Based on N_need and N_max, determine the actual expansion quantity N_scale = min(N_need, N_max); min() is the preset minimum value function; and expand the RA for dedicated execution instances based on N_scale.
[0253] Perform a capacity expansion operation:
[0254] API call: Call the Kubernetes API via a Kubernetes client library (such as client-go) to update the `replicas` field of the Deployment corresponding to the RA queue. New replica count = current replica count + N_scale.
[0255] Elegant creation: The Kubernetes scheduler will schedule the new Pod to the appropriate node and start the container based on the constraints (affinity, resources, etc.) considered in step Q513.
[0256] Instance registration: After the dedicated Jenkins Agent in the new Pod starts, it registers with the JenkinsMaster with a unique tag for the RA queue and starts a consumer service that only listens to the RA queue.
[0257] State synchronization and monitoring: Update the instance mapping in the central scheduler and start monitoring the startup status of new Pods to ensure they are successfully ready.
[0258] This step completes the final closed loop from intelligent decision-making to secure, automated execution. The `min()` function ensures that the scaling operation is both proactive (striving to achieve the target SLO) and absolutely robust (never crossing the cluster's security red lines). Automated API calls instantly translate precise calculations into actual resource provision, significantly shortening the cycle from demand identification to resource readiness. This guarantees that under high load scenarios, the system can obtain the required processing power in the fastest and safest way, thereby reliably achieving the preset service level objectives.
[0259] Q600: If the number of configuration information in RA is less than or equal to the second preset queue length threshold and RA has a dedicated execution instance, then the dedicated execution instance corresponding to RA will be scaled down.
[0260] If the length of queue RA is consistently lower than the second preset queue length threshold (and RA currently has a dedicated execution instance), a scaling-down assessment will be triggered.
[0261] The scheduler selects one or more instances from the RA's dedicated instance pool as scaling-down candidates. The selection strategy can be based on: instance idle time (prioritizing the longest idle time), instance creation time (prioritizing newer or older instances, depending on the strategy), instance resource specifications (prioritizing the release of instances with high resource consumption), etc.
[0262] Furthermore, the second preset queue length threshold is determined through the following steps:
[0263] Q610: Obtain the typical idle queue level of RA during the same historical period as a basic reference value.
[0264] Extract the queue length data of the subtask queue RA from historical monitoring data for each same time segment (such as Wednesday morning 10:00-10:15) within multiple complete cycles (such as the past 4 weeks).
[0265] Statistical analysis is performed on the historical queue length sequence within this time period. To remove outliers, statistical methods are typically used (such as calculating the 10th percentile or taking the average after removing the highest and lowest 5%) to determine the typical idle queue level. This value represents the common "low waterline" of the RA queue during the same historical period when the system load is calm.
[0266] Example: Suppose the queue lengths at 10 AM on the past four Wednesdays are [0, 1, 0, 2]. Taking the 10th percentile (or a reasonable lower limit) might be 0, but to avoid being too sensitive, it can be slightly increased to 1.
[0267] This step uses historical data from the same period to ensure that the baseline matches the current business cycle (e.g., weekday / weekend), avoiding misleading comparisons across time periods. This baseline reflects the inherent "background noise" level of the system under no-stress conditions, providing a stable and reliable starting point for subsequent adjustments.
[0268] Q611, Based on the average resource utilization of the dedicated execution instance cluster corresponding to RA, the basic reference value is positively adjusted: the lower the average resource utilization, the higher the adjusted value.
[0269] Collect the CPU and memory usage of all dedicated execution instance Pods corresponding to the RA queue within the current statistical period (e.g., the last 15 minutes). For each resource type, calculate the average utilization of all instances, and then weight them (e.g., CPU weight 0.7, memory weight 0.3) to obtain the comprehensive average resource utilization (U_avg), which is between 0 and 1.
[0270] The lower the U_avg value, the more idle the dedicated resources allocated to the RA are, and the more significant the cost waste. Therefore, scaling down should be triggered more aggressively. To achieve "more aggressive" scaling down, the scaling down threshold (L_threshold) needs to be increased, making it easier for the current queue length to fall below this threshold and thus meet the scaling down conditions.
[0271] Adjustment formula: Adjustment value 1 = base reference value × α, where adjustment factor α = 1 + k1 × (1 - U_avg). k1 is a preset positive adjustment coefficient (0 < k1 < 1, for example 0.5).
[0272] Example: The base reference value is 1, U_avg=0.2 (utilization rate 20%), k1=0.5, then α=1+0.5×(1-0.2)=1.4, and the adjustment value 1=1×1.4=1.4.
[0273] This step directly transforms the cost of "resource idleness" into a technical decision parameter (raising the threshold), driving the system to automatically adopt a more aggressive resource recovery strategy when resource utilization is low, thereby effectively reducing unnecessary resource holding costs.
[0274] Q612, based on the availability ratio of all-purpose execution instances in the current K8s cluster, the adjusted value is adjusted a second time in a positive direction; the higher the availability ratio, the higher the value after the second adjustment.
[0275] Query the real-time status (busy / idle) of all omnipotent execution instances. Calculate the number of omnipotent instances currently in the "idle" state, divide it by the total number of omnipotent instances, and obtain the omnipotent instance availability ratio (P_avail).
[0276] Adjustment logic: A higher P_avail means a more abundant global shared resource pool and stronger elastic buffering capability. In this case, even if the dedicated instances of RA are scaled down, there are still enough omnipotent instances that can be quickly scheduled as an emergency measure should there be a subsequent surge in load. Therefore, the scale-down threshold can be further increased aggressively.
[0277] Adjustment formula: Adjustment value 2 = Adjustment value 1 × β, where the secondary adjustment factor β = 1 + k2 × P_avail. k2 is another preset positive adjustment coefficient (0 < k2 < 1, for example 0.3).
[0278] Example: If adjustment value 1 is 1.4, P_avail=0.8 (80% of omnipotent instances are idle), k2=0.3, then β=1+0.3×0.8=1.24, and adjustment value 2=1.4×1.24≈1.74.
[0279] This step is a key aspect of the solution's innovation. It recognizes that scaling down decisions cannot be made in isolation and must consider the system's overall "backup" capability. By using the abundance of global backup resources (the omnipotent instance pool) as an adjustment factor, the system can dynamically assess the risks of scaling down operations: if backup resources are sufficient, scaling down is aggressive (high threshold); if backup resources are scarce, scaling down is cautious (low threshold). This achieves intelligent linkage between local cost optimization and global risk control.
[0280] Q613, the maximum value between the final adjusted value and the preset system safety lower limit value is determined as the second preset queue length threshold.
[0281] The system safety minimum limit (L_safe) is a static parameter preset based on business continuity and system reliability requirements. For example, it can be set to 1 to ensure that scaling down is only considered when the queue is completely empty; or it can be set to (minReplicas × buffer coefficient), where minReplicas is the minimum number of dedicated instances that the RA queue is required to retain.
[0282] Execute the final decision: Compare the adjusted value 2 (after secondary adjustment) with the safety lower limit L_safe, and take the maximum value as the final determined second preset queue length threshold. That is, final threshold = max(adjustment value 2, L_safe).
[0283] As a final verification, it should also be ensured that the calculated final threshold is less than the first preset queue length threshold to prevent the shrinking line from overlapping or being reversed with the expansion line, thus avoiding logical conflicts.
[0284] No matter how low the threshold calculated by the preceding intelligent algorithm (e.g., due to abnormal historical data or extreme parameters), the safety lower bound L_safe acts like a "hard brake," ensuring that the scaling-down trigger point will not fall below the minimum required for safe system operation. This fundamentally prevents the risk of excessive scaling-down due to model fluctuations or algorithm defects, guarantees the basic availability and responsiveness of the service, and is the ultimate guarantee of system robustness.
[0285] Furthermore, step Q400 includes the following steps:
[0286] Q420, if the number of configuration information in RA is greater than or equal to the first preset queue length threshold, then obtain the first duration for which the number of configuration information in RA is greater than or equal to the preset queue length threshold.
[0287] This step is the same as step S400 in Embodiment 2, and will not be described again here.
[0288] Q421. If the first duration is greater than or equal to the first preset duration, check whether there are any idle execution instances in the execution instances corresponding to other subtask queues besides RA.
[0289] The check operation is performed by the central resource scheduler. The scheduler queries its instance status database, which maintains the real-time status (idle / busy) of all execution instances, including all omnipotent execution instances and dedicated execution instances of all other subtask queues.
[0290] This step is the master switch for the two-tiered elastic response strategy (schedule first, then scale up). Only after confirming that the RA queue is facing continuous and genuine pressure will the lower-cost, faster resource scheduling path be initiated. This ensures that high-value global scheduling operations only occur when necessary, guaranteeing timely relief for high-load queues while minimizing disruption to other queues and shared resource pools, demonstrating the precision and economy of resource allocation. Simultaneously, it establishes a pre-approval threshold for potential subsequent dedicated instance scaling, further optimizing the resource allocation decision-making process.
[0291] Furthermore, step Q600 includes the following steps:
[0292] Q620, if the number of configuration information in RA is less than or equal to the second preset queue length threshold, then obtain the second duration for which the number of configuration information in RA is less than or equal to the second preset queue length threshold.
[0293] This step is the same as step S600 in Embodiment 2, and will not be described again here.
[0294] Q621, if the second duration is greater than or equal to the second preset duration, then the dedicated execution instance corresponding to RA is scaled down.
[0295] In this step, the method described in steps S710-S730 of Embodiment 2 can be used for scaling down, which will not be elaborated here. Optionally, all dedicated execution instances corresponding to RA can be directly canceled to achieve rapid scaling down.
[0296] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0297] Embodiments of the present invention also provide a non-transitory computer-readable storage medium that can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiments, wherein the at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiments.
[0298] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0299] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0300] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0301] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0302] Embodiments of the present invention also provide an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0303] The electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments in this application.
[0304] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and a bus connecting different system components (including memory and processor).
[0305] The memory stores program code that can be executed by the processor, causing the processor to perform the steps in the various embodiments described in this specification.
[0306] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).
[0307] The memory may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0308] A bus can represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus structures.
[0309] Electronic devices can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable user interaction with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be achieved through input / output (I / O) interfaces. Furthermore, electronic devices can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via network adapters. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0310] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0311] Embodiments of the present invention also provide a computer program product including program code, which, when the program product is run on an electronic device, causes the electronic device to perform the steps of the methods described above in various exemplary embodiments of the present invention.
[0312] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention.
Claims
1. A method for scaling up and down task execution resources based on Jenkins and Kubernetes, characterized in that, Includes the following steps: S100, obtain each preset subtask queue; wherein, each subtask queue is used to store subtask configuration information of a pipeline task type, and each subtask queue corresponds to several execution instances for executing that type of subtask. S200 responds to any subtask QR corresponding to the Jenkins build-to-deploy task after it has been completed and stores the configuration information corresponding to the QR into the subtask queue corresponding to the QR. S300: For any subtask queue RA, if there is an idle execution instance among the several execution instances corresponding to RA, the configuration information in RA is sent to the idle execution instance, so that the idle execution instance executes the subtask corresponding to the received configuration information. S400, if the number of configuration information in RA is greater than or equal to the first preset queue length threshold, then obtain the first duration for which the number of configuration information in RA is greater than or equal to the preset queue length threshold. S500, if the first duration is greater than or equal to the first preset duration, then the execution instance corresponding to RA is expanded; S600, if the number of configuration information in RA is less than or equal to the second preset queue length threshold, then obtain the second duration for which the number of configuration information in RA is less than or equal to the second preset queue length threshold; S700, if the second duration is greater than or equal to the second preset duration, then the execution instance corresponding to RA is scaled down.
2. The method for scaling up and down task execution resources based on Jenkins and Kubernetes according to claim 1, characterized in that, The pipeline task types include: production release, canary release, test release, and continuous integration build.
3. The method for scaling up and down task execution resources based on Jenkins and Kubernetes according to claim 1, characterized in that, The first preset queue length threshold is determined through the following steps: S410: Obtain all subtasks completed by RA within the past statistical period T, and calculate the corresponding average execution time t_task; S420 obtains the average cold start time t_startup required to start a corresponding execution instance for RA from the K8s cluster monitoring data; S430, based on t_task and t_startup, determine the base threshold N_base = ceil(t_startup / t_task); where ceil() is the floor function; S440, get the average number of tasks N_blocked in RA due to waiting for the execution instance to start within T; S450, based on N_base and N_block, determine the cold start compensation factor α = 1 + (N_block / N_base); S460, based on α and N_base, determine the first preset queue length threshold L_threshold = α × N_base.
4. The method for scaling up and down task execution resources based on Jenkins and Kubernetes according to claim 1, characterized in that, The first preset duration is determined through the following steps: S510, Based on the historical sequence of queue length of RA in the most recent M statistical windows, determine the average rate of change V_avg and instantaneous volatility σ of the queue length corresponding to RA; S511, based on V_avg, determine the initial reference duration T_base=β / (|V_avg|+ε); where β is the preset system response coefficient and ε is the preset constant; S512, based on σ, determine the fluctuation suppression factor γ = 1 + (σ / σ_ref); where σ_ref is the reference benchmark value for queue length fluctuation; S513, based on the ratio of remaining schedulable resources available for RA in the current K8s cluster, R_available, determine the resource pressure coefficient η=max(0.5,R_available); max() is a preset function to find the maximum value; S514, determine the first preset duration T_threshold=(T_base×γ) / η based on T_base, γ and η; S515, constrain T_threshold so that T_threshold is within the feasible operating range [T_min, T_max]; T_min is the first preset minimum duration, and T_max is the first preset maximum duration.
5. The method for scaling up and down task execution resources based on Jenkins and Kubernetes according to claim 1, characterized in that, Expanding the execution instance corresponding to RA includes the following steps: S520, based on the pipeline task type LA corresponding to RA, obtains the resource weight coefficient W and the number of parallel tasks per instance P of LA from the preset configuration table; S521, Based on the length of RA L_current and the first preset queue length threshold L_threshold, determine the number of tasks to be resolved based on queue pressure ΔQ=L_current-L_threshold; S522, based on ΔQ and P, determine the theoretical number of new instances N_need required to meet the current queue pressure: N_need = ceil(ΔQ / P); ceil() is the floor function. S523, query the total amount of idle resources currently available for LA in the K8s cluster, and calculate the maximum number of instantaneous expansion instances N_max allowed under the current resource conditions, in combination with the resource weight coefficient W; S524, based on N_need and N_max, determine the number of expanded instances N_scale=min(N_need,N_max), and call the K8s cluster interface to create N_scale new execution instances for RA; min() is the preset minimum value function.
6. The method for scaling up and down task execution resources based on Jenkins and Kubernetes according to claim 1, characterized in that, The second preset queue length threshold is determined through the following steps: S610, obtain the average resource utilization rate of the execution instance cluster corresponding to RA in the current statistical period; S611, Determine the resource release conservatism coefficient based on the average resource utilization rate; wherein, the lower the average resource utilization rate, the larger the resource release conservatism coefficient. S612, obtain the typical minimum queue length of RA during the historical steady state period; S613, Determine the basic shrinkage threshold based on the typical minimum queue length and the resource release conservative coefficient; S614. Determine the cost efficiency factor based on the average task digestion rate corresponding to RA and the keep-alive maintenance cost index of the execution instance. S615, combine the basic shrinkage threshold with the cost efficiency factor to obtain a preliminary queue length threshold; S616, the maximum value between the initial queue length threshold and the preset absolute safety lower limit is determined as the second preset queue length threshold.
7. The method for scaling up and down task execution resources based on Jenkins and Kubernetes according to claim 1, characterized in that, The second preset duration ranges from 3 minutes to 6 minutes.
8. The method for scaling up and down task execution resources based on Jenkins and Kubernetes according to claim 1, characterized in that, Scaling down the execution instance corresponding to RA includes the following steps: S710, for a plurality of execution instances corresponding to RA, identify at least one target execution instance that is in an idle state and whose continuous idle time exceeds a third preset duration; S720, determine the number of instances to be released from the target instances; S730, call the K8s cluster interface to delete the execution instances corresponding to the number of instances to be released, and update the mapping relationship between RA and execution instances.
9. A non-transitory computer-readable storage medium, wherein the storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the task execution resource scaling method based on Jenkins and Kubernetes as described in any one of claims 1-8.
10. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 9.
Citation Information
Patent Citations
Dynamic capacity expansion or shrinkage method, device, equipment and medium
CN115632949A
Method, device and equipment for pre-expanding and shrinking capacity of K8s cluster and storage medium
CN121173683A