Method and device for supporting elastic scaling capability of intelligent computing center computing power resources
Through a three-level cache system and a dynamic hierarchical resource scheduling mechanism, the problems of high loading delay and low resource utilization of large AI models in intelligent computing centers are solved, rapid fault recovery and refined resource management are achieved, and cross-node scheduling is optimized.
Patent Information
- Application Number
- CN202511049322.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing technologies in intelligent computing centers suffer from high loading latency of large AI models, low resource utilization, and cross-node scheduling issues. Furthermore, the lack of effective cache management strategies leads to severe cold start latency and insufficiently refined resource scheduling.
It adopts a three-level cache system and a dynamic hierarchical resource scheduling mechanism. By real-time monitoring of QPS indicators, it dynamically divides service levels, and combines pre-caching and LRU-LFU hybrid elimination algorithms to optimize cross-node resource scheduling and cache management to achieve rapid fault recovery.
It significantly reduced cold start latency, improved resource utilization, shortened fault recovery time, and solved engineering problems in cross-node scheduling, enabling more refined resource management.
Smart Images

Figure CN120560858B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computing power infrastructure, in particular to a method and device for supporting elastic scaling capability of computing power resources in intelligent computing center. BACKGROUND
[0002] The intelligent computing center supports AI applications such as large model inference by integrating CPUs, GPUs and other heterogeneous computing power. Although the Serverless architecture provides automatic scaling capability, it faces the serious challenge of extremely high cold start delay in the AI scenario. When loading a model file (such as LLaMA-70B) of hundreds of GB and a container image of GB, the network transmission and initialization time consumption can reach 30 minutes. For example, it takes about 80 seconds to transfer 100 GB at 10 Gbps bandwidth, and additional time is consumed for decompression.
[0003] To solve the technical problem, the patent document with publication number CN120086003A discloses a method and device for supporting elastic scaling capability of computing power resources in intelligent computing center. The method includes monitoring the call frequency of the target service, determining the service level to which the target service currently belongs according to the call frequency, wherein the service level includes high-frequency service, medium-frequency service, low-frequency service and no-cache service, and running the target service according to the service level to which the target service currently belongs. The intelligent computing center can dynamically adjust the allocation of computing power resources according to the actual demand of the service, realize the elastic scaling capability of the computing power resources, avoid the start delay of the service, and improve the user experience.
[0004] However, the above-mentioned scheme has the following defects: 1. Only the cold start problem is mentioned generally, and the root cause of the large AI model loading delay is not quantified; 2. Lack of cache management strategy (such as triggering LRU eviction when the disk is full 65%); 3. The complete life cycle of the specific scene from no cache to high-frequency service is not shown; 4. The actual constraints are not considered, and the engineering problems such as cross-node resource scheduling, network delay and fault recovery are ignored. SUMMARY
[0005] The present application provides a method and device for supporting elastic scaling capability of computing power resources in intelligent computing center, which uses a three-level cache system to accelerate cold start, effectively shortens the fault recovery time, solves the problems of high AI model loading delay, low resource utilization and cross-node scheduling, and solves the problems raised in the background technology.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a method for supporting elastic scaling capability of computing power resources in intelligent computing center, characterized in that it comprises the following steps:
[0007] S1: Real-time monitoring of QPS indicators of all services deployed on the Kubernetes cluster, and collecting original QPS data through Prometheus;
[0008] S2: Calculate the set time sliding window average of the original QPS data to smooth the instantaneous flow fluctuations, and store the processed indicators in the time series database;
[0009] S3: Dynamic division of service state level based on preset QPS threshold: low frequency, medium frequency, high frequency;
[0010] S4: Dynamic enhancement adjustment: threshold is not completely static, dynamically fine-tune the threshold according to historical data, prediction model or current cluster overall load, determine the current load level of each service to make decision output;
[0011] S5: Perform hierarchical resource scheduling according to service level: HF service: deploy ≥2 container instances, dynamically scale based on HPA according to QPS; MF service: keep 1 running instance, trigger promotion or demotion according to QPS change; LF service: release running instance, only keep SSD node pre-cache resource;
[0012] S6: Perform queue state migration based on promotion, demotion and demotion paths.
[0013] Preferably, the following steps are further included:
[0014] S101: Step S1 further includes verifying the adjustment factor:
[0015] Perform request delay (P50, P90, P99 delay) to determine whether scaling effectively improves performance, or as a trigger condition in addition to QPS;
[0016] Verify whether resource bottleneck is indeed caused by QPS through container resource utilization, and prevent over-scaling;
[0017] Monitor error rate as a supplementary condition for scaling or alarm.
[0018] Preferably, the following steps are further included:
[0019] S201: Step S2 further includes that the sliding window average QPS calculated in S2 is used to determine the reasonable threshold range of different services based on historical data analysis, and the preset threshold is determined. Analysis should include: typical business cycle traffic pattern, service performance baseline, cost constraint target.
[0020] Preferably, the following steps are further included:
[0021] S501, in step S5, LF service does not run container instance for a long time, and preloads and caches the key resources required by the running instance to the local storage of the selected node;
[0022] When a low-frequency service receives the first request, the system will start a container instance trigger cache according to the standard process, and after the container instance successfully runs and processes the request, the system will automatically cache the required resources of the container instance to the SSD of the node where the container instance is located;
[0023] For services that are known to be used in advance, actively cache, reserve cache, and if there is no request to use the cache after a certain time, mark it as evictable; periodically scan and clean up caches that exceed TTL and have not been accessed.
[0024] Preferably, the scale-out control includes:
[0025] S502, step S5 sets the scale-out cooling period and the scale-in cooling period; performs maximum / minimum replica number restriction, sets MinReplicas and MaxReplicas for each service to prevent uncontrolled scale-out from exhausting resources; performs load balancing to distribute requests to healthy container instances.
[0026] Preferably, it further includes the following steps:
[0027] S601: Step S6 also includes an accelerated startup process:
[0028] When a low-frequency service receives a new request, the scheduler first checks the pre-cache pool. If there is valid cache, the scheduler preferentially schedules a new container instance to the SSD node that owns the service cache. Since the required resources are cached in the local SSD, the container instance startup process skips or greatly shortens the time of pulling images from the image repository and downloading large models / data from object storage;
[0029] If there is no valid cache, start the container instance according to the standard process, and trigger resource caching after successful startup. When there is no request, no container instance runs.
[0030] Preferably, it further includes the following steps:
[0031] S602: Step S6 also sets the cache eviction strategy, including the main eviction strategy LRU, which preferentially deletes the longest unused cache resources; the auxiliary strategy LFU, which deletes selected cache resource files and updates the cache metadata record.
[0032] An intelligent computing center computing resource supporting elastic scaling device, comprising a monitor: integrating Prometheus to collect QPS, Node Exporter to collect disk usage, and real-time calculation of sliding window mean;
[0033] Scheduler: built-in dynamic grading module, pre-cache management module;
[0034] Router: query resource cache location based on service ID, route request to SSD node, update node mapping table when failover;
[0035] Three-level cache system: memory, local SSD, distributed data archive;
[0036] Fault-tolerant processing unit: restart and update routing when instance crashes on same availability zone node;
[0037] Processor and memory, memory stores program code, when executed, realizes any of the above methods.
[0038] Compared with the prior art, the beneficial effects of the present application are:
[0039] 1. Cold start delay is reduced, and the cold start time of AI service is greatly reduced by pre-caching a hundred GB level model through SSD node, solving the large model loading bottleneck of the prior art which is not quantified.
[0040] 2. Resource utilization is improved, and LF service running instance is dynamically released, combined with LRU+TLL cache eviction, so that the resource utilization is improved compared with the prior art.
[0041] 3. State migration quantization rules are defined, and a QPS-time dual factor triggering mechanism is proposed, such as LF→MF needs to meet 5 minute QPS≥1, solving the problem of "general classification" leading to mis-scaling in the prior art.
[0042] 4. An engineering-level cache management scheme is proposed, an LRU-LFU hybrid eviction algorithm (threshold 65%) triggered by disk pressure is designed, avoiding cache invalidation caused by disk full in the prior art.
[0043] 5. Cross-node resource scheduling is supported, the scheduler makes node allocation decision (weight optimization) by comprehensively considering network delay, SSD space, and GPU utilization, solving the cross-node engineering problem ignored in the prior art.
[0044] 6. The fault recovery time is shortened to 15 seconds, and the routing table is updated through the same availability zone instance migration, which is about 5 times faster than the traditional scheme. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The method flowchart of the present application;
[0046] Figure 2 The monitor method flowchart of the present application;
[0047] Figure 3 The service state migration timing diagram of the embodiment of the present application;
[0048] Figure 4 The device structure diagram of the present application. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the protection scope of the present application.
[0050] Embodiment 1, please refer to Figures 1-3 The present application provides a method for supporting elastic scaling capability of intelligent computing center computing power resources, comprising:
[0051] Step S1: Real-time monitoring and index collection:
[0052] Monitor all services deployed on the Kubernetes group of the intelligent computing center, collect QPS through Prometheus (or other compatible monitoring systems such as VictoriaMetrics, Thanos); perform QPS aggregation calculation for each service ID, which is the core driving index of elastic scaling.
[0053] Step S2: Data processing:
[0054] Calculate the average value of the 5-minute sliding window for the original QPS data (usually one point per second), effectively smooth the transient traffic peaks (such as occasional retries, crawlers) and temporary troughs, and avoid false scaling decisions based on transient noise.
[0055] For example, given the QPS sequence [0, 3, 12, 8] (assuming 4 sampling points, but actually a continuous time sequence), the window average (0+3+12+8) / 4 = 5.75. In actual implementation, the data points contained in the window are much more than 4 (usually 300, one point per second).
[0056] The processed index (window average QPS, etc.) is stored in the Prometheus TSDB or similar time series database.
[0057] Step S3: Dynamic classification:
[0058] Mainly based on the 5-minute sliding window average QPS calculated in S2, determine the reasonable threshold range preset threshold for different services based on historical data analysis, the analysis should include: typical business cycle (day, week, month) traffic pattern, service performance baseline (QPS load capacity under SLA), cost constraint target.
[0059] Preset QPS threshold for each service (or service type), define service state level:
[0060] Low frequency: Avg QPS < LF_Threshold (e.g. LF_Threshold = 1 QPS);
[0061] Medium frequency: LF_Threshold <= Avg QPS < MF_Threshold (e.g. MF_Threshold = 5 QPS);
[0062] High frequency: Avg QPS >= MF_Threshold;
[0063] Step S4: Dynamic enhancement adjustment:
[0064] Threshold is not completely static, according to historical data (such as last week the same day the same hour), prediction model (simple like moving average prediction, complex like machine learning model) or current cluster overall load, within a certain range, dynamically fine-tune the threshold to adapt to business growth or special activities, determine the current load level (LF, MF, HF) of each service Decision output.
[0065] Step S5: Hierarchical resource scheduling:
[0066] High frequency service processing:
[0067] The minimum instance requirement must be deployed at least 2 container instances, which is the basis for realizing high availability (HA) and load balancing, avoiding single point of failure.
[0068] Dynamic scaling rules based on the current average QPS determined in S2; The target value is to set the target average QPS that each instance can handle, for example, according to the stress test to determine the stable processing capacity of a single container instance is 10 QPS, and then calculate the required number of replicas.
[0069] Medium frequency service processing:
[0070] Keep 1 running container instance, make a compromise between cost and response ability, monitor QPS, and upgrade to HF or downgrade to LF according to the rules of S3.
[0071] Low frequency service optimization and pre-cache resource pool:
[0072] LF service request is rare, long-running instance wastes resources; Cold start time-consuming (pull image, initialization) causes slow response for the first / low frequency request.
[0073] Pre-cache resource pool, no long-running container instance, but pre-load the key resources (container image, model file, dependent library, configuration file) required by the running instance into the local storage of the selected node.
[0074] Cache content: including container images, large model files (such as medical imaging model svc_ctscan), frequently accessed static data / dependent libraries, service configuration files; storage location priority SSD storage node, SSD provides extremely fast read speed, significantly shortens the data loading time when the container instance starts.
[0075] Cache strategy: first request: when a low-frequency service receives the first request, the system will start a container instance according to the standard process to trigger caching. After the container instance successfully runs and processes the request, the system automatically caches the required resources of the container instance to the SSD of the node (or specified node pool) where the container instance is located.
[0076] For services that are known to be used soon (such as scheduled tasks), proactive caching is performed in advance. The retention time in this embodiment is selected to be 24 hours after the last access to the cached resource. If there is no request to use the cache within 24 hours, it is marked as eliminable.
[0077] Cache management: maintain a cache directory or database to record cached resources (service ID, resource ID / path, cache node, cache time, last access time). Periodically scan and clean up caches that exceed TTL and have not been accessed.
[0078] Step S6: queue state migration and service life cycle management
[0079] Upgrade path
[0080] No cache state -> low frequency (LF): the service receives the first request. The system starts a container instance according to the standard process. After the request is successfully processed, the system caches the resources required by the container instance to the node (or specified node pool), and the service state enters LF (cache exists, no running container instance).
[0081] Low frequency (LF) -> medium frequency (MF): for example, when the 5-minute sliding window average QPS of the service is >= LF_Threshold (for example, 1 QPS) for 5 minutes. The system quickly starts 1 container instance on the node with the cache, and the service state migrates to MF (1 running container instance).
[0082] Medium frequency (MF) -> high frequency (HF): when the 5-minute sliding window average QPS of the service is >= MF_Threshold (for example, 5 QPS) for 2 minutes. Trigger HPA scaling logic to scale the number of container instances to at least 2, and the service state migrates to HF (>= 2 running instances, HPA takes effect).
[0083] Downgrade path
[0084] High Frequency (HF) -> Medium Frequency (MF): When the 5-minute sliding window average QPS of the service stays < MF_Threshold (e.g. 5 QPS) for 10 minutes. HPA scale-in logic reduces the number of container instances to 1, service state migrates back to MF (1 running container instance).
[0085] Medium Frequency (MF) -> Low Frequency (LF): When the 5-minute sliding window average QPS of the service stays < LF_Threshold (e.g. 1 QPS) for 30 minutes. System releases the currently running 1 container instance. Service state migrates back to LF (no running container instance, cache resources reserved into TTL countdown).
[0086] As a further embodiment, the following steps are also included:
[0087] S101: In step S1, the adjustment factor is also verified:
[0088] Request latency (P50, P90, P99 latency) is performed to determine whether scaling effectively improves performance or as a trigger condition other than QPS (e.g. QPS does not reach threshold but latency soars).
[0089] Resource bottleneck is verified by container resource utilization (CPU utilization, memory utilization) to determine whether the bottleneck is indeed caused by QPS, and to prevent over-scaling (e.g. CPU is close to the upper limit, simply adding instances may be ineffective).
[0090] Error rate (HTTP 5xx error rate or service-specific error code) is monitored, and high error rate may mean resource shortage or service itself problem, as a supplementary condition for scaling or alerting.
[0091] As a further embodiment, the following steps are also included:
[0092] S502: In step S5, when Desired Replicas > Current Replicas, scaling is triggered, such as QPS increment of 10 (e.g. 15 -> 25), and 1 container instance is scaled. A more general approach is to directly set Desired Replicas, and set a scaling cooling period, for example 2-5 minutes, to prevent aggressive scaling caused by short-term fluctuations.
[0093] When Desired Replicas < Current Replicas, scale-in is triggered, and a longer scale-in cooling period is set, for example 10-15 minutes or even longer, to ensure that traffic has indeed decreased and stabilized, and to avoid premature scale-in causing jitter.
[0094] Max / Min replica constraints are enforced, setting MinReplicas (>=2 for HF) and MaxReplicas for each service, preventing runaway scaling that can exhaust resources.
[0095] Load balancing is performed, with Kubernetes Service (often in conjunction with Ingress Controller) automatically implementing round-robin or other strategies (such as Least Connections) to distribute requests to healthy container instances.
[0096] As a further embodiment, the method further comprises the steps of:
[0097] S602: The method of step S6 further comprises accelerating the startup process:
[0098] When a new request arrives for a low-frequency service, the scheduler first checks the pre-cached pool. If there is a valid cache, the scheduler preferentially schedules a new container instance to the SSD node that owns the cache for that service. Since the required resources are cached locally on the SSD, the container instance startup process skips or greatly shortens the time for pulling the image from the image repository and downloading large models / data from object storage.
[0099] For example, the resource for the medical image model svc_ctscan is pre-cached on the SSD of Node-7. When a request arrives, a new container instance is scheduled to Node-7, and the startup time is reduced from tens of seconds or even minutes to about 8 seconds.
[0100] If there is no valid cache, the container instance is started according to the standard process (which is slower), and after successful startup, the resource caching is triggered. When there is no request, no container instance is running, saving computing resource costs.
[0101] As a further embodiment, the method further comprises the steps of:
[0102] S603: The method of step S6 further comprises cache eviction
[0103] Based on TTL: Triggered when a cached resource has not been accessed for more than 24 hours.
[0104] Based on disk pressure: Global disk usage monitoring. When the overall usage of the disk partition where the pre-cached pool is located is >=65%, active cleaning is triggered.
[0105] Eviction strategy: LRU (Least Recently Used): Preferentially delete the least recently accessed cache resources. This is the most commonly used strategy.
[0106] LFU (Least Frequently Used) / Size-Aware: When the disk pressure is extremely large, the cleaning operation can be triggered by considering the resource size (deleting large volume but less access) or access frequency (deleting the least frequently accessed) background task or monitoring system, deleting selected cache resource files, and updating cache metadata records.
[0107] Embodiment 2, see Figure 4 The application provides an apparatus for supporting elastic scaling capability of intelligent computing center computing power resources, comprising a monitor: integrating Prometheus to collect QPS, Node Exporter to collect disk usage, and calculating 5-minute sliding window mean value in real time;
[0108] Scheduler: built-in dynamic grading module (FPGA implementation threshold comparison), pre-cache management module (LRU controller + checksum generator);
[0109] Router: based on service ID to query resource cache location, route request to SSD node, update node mapping table when fault switching;
[0110] Three-level cache system: memory (LRU cache hotspot model), local SSD (24-hour access resource), distributed S3 (cold data archiving).
[0111] The scheduler further comprises:
[0112] Service queue management unit: maintaining HF / MF / LF queue, recording service ID and node mapping relationship;
[0113] Fault tolerance processing unit: when the instance crashes, restart and update the route within 15 seconds on the same availability zone node.
[0114] The processor and the memory, the memory stores program code, when executing, realizing any of the above-mentioned methods, wherein:
[0115] FPGA chip is adopted to accelerate QPS threshold comparison operation;
[0116] SSD storage pool provides >3GB / s read bandwidth through NVMe protocol.
[0117] Although the embodiments of the application have been shown and described, it can be understood by those of ordinary skill in the art that various changes, modifications, replacements and variations can be made to these embodiments without departing from the principles and spirits of the application, and all are within the scope of the application.
Claims
1. A method for supporting elasticity and scalability of computing power resources in an intelligent computing center, characterized in that, Comprising the following steps: S1: Real-time monitoring of QPS indicators of all services deployed on the Kubernetes cluster, and collecting raw QPS data through Prometheus, and performing QPS aggregation calculation for each service ID; S2: Calculate the average value of the original QPS data in a sliding time window to smooth the instantaneous traffic fluctuations, and store the processed indicators in a time series database; S3: Based on the sliding window average QPS calculated in S2, determine the reasonable threshold range of different services based on historical data analysis, including: traffic pattern of typical business cycle, service performance baseline, cost constraint target, and dynamically divide service state level based on preset QPS threshold: low frequency, medium frequency, high frequency; S4: Dynamic enhancement adjustment: the threshold is not completely static, and the threshold is dynamically adjusted according to historical data, prediction model or current cluster overall load, to determine the current load level of each service for decision output; S5: Dynamic expansion and contraction rules based on the current average QPS determined in S2, perform hierarchical resource scheduling according to service level: HF service: deploy ≥2 container instances, dynamically expand and contract based on HPA according to QPS; MF service: keep 1 running instance, trigger upgrade and downgrade according to QPS change; LF service: release running instance, only keep SSD node pre-cache resource, do not run container instance for a long time, but pre-load and cache the key resources required by the running instance to the local storage of the selected node; S6: Perform queue state migration based on the upgrade and downgrade path.
2. The method for supporting elastic scaling capability of computing center computing power resources according to claim 1, characterized in that, Further comprising the following steps: S101: Step S1 also includes verifying the adjustment factor: Perform request delay to determine whether expansion effectively improves performance, or as a trigger condition in addition to QPS; Verify whether resource bottleneck is indeed caused by QPS through container resource utilization, and prevent over-expansion; Monitor error rate as a supplementary condition for expansion or alarm.
3. The method of claim 1, wherein, Further comprising the following steps: S501, in step S5, LF service does not run container instance for a long time, and pre-loads and caches the key resources required by the running instance to the local storage of the selected node; When a low-frequency service receives the first request, the system will start a container instance according to the standard process to trigger caching, and after the container instance successfully runs and processes the request, the system automatically caches the required resources to the SSD of the node where the container instance is located; For services that are known to be used in advance, actively cache, keep the cache, and if there is no request to use the cache after a certain time, mark it as eliminable; periodically scan and clean up caches that exceed TTL and have not been accessed.
4. The method of claim 3, wherein, Expansion and contraction control includes: S502, step S5 sets expansion cooling period and contraction cooling period; perform maximum / minimum replica number constraint, set MinReplicas and MaxReplicas for each service to prevent uncontrolled expansion from exhausting resources; perform load balancing to distribute requests to healthy container instances.
5. The method of claim 1, wherein, Further comprising the following steps: S601: Step S6 also includes an accelerated startup process: When a new request is received by the low-frequency service, the scheduler first checks the pre-cached pool. If there is a valid cache, the scheduler preferentially schedules the new container instance to the SSD node that owns the service cache. Since the required resources are cached locally on the SSD, the container instance startup process skips or greatly shortens the time for pulling images from the image repository and downloading large models / data from object storage. If there is no valid cache, the container instance is started according to the standard process, and resource caching is triggered after successful startup. When there is no request, no container instance is running.
6. The method of claim 5, wherein, Further comprising the following steps: S602: In step S6, a cache eviction policy is also set, including a main eviction policy LRU, which preferentially deletes the longest unused cache resource; An auxiliary policy LFU, which deletes selected cache resource files and updates cache metadata records.
7. A device for supporting elastic scalability of computing resources in an intelligent computing center, characterized in that: A monitor: integrates Prometheus to collect QPS, Node Exporter to collect disk usage, and real-time calculation of sliding window mean; A scheduler: built-in dynamic grading module, pre-cached management module; A router: based on service ID to query resource cache location, route requests to SSD nodes, and update node mapping table when fault switching; A three-level cache system: memory, local SSD, and distributed data archive; A fault-tolerant processing unit: when an instance crashes, restart on the same availability zone node and update the routing; A processor and a memory, the memory storing program code, which when executed implements the method of any one of claims 1-6.
Citation Information
Patent Citations
Elastic scaling method and device for container cloud service
CN115469966A
Cloud collaboration container elastic expansion and contraction method and device and electronic equipment
CN118041787A
Method and device for supporting elastic expansion capability by computing power resource of intelligent computing center
CN120086003A