A method and system for hybrid deployment of deep learning tasks

By introducing components such as task scheduling optimizer, SLO manager, Koordlet component and traffic security monitor into the deep learning task hybrid deployment system, the problems of low resource utilization and insufficient balance of task isolation and sharing in the existing technology are solved, and efficient hybrid deployment of deep learning tasks and stable and efficient operation of the system are achieved.

CN119987974BActive Publication Date: 2025-06-17HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510445993.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-17
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

When faced with complex task loads and diversified hardware environments, existing deep learning task hybrid deployment technology has problems such as low resource utilization, insufficient balance of task isolation and sharing, and insufficient adaptability of system scalability and dynamic environments.

Method used

A hybrid deployment method and system for deep learning tasks is proposed. Tasks are submitted through Kubernetes native interfaces, and components such as task scheduling optimizer, SLO manager, Koordlet components and traffic security monitor are used to realize dynamic resource scheduling, task priority management, resource sharing and security monitoring, ensuring the service level of online tasks and the efficient utilization of data center resources.

Benefits of technology

Through dynamic scheduling and optimization strategies, resource utilization and task throughput are improved, ensuring the stability and efficient operation of the system under high load conditions, and adapting to complex task loads and diversified hardware environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987974B_ABST
    Figure CN119987974B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of computer deep learning, and particularly relates to a method and system for hybrid deployment of deep learning tasks. The method includes: S1. A user submits a task through the native interface of Kubernetes; S2. A task scheduling optimizer analyzes the task according to resource requirements and service levels and allocates it to a suitable node; S3. An SLO manager tracks the real-time data of various performance indicators of the system and compares the real-time data with the pre-set service level objectives; S4. The Koordlet component performs specific resource allocation operations on the target node; S5. A traffic security monitor monitors the data flow between nodes in real time. By analyzing the periodic laws of resource usage of deep learning tasks, the present invention proposes a hybrid deployment strategy for different types of resources, realizes dynamic resource sharing between online and offline tasks, and at the same time ensures that the system can still operate stably and efficiently under high load conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer deep learning, and particularly relates to a method and system for hybrid deployment of deep learning tasks. Background Art

[0002] With the rapid development of deep learning technology, its wide applications in fields such as image recognition, natural language processing, and recommendation systems have continuously increased the demand for computing resources. Facing the significant differences in resource requirements between online real-time inference and offline batch training deep learning tasks, the hybrid deployment method has become a key research direction for optimizing the resource utilization rate of cloud data centers and improving task performance.

[0003] The current hybrid deployment technologies mainly focus on the following aspects:

[0004] 1) Task characteristic analysis and classification: By analyzing the low-latency requirements of online tasks and the high-throughput requirements of offline tasks, existing methods classify based on task computational intensity, communication mode, and latency sensitivity, providing a scheduling basis for hybrid deployment. For example, the real-time requirements of online tasks are preferentially satisfied, while offline tasks are processed using resource gaps.

[0005] 2) Resource scheduling and allocation: Resource scheduling is the core of hybrid deployment. The efficiency of static resource partitioning methods is limited, while dynamic scheduling algorithms adjust resource allocation in real time according to task load, improving resource utilization. Intelligent scheduling methods based on reinforcement learning and deep learning optimize decisions through task characteristics and historical data, further enhancing system performance.

[0006] 3) Coexistence optimization of online and offline tasks: Through task priority adjustment and resource isolation mechanisms, resource competition between tasks is optimized. For example, containerization technology realizes task resource isolation, while dynamic resource concession strategies balance task performance and system throughput.

[0007] 4) Cooperative scheduling of heterogeneous hardware: By matching task characteristics with hardware characteristics such as GPUs and TPUs for allocation, computing efficiency is improved. Lightweight tasks are preferentially allocated to low-power hardware, while computationally intensive tasks rely more on GPUs, achieving dual optimization of performance and energy consumption.

[0008] The current research on deep learning task hybrid deployment methods mainly focuses on four directions: task classification, resource scheduling, coexistence optimization, and heterogeneous hardware cooperation. These technologies have promoted the further popularization of deep learning in large-scale applications by optimizing the resource utilization rate of cloud data centers and task performance from multiple angles. However, facing the increasingly complex future task load scenarios and diverse hardware environments, hybrid deployment technologies still need to conduct more in-depth exploration in terms of refined scheduling, energy consumption optimization, and intelligent decision-making.

[0009] The current hybrid deployment technology has the following technical defects:

[0010] 1) Hybrid deployment data center resource allocation and scheduling scheme:

[0011] Although the existing resource allocation schemes for deep learning task deployment have improved resource utilization to some extent, they also have many deficiencies. The static allocation method, although simple in structure, lacks flexibility and is difficult to cope with the fluctuations of dynamic task loads, easily leading to resource waste or increased task latency. The dynamic allocation method, on the other hand, improves efficiency by adjusting resource configurations in real time, but it has a high computational complexity and is prone to system performance fluctuations due to frequent resource reallocations. In addition, the existing dynamic scheduling algorithms lack comprehensive consideration of task priorities and resource requirements in multi-task mixed load scenarios, making it difficult to guarantee the service level objectives (Service Level Object, SLO) of online tasks. Even more complex scheduling strategies (such as reinforcement learning-driven schemes), although excellent in theoretical performance, have high deployment costs in actual production environments and face the problem of strong sensitivity to tuning parameters.

[0012] 2) Task hybrid deployment system:

[0013] Although the task hybrid deployment system has achieved certain results in practical applications, there are still key challenges. First of all, the current systems generally lack the ability to deeply perceive different task types and allocate resources only through simple scoring rules or static policies, making it difficult to adapt to complex dynamic environments. In addition, many systems have deficiencies in the balance between task isolation and resource sharing. Overemphasizing isolation may lead to resource waste, while excessive sharing will cause resource contention and affect the stability of online tasks. Even relatively advanced systems (such as Google Borg and Tencent Caelus) still have problems of increased task scheduling latency and decreased throughput in large-scale task load environments. Another prominent problem is the insufficient control of containerization technology overhead, especially in scenarios of high-frequency task migrations.

[0014] 3) Optimization strategies for deep learning task service systems:

[0015] The optimization of deep learning task service systems has achieved remarkable results in improving resource utilization and performance, but there are still various limitations. Existing methods mainly focus on latency optimization and throughput improvement, but it is difficult to balance both in the scenario of multi-task concurrency. For example, although InferLine and INFaaS systems perform well in latency control, their scalability is weak in the concurrent deployment of large-scale models. In addition, the efficient utilization of heterogeneous hardware resources such as GPUs remains a challenge. Existing methods mostly rely on preemption and concurrency control, but bottlenecks may be caused by the high utilization rate of hardware resources. Although the dynamic scaling and resource allocation schemes of model services can improve performance, it is difficult to ensure stability in the face of bursty traffic, resulting in a significant increase in task latency and failure rate. At the same time, for the service requirements of diverse deep learning models, existing optimization strategies lack sufficient flexibility and generalization ability, and it is difficult to meet the complex requirements of different task loads and hardware environments.

[0016] Although existing technologies have made significant progress in resource allocation, task scheduling, and system optimization, there is still room for improvement in aspects such as resource utilization efficiency, the balance between task isolation and sharing, system scalability, and adaptability to dynamic environments. Future research needs to further combine intelligent scheduling algorithms, refined resource management, and load-aware technologies to improve the performance and stability of the hybrid deployment method for deep learning tasks.

[0017] Currently, cloud data centers are widely used in the deployment and execution of deep learning tasks. However, the independent deployment mode of online tasks and offline tasks leads to low resource utilization. Especially in the context where deep learning tasks highly rely on GPU resources, the problems of resource allocation and scheduling are particularly prominent. To address this issue, cloud service providers have begun to attempt hybrid deployment strategies for online and offline tasks, aiming to improve the overall resource utilization by allowing offline tasks to utilize the idle resources of online tasks. However, this hybrid mode also brings the risk of violating the service level agreement of online tasks. At the same time, it is still difficult to balance system stability and resource utilization efficiency in the face of bursty traffic and complex computational dependencies. Summary of the Invention

[0018] The present invention provides a method and system for hybrid deployment of deep learning tasks, aiming to optimize resource scheduling and task management, and improve the resource utilization and task throughput of the data center while ensuring the service level of online tasks.

[0019] The present invention provides a method for hybrid deployment of deep learning tasks, including:

[0020] S1. The user submits a task through the Kubernetes native interface;

[0021] S2. The task scheduling optimizer parses the tasks submitted by users, extracts key information such as task type, resource requirements, priority, service level, and data locality, and evaluates the resource consumption and execution duration of the tasks based on historical data and the current system state. Combining the real-time resource usage and load status of the nodes, it uses a multi-dimensional scoring model to evaluate each node and finally determines one or more target nodes suitable for task execution;

[0022] S3. The SLO manager tracks the real-time data of various system performance metrics and compares the real-time data with the pre-set service level targets. If the real-time data of the performance metrics of a certain task deviates, it starts the dynamic resource scheduling process and reallocates resources by calling the optimized scheduling strategy;

[0023] S4. After determining the nodes, the scheduling strategy is sent to the Koordlet component. The Koordlet component performs specific resource allocation operations on the target nodes, continuously monitors the resource usage of the nodes, and feeds back the resource usage and load status of the nodes after task execution to the task scheduling optimizer;

[0024] S5. The traffic security monitor monitors the data flow between nodes in real time. By comparing with the historical normal state data, when abnormal traffic patterns or sudden traffic surges are identified, the detected abnormal information is fed back to the SLO manager in real time.

[0025] As a further improvement of the present invention, in step S2, the execution process of the task scheduling optimizer further includes a rescheduling component optimizing the task distribution process according to the node load status:

[0026] During the task execution process, the rescheduling component dynamically evaluates the current task distribution based on the real-time collected node load status and task performance metrics. If it is found that a node is overloaded or the task running efficiency is low, the rescheduling component triggers the rescheduling process and migrates some tasks to nodes with sufficient resources and low load.

[0027] As a further improvement of the present invention, in step S2, different tasks are divided into service levels, an association between the service level and the task scheduling priority is established, the online inference task is given a high service level, and the offline training task is given a low service level. When there is competition for resources in the task scheduling optimizer, the online inference task with a high service level is preferentially executed.

[0028] As a further improvement of the present invention, in step S2, the service level quality types of the tasks include:

[0029] SYSTEM: In the system process, strict resource limit design is carried out according to the resource usage corresponding to the lowest latency time of the SYSTEM response, so as to ensure that while the system process obtains the necessary resources to ensure low-latency response, it will not occupy too many resources;

[0030] LSE: Develop dedicated resource pools and resource reservation machines to ensure that these tasks can be executed stably and with low latency in a low-interference environment;

[0031] LSR: Introduce stricter resource isolation policies and priority scheduling; the resource isolation policy reserves necessary resources for these LSR tasks, and the priority scheduling ensures that the priority of LSR tasks will always be in a higher priority sequence;

[0032] LS: Adopt an elastic resource sharing mode. The elastic resource sharing mode means that when the system resources are sufficient, more resources are allocated to tasks; when resource contention occurs between tasks or the user's request traffic suddenly surges, the resources occupied by tasks are reduced;

[0033] BE: Adopt a best-effort strategy. The best-effort strategy is to make full use of the idle resources in the cluster on the premise that the resource requirements of high-priority tasks are not affected.

[0034] As a further improvement of the present invention, the optimized scheduling strategy includes the CPUBurst strategy, and the CPUBurst strategy specifically includes:

[0035] a1. Adjustment of CPUBurst time slices based on task history: Dynamically adjust the length of the time slices allocated to tasks in the future according to the historical CPUBurst time of tasks. For tasks with shorter CPUBurst, select to allocate shorter time slices, and for tasks with longer CPUBurst, allocate longer time slices;

[0036] a2. Multi-level feedback queue scheduling strategy: Different tasks are assigned to different queues, and the queue priorities are continuously adjusted according to the task execution situation. Tasks enter the corresponding priority queues according to the initial CPUBurst time and scheduling strategy. Shorter tasks are placed in high-priority queues and use shorter time slices, while longer tasks are assigned to low-priority queues and use longer time slices.

[0037] As a further improvement of the present invention, the optimized scheduling strategy includes a memory over-issuance design strategy, and the memory over-issuance design strategy specifically includes:

[0038] b1. Rational configuration of resource requests and limit values: Configure the resources request and resources limit for each task. For online deep learning tasks, set higher resources request and resources limit. When scheduling online deep learning tasks, preferentially select nodes with abundant resources for deployment; for offline deep learning tasks, set lower resources request and higher resources limit. Offline deep learning tasks utilize more memory resources when the cluster resources are abundant and are preferentially restricted or evicted when resources are scarce.

[0039] b2. Task management based on QoS categories: Classify and manage different tasks in the Kubernetes cluster in combination with QoS categories:

[0040] b21. Resource management for SYSTEM tasks: Set appropriate resources limit for SYSTEM type tasks to ensure that SYSTEM type tasks run within the control range;

[0041] b22. Latency-sensitive exclusive tasks (LSE): Reduce the memory limit of LSE type tasks to release a part of the memory for temporary use by other non-critical tasks;

[0042] b23. Latency-sensitive reserved tasks (LSR): Bind the LSR type tasks to CPU cores, allocate the LSR type tasks in a memory pool with higher priority, and monitor the memory usage pattern of the LSR type tasks;

[0043] b24. Latency-sensitive tasks (LS): Allow LS type tasks to obtain a higher resource utilization right when the cluster resources are abundant and share resources when resources are scarce; occupy the unused memory of other tasks when the load is low and automatically reduce the memory usage when the load increases;

[0044] b25. Best-effort tasks (BE): Do not set specific resource requests and limit values.

[0045] As a further improvement of the present invention, the optimized scheduling strategy includes a GPU time slice sharing strategy, and the GPU time slice sharing strategy specifically includes:

[0046] c1. Partition the GPU using MIG technology, and register each partition as an independent device to the cluster through DevicePlugin;

[0047] c2. When a user submits a GPU task, the scheduler makes a decision based on the task requirements and node resource status. When the Pod is started, container-toolkit is responsible for mapping the specified MIG instance to the container and setting the corresponding environment to ensure that the container can only access the predetermined GPU resources.

[0048] c3. Through continuous monitoring and closed-loop feedback mechanisms, the graphics memory allocation is dynamically adjusted according to the real-time load, and GPU resources are efficiently shared in hybrid deployment scenarios.

[0049] As a further improvement of the present invention, the optimization scheduling strategy includes a multi-resource-aware scheduling strategy, and the multi-resource-aware scheduling strategy specifically includes:

[0050] d1. Input online reasoning task set T online , offline training task set T office , node resource status R i ( i =1,…, N ), task priority function P ( t j );in, t j Indicates j tasks;

[0051] d2. Initialization and monitoring: Initialize all nodes R i =( C i , M i , G i ), monitor the resource usage of all nodes in real time and calculate the load status; C i Indicates i The CPU resources of each node, M i Indicates i The memory resources of each node, G i Indicates i GPU resources of each node;

[0052] d3.Task classification and priority allocation: Combine tasks T online ∪ T office ,according to P ( t j)Sorted by priority;

[0053] d4. Resource allocation and scheduling:

[0054] For online inference tasks t j ∈ T online , find available nodes i , meeting the resource requirements , , , if a suitable node is found, the task will be scheduled to that node and the resource status of the node will be updated. A t j = n i , allocate the task t j to the node n i , update the node resources , if the node resources are insufficient, wait for the resources to be released; where represents the CPU resource requirement of the j th task; represents the memory resource of the j th task; represents the GPU resource of the j th task; n i represents the i th node; A t j represents the node allocation situation for the j th task;

[0055] For offline training tasks, t j ∈ T office , if the node has sufficient resources, schedule the task t j to a node that meets the constraints and update the node resources , when the node resources are insufficient, introduce a dynamic adjustment mechanism to calculate ∆ R Recover resources from low-priority tasks and update , , reallocate the resources to online inference tasks; where ∆ R represents the change in node resources, represents the current resources of the node where the task j is located, represents the current resources of the node where the task l is located.​​ Indicates the task k The current resources of the node where it is located;

[0056] d5. Dynamic adjustment and iterative optimization: Real-time monitor resource utilization and task completion status, dynamically adjust resource allocation according to node load, optimize the scheduling strategy, and finally output the mapping A of the task deployment plan to the node as the task scheduling plan.

[0057] The present invention also provides a deep learning task hybrid deployment system, including:

[0058] A task scheduling optimizer for parsing the tasks submitted by users, extracting key information such as task type, resource requirements, priority, service level, and data locality, evaluating the resource consumption and execution duration of tasks according to historical data and the current system state, combining the real-time resource usage and load status of nodes, and using a multi-dimensional scoring model to evaluate each node, and finally determining one or more target nodes suitable for task operation;

[0059] An SLO manager for tracking the real-time data of various system performance indicators, comparing the real-time data with the pre-set service level objectives, and if the real-time data of the performance indicators of a certain task deviates, starting the dynamic resource scheduling process and reallocating resources by calling the optimized scheduling strategy;

[0060] The Koordlet component is used to perform specific resource allocation operations issued by the scheduling strategy on the target node, continuously monitor the node resource usage, and feedback the node resource usage and load status after task execution to the task scheduling optimizer;

[0061] A traffic security monitor for real-time monitoring of the data flow between nodes. By comparing with historical normal state data, when an abnormal traffic pattern or sudden traffic surge is identified, the detected abnormal information will be fed back to the SLO manager in real time.

[0062] As a further improvement of the present invention, the SLO manager includes: a hybrid configuration component for users to perform simple hybrid configuration; a resource over-issuance management component for dynamically adjusting the resource over-issuance ratio according to the real-time state of nodes; a workload statistics component for analyzing and predicting resource requirements using histograms and optimizing the load distribution;

[0063] The task scheduling optimizer includes: a hybrid scheduling component for balancing node loads according to the load awareness strategy, providing fine-grained resource isolation strategies for different tasks, supporting dynamic allocation of heterogeneous resources, and a rescheduling component for dynamically evaluating the current task distribution according to the real-time collected node load status and task performance indicators to re-adjust the resource allocation strategy.

[0064] The beneficial effects of the present invention are as follows: By analyzing the periodic patterns of resource usage in deep learning tasks, a hybrid deployment strategy for different types of resources is proposed to achieve dynamic resource sharing between online and offline tasks, while ensuring that the system can still operate stably and efficiently under high-load conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is an architecture diagram of the hybrid deployment system for deep learning tasks of the present invention;

[0066] Figure 2 is a strategy diagram for dividing the total resource requirements of the present invention;

[0067] Figure 3 is a schematic diagram of the principle of the CPU Burst strategy of the present invention;

[0068] Figure 4 is a schematic diagram of the memory over-issuance allocation scheme of the present invention;

[0069] Figure 5 is a schematic diagram of the GPU sharing scheme of the present invention;

[0070] Figure 6 is an experimental result diagram of the utilization rate of each resource of the system when different types of deep learning training tasks and hybrid tasks are performed in the present invention;

[0071] Figure 7 is a result diagram of the throughput of the ResNet50 inference task within one hour in the present invention;

[0072] Figure 8 is a result diagram of the performance of the ResNet50 task in hybrid deployment under different loads in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0074] I. Method and System for Hybrid Deployment of Deep Learning Tasks Based on Kubernetes.

[0075] 1.1 Description of the method and system for hybrid deployment of deep learning tasks:

[0076] As Figure 1 shown, a hybrid deployment system for deep learning tasks of the present invention aims to achieve efficient collaborative deployment of online tasks and offline deep learning tasks by expanding the native functions of Kubernetes, so as to maximize resource utilization and service quality.

[0077] The core objectives of the system include the following aspects:

[0078] High reusability: Make the most of the native features of Kubernetes as much as possible to avoid making a large number of modifications to the existing platform;

[0079] Scalability: Adopt a modular design to support flexible expansion and maintenance;

[0080] Resource optimization: Optimize the use of key resources such as CPU, GPU, and memory through dynamic scheduling strategies;

[0081] Service quality assurance: Through the management of Service Level Objectives (SLO), ensure the performance requirements of online tasks and effectively isolate offline tasks to avoid resource contention.

[0082] The system architecture consists of native components and extended components of Kubernetes. The native components are responsible for the basic functions of the cluster. For example, Kubernetes-APIServer on the master node is responsible for managing the status of cluster nodes, the Kubernetes scheduler is responsible for task scheduling control, and Kubelet on the worker node is responsible for managing the status of the worker node; the extended components include the node center control component (i.e., the corresponding SLO manager), the traffic security monitor, Koordlet on the worker node is responsible for managing the status of the worker node, and other partially customized components are used to implement the advanced functions of the system.

[0083] Specifically, the deep learning task hybrid deployment system includes:

[0084] Task scheduling optimizer, which is used to parse the tasks submitted by users, extract key information such as task type, resource requirements, priority, service level, and data locality, evaluate the resource consumption and execution duration of tasks based on historical data and the current system status, combine the real-time resource usage and load status of nodes, and use a multi-dimensional scoring model to evaluate each node, and finally determine one or more target nodes suitable for task execution;

[0085] SLO manager, which is used to track the real-time data of various performance indicators of the system and compare the real-time data with the pre-set service level objectives. If the real-time data of the performance indicators of a certain task deviates, start the dynamic resource scheduling process and reallocate resources by calling the optimized scheduling strategy;

[0086] Koordlet component, which is used to execute the specific resource allocation operations issued by the scheduling strategy on the target node, continuously monitor the resource usage of the node, and feedback the resource usage and load status of the node after task execution to the task scheduling optimizer;

[0087] A traffic security monitor is used to monitor the data flow between nodes in real time. By comparing with historical normal state data, when an abnormal traffic pattern or a sudden surge in traffic is identified, the detected abnormal information will be fed back to the SLO manager in real time;

[0088] Task scheduling optimizer: Design a refined scheduling plan according to task characteristics and dynamically adjust the task allocation strategy; Traffic security monitor: Monitor the system traffic status in real time, detect abnormal traffic patterns and prevent potential attacks; SLO manager: Dynamically adjust the resource allocation strategy according to the service level requirements of tasks; Container runtime trigger in the Koordlet component, i.e., the resource isolation module: Ensure the coordinated operation of online tasks and offline tasks through fine-grained resource isolation technology.

[0089] The node central control component is deployed in the cluster in the form of a Deployment, consisting of a main instance (leader) and a backup instance (backup). Its main functions include: Hybrid deployment configuration component: Support hybrid deployment configuration, and users can achieve efficient integration of tasks through simple configuration; Resource overcommitment management component: Dynamically adjust the resource overcommitment ratio according to the real-time status of nodes to maintain service stability; Workload statistics component: Use histogram analysis and prediction of resource requirements to optimize load distribution.

[0090] The hybrid deployment scheduler component is used for task scheduling and resource isolation on Kubernetes, enhancing the functions of the Kubernetes native scheduler and supporting the following capabilities: QoS-aware scheduling: Balance node loads according to the load-aware strategy; Differentiated SLO: Provide fine-grained resource isolation strategies for different tasks; Elastic resource management: Support dynamic allocation of heterogeneous resources to improve system throughput. The hybrid deployment scheduler also provides resource reservation and node fragmentation reorganization functions to further improve the scheduling efficiency.

[0091] The rescheduling component optimizes the task distribution through a load-aware scheduling framework to avoid node hotspots. This component can dynamically adjust the task distribution of nodes to improve the stability and performance of the system.

[0092] The Koordlet component is deployed as a DaemonSet and is used to support resource management under hybrid deployment, including the following functions: Container runtime trigger: Estimate the resource usage of Pods in real time, support resource overcommitment and lifecycle monitoring; QoS management controller: Dynamically adjust the node resource allocation strategy to suppress interference that may affect service quality; Node storage and resource optimization: Provide optimized management for CPU, memory, and video memory.

[0093] Load balancing is achieved through the collaboration of the task scheduling optimizer and the Koordlet component. The scheduler optimizes task allocation globally, while the Koordlet component executes resource allocation policies within the node, ensuring the efficient operation of the system through a multi-level scheduling mechanism.

[0094] Based on this deep learning task hybrid deployment system, the present invention also provides a deep learning task hybrid deployment method, including:

[0095] S1. The user submits a task through the Kubernetes native interface;

[0096] S2. The task scheduling optimizer parses the task submitted by the user, extracts key information such as task type, resource requirements, priority, service level, and data locality, and evaluates the resource consumption and execution duration of the task based on historical data and the current system state. Combining the real-time resource usage and load status of the nodes, a multi-dimensional scoring model is used to evaluate each node, and finally one or more target nodes suitable for task execution are determined;

[0097] The rescheduling component optimizes the task distribution process according to the node load status: during the task execution process, the rescheduling component dynamically evaluates the current task distribution based on the real-time collected node load status and task performance metrics. If it is found that a node is overloaded or the task running efficiency is low, the rescheduling component triggers a rescheduling process and migrates some tasks to nodes with sufficient resources and low load;

[0098] S3. The SLO manager tracks the real-time data of various system performance metrics and compares the real-time data with the pre-set service level objectives. If the real-time data of the performance metrics of a certain task deviates, a dynamic resource scheduling process is started, and the resources are reallocated by calling the optimized scheduling strategy;

[0099] S4. After determining the node, the scheduling policy is sent to the Koordlet component. The Koordlet component performs specific resource allocation operations on the target node, continuously monitors the node resource usage, and feeds back the node resource usage and load status after task execution to the task scheduling optimizer;

[0100] S5. The traffic security monitor monitors the data flow between each node in real time. By comparing with the historical normal state data, when an abnormal traffic pattern or sudden traffic surge is identified, the detected abnormal information is fed back to the SLO manager in real time.

[0101] The task scheduling optimizer aims to achieve efficient scheduling and optimize resource utilization in the hybrid deployment environment of offline deep learning tasks. First, it parses the tasks submitted by users, extracts key information such as task type, resource requirements, priority, service level requirements, and data locality, and constructs a task feature vector. By leveraging historical task data, the system can estimate the execution duration and resource consumption of tasks, providing a basis for task classification and stratification. Meanwhile, it identifies the dependencies between tasks to ensure a reasonable scheduling order and avoid resource conflicts.

[0102] On this basis, the system collects the resource information and health status of each node in real time. Through monitoring tools such as Kubernetes API and Prometheus, it obtains the usage of CPU, memory, GPU, storage, and network bandwidth of each node, and combines the load level and historical stability of the node to mark or temporarily exclude abnormal nodes. The resource capacity model of each node not only records the total resources and idle resources, but also includes data locality information, enabling tasks to be scheduled to nodes as close as possible to the location of the required data to reduce data transmission latency.

[0103] The task scheduling decision is made by constructing a matching matrix between task requirements and node resource capabilities. First, it filters out the nodes that meet the minimum resource requirements, and then comprehensively evaluates these nodes using a multi-dimensional scoring mechanism. This scoring mechanism considers multiple factors such as resource matching degree, node load, data locality, and SLO satisfaction. By setting a weighted scoring function to quantify each indicator, the node with the highest score is finally selected as the candidate target. According to the priority and urgency of tasks, the system will also adopt reservation, preemption, or task migration strategies in scenarios of resource shortage or high concurrency to ensure that the optimal scheduling decision can be implemented on specific nodes.

[0104] In terms of resource reservation, the scheduling optimizer reserves the required resources after determining the target node through the scheduling extension of Kubernetes or the custom scheduler interface, preventing the resources from being occupied by other tasks during the scheduling process. After the task scheduling decision result is passed to Kubernetes, the system can successfully start the task on the reserved resources and monitor the startup and execution of the task in real time. After the task execution ends, the system collects performance data including execution duration, resource utilization rate, and task success rate, providing feedback for the optimization of subsequent scheduling strategies and resource allocation models.

[0105] This solution also constructs a closed-loop feedback mechanism. By feeding back the data after task execution to the scheduling optimizer, it realizes the dynamic adjustment of future scheduling strategies. For long-running tasks or sudden load changes, the system can evaluate the task status in real time and trigger task migration or resource reallocation when necessary to ensure the achievement of service level objectives. Furthermore, the system introduces a machine learning model to predict task execution time, resource consumption, and node load through historical scheduling data, thereby optimizing scheduling decisions in advance. At the same time, it adopts reinforcement learning algorithms to continuously self-adjust during the scheduling process, forming an adaptive intelligent optimization mechanism. To balance resource utilization, task execution efficiency, and SLO satisfaction, the system also constructs a multi-objective optimization model and combines heuristic search, greedy algorithms, or genetic algorithms to quickly solve the optimal or near-optimal scheduling solution under resource constraints.

[0106] The traffic security monitor and the SLO manager play dual roles of real-time security protection and performance guarantee in the system, and they work together to dynamically adjust the resource allocation strategy. After the system starts, the traffic security monitor begins to monitor the data flow between nodes in real time. By comparing with historical normal state data, it identifies any abnormal traffic patterns or sudden traffic surges. This process relies on preset security rules and advanced data analysis algorithms. Once potential attack risks or abnormal behaviors are detected, the monitor will immediately trigger resource adjustment strategies, such as reducing the resource supply to the affected nodes, or redirecting some traffic to a dedicated isolation environment for in-depth detection and processing, thereby effectively reducing the spread of security risks.

[0107] At the same time, the SLO manager continuously tracks various performance indicators of the system, such as response time, throughput, and latency, and compares these real-time data with the pre-set service level objectives. If the performance indicators of a certain service or task deviate, the SLO manager will immediately initiate the dynamic resource scheduling process and reallocate resources by calling the internal scheduling optimization module of the system. For example, when it is detected that the response time of a certain task exceeds the standard, the system may automatically increase the CPU or memory resources for this task, or adjust the task priority to ensure that critical tasks are prioritized in resource competition. This process not only relies on real-time monitoring data but also uses historical operation data and prediction models to anticipate resource requirements, thus achieving precise scheduling.

[0108] The entire adjustment mechanism forms a closed-loop feedback system: the abnormal information detected by the traffic security monitor is fed back to the SLO manager in real time. The latter combines the performance deviation data to initiate a resource adjustment strategy and notifies the scheduling optimizer to reconfigure tasks and nodes. After the adjustment is completed, the system continuously monitors the resource allocation and task running status to ensure that the security protection measures and service quality requirements are always in balance, and further iteratively optimize when necessary. Through this specific implementation plan, the system can quickly respond in the face of sudden security incidents and performance bottlenecks, dynamically balance the security and performance requirements, and ultimately ensure the stable and efficient operation of the overall service.

[0109] During the entire system operation, first, the user submits a task through the Kubernetes native interface. At this time, the task contains resource requirements, service level objectives, and other task metadata. After receiving the task, the task scheduling optimizer will first parse and preprocess the task, extract the key attributes, and evaluate the resource consumption and priority of the task based on historical data and the current system state. Based on this information, the scheduler combines the real-time resource status and load status of the nodes, and uses a multi-dimensional scoring model to evaluate each node, and finally determines one or more target nodes that are most suitable for the task to run.

[0110] After determining the nodes, the scheduling decision is sent to the Koordlet component. Koordlet is responsible for performing specific resource allocation operations on the target nodes. It pre-reserves the required resources, starts the task containers, and ensures the smooth operation of the tasks in the allocated resource environment by deeply integrating with the Kubernetes API. At the same time, Koordlet also continuously monitors the resource usage of the nodes, such as CPU, memory, network bandwidth, and storage, etc., in order to timely detect resource bottlenecks or abnormal loads during the task running process.

[0111] At the same time, the traffic security monitor inside the system monitors the network traffic and data exchange status of the tasks in real time with millisecond-level precision. By comparing with the preset security rules and historical traffic models, once abnormal traffic or potential attack signs are detected, it will quickly issue an alarm signal and work together with the SLO management module to adjust the traffic of the tasks or nodes, such as restricting traffic, isolating abnormal nodes, or guiding the traffic to the security inspection area, so as to ensure the overall security and stability of the system.

[0112] During the task execution process, the system also has a built-in rescheduling component. It dynamically evaluates the current task distribution based on the node load data and task performance metrics collected in real time. If it is found that some nodes are overloaded or the task running efficiency is low, the rescheduling component will trigger the rescheduling process and migrate some tasks to nodes with sufficient resources and lower load. This dynamic rescheduling mechanism can not only relieve the pressure on a single node in a timely manner but also ensure that the system is always in the best operating state and the service level objectives are continuously met.

[0113] The entire process forms a closed-loop feedback system. The tasks are continuously iteratively optimized in submission, scheduling, resource allocation, real-time monitoring, and dynamic rescheduling, ensuring that the system can quickly respond to sudden traffic changes and resource bottlenecks and continuously improve the overall performance and security.

[0114] 1.2 Service level division under the hybrid system:

[0115] In a hybrid deployment system, the resource requirement analysis and service level division of online and offline deep learning tasks are the basis for achieving efficient scheduling and optimal resource utilization. In this environment, online inference tasks and offline training tasks often need to run simultaneously, but their resource requirement characteristics and service requirements are very different. Online inference tasks are usually highly sensitive to latency, so low latency and high response speed must be prioritized to meet the performance requirements of users; while offline training tasks focus more on overall computational efficiency and resource utilization and can tolerate a certain degree of scheduling latency and resource competition.

[0116] For a hybrid deployment system, different tasks need to be divided into service levels, and the reasonable allocation and use of resources are achieved through the association between service levels and task scheduling priorities. Online inference tasks are usually given a higher service level because their latency directly affects the user experience and the overall performance of the system. Offline training tasks may be at a lower service level so that they can obtain more resources when the cluster resources are sufficient and actively release resources when resources are scarce to prioritize the needs of online inference tasks. By defining the service levels of tasks, it can be ensured that the resource scheduler gives priority to meeting the needs of critical tasks in the face of resource competition, thereby improving the overall system QoS (Quality of Service). At the same time, the division of service levels also provides clear priority guidance for the scheduling strategy, enabling the scheduler to make dynamic adjustments according to the importance and urgency of different tasks.

[0117] Through a detailed discussion of the resource requirements and service level division of offline deep learning tasks, theoretical support is provided for the scheduling mechanism of the entire hybrid deployment platform. Such analysis will provide a data basis for formulating efficient resource scheduling solutions, so that the response time and latency of online reasoning tasks can be effectively controlled, while also ensuring that offline training tasks obtain as many resources as possible without affecting online services, achieving the optimal configuration of overall performance. This in-depth understanding of resource requirements and service levels can help the hybrid deployment platform maximize resource utilization and fine-tune task scheduling while ensuring service quality.

[0118] In the present invention, the total amount of resources is divided into four indicator lines according to the proportion: limit, usage, short-term reservation, and long-term reservation. The basic idea is to use those allocated but unused resources to run low-priority pods, such as Figure 2 shown.

[0119] Limit in the figure: The gray line indicates the amount of resources requested by the high-priority Pod, that is, the requested amount of resources for the highest-priority inference deep learning task, which corresponds to the native Burstable type Pod request of Kubernetes.

[0120] Usage: The red line indicates the actual amount of resources used by the Pod. The horizontal axis is the timeline, and the red line is the fluctuation curve of the Pod load over time.

[0121] Short-term reservation: The dark blue line is an estimate of resource usage in the future based on resource usage in the past (short) period. The difference between reservation and limit is that the allocated unused (resources that will not be used in the future) can be used to run batch Pods for short-term execution.

[0122] Long-termreservation: The light blue line is similar to short-termreservation, but the estimated historical usage period is longer. Resources from reserved to limited can be used for Pods with longer life cycles, and compared with short-term forecast values, there are fewer available resources, but they are more stable.

[0123] At the same time, since the granularity of Kubernetes' native service level quality (Quality of Service, QoS) division is too large, it cannot meet the fine-grained task priority differentiation in the co-location scenario. The present invention redesigns five new QoS types supported by the scheduling system in the co-location system.

[0124] (1) SYSTEM: Tasks of the SYSTEM type are system processes with restricted resource usage. For example, system services such as DaemonSets. Although it is necessary to ensure the latency of system services, it is also necessary to limit the resource usage of these system service containers on the node to ensure that they do not occupy excessive resources. For SYSTEM type tasks, a strict resource limit design is adopted. It is mainly based on the required resource usage corresponding to the minimum latency time of its response in the system process for strict resource limit design, ensuring that system processes (such as DaemonSets) can obtain the necessary resources to ensure latency response while not occupying excessive resources, thus guaranteeing the stable operation of the entire system.

[0125] (2) LSE (LatencySensitiveExclusive): LSE tasks require reserved resources and organize pods with the same QoS to share resources. It is generally used for deep learning inference tasks with special latency response requirements and is used in an independent resource pool, corresponding to a part of the Guaranteed type tasks in Kubernetes natively. The core design of LSE tasks lies in resource exclusivity. Through dedicated resource pools and resource reservation mechanisms (such as CPU binding and exclusive memory allocation), it is ensured that these tasks can obtain stable and low-latency execution in a low-interference environment.

[0126] (3) LSR (LatencySensitiveReserved): LSR tasks need to reserve resources to obtain better accuracy. Strategies such as CPU core binding are adopted to ensure execution efficiency, corresponding to a part of the Guaranteed type tasks in Kubernetes natively. LSR tasks further emphasize the accuracy and stability of task execution on this basis. Not only are necessary resources reserved, but more strict resource isolation policies and priority scheduling are also introduced in the scheduling process. The more strict resource isolation policy means that the system will reserve some resources for these LSR tasks, and these resources will not be occupied by other tasks. In this way, when LSR tasks need them, they can quickly use these resources; the more strict priority scheduling means that the priority of these LSR tasks will always be in a higher priority sequence, ensuring their priority execution, thus achieving double guarantee of the task execution quality while maintaining low-latency response.

[0127] (4) LS (Latency Sensitive): LS tasks require shared resources to ensure better resilience to burst traffic. It corresponds to deep learning inference workloads relying on microservices, thus achieving better resource resilience and more flexible resource adjustment capabilities, corresponding to the Guaranteed and Burstable types of tasks that are part of Kubernetes natively. For LS tasks that share resources, an elastic resource sharing mode is designed. The elastic resource sharing mode is that when system resources are sufficient, more resources will be allocated to tasks; when resource contention occurs among tasks, resulting in performance degradation, or when the user's request traffic suddenly surges, the occupied resources of tasks will be reduced to reduce resource contention.

[0128] In this mode, LS tasks can share resources with other tasks during normal operation, but when the system monitors burst traffic or performance degradation, the scheduling system can quickly adjust resource allocation and dynamically expand the resource supply of LS tasks to meet temporary high-concurrency requirements. This design relies on real-time monitoring data and a closed-loop feedback mechanism, enabling precise control of resources in a co-location scenario.

[0129] (5) BE (Best Effort): The resource operation quality of BE tasks is limited and may even be killed in extreme cases. It is the typical QoS level of batch jobs, with stable computing throughput within a certain period, low-cost resources, corresponding to the BestEffort type of tasks in Kubernetes natively. Regarding BE tasks, a best-effort strategy is adopted. On the premise of ensuring that the resource requirements of high-priority tasks are not affected, the idle resources in the cluster are fully utilized. Through dynamic detection and automatic recycling mechanisms, the system allows BE tasks to run when resources are abundant, but gives priority to reservation and migration when resources are scarce, and even allows them to be terminated in extreme cases to maintain the balance of overall resource scheduling.

[0130] Overall, this design scheme constructs a fine-grained and multi-dimensional resource management model. From task metadata extension, scheduling policy adjustment to real-time monitoring and dynamic feedback, it closely focuses on the characteristics of different QoS types. Through deep integration with the Koordlet component, the system can flexibly adopt dedicated, shared, or reserved resource allocation strategies according to the different requirements of each QoS type during task scheduling, making full use of the idle resources in the cluster and ensuring resource guarantee for high-priority tasks, thus achieving the efficiency and flexibility of resource scheduling in a co-location environment.

[0131] Through redesigning the Quality of Service (QoS) types and resource partitioning strategies, the present invention effectively realizes the coordinated scheduling of tasks with different priorities in a hybrid deployment environment through the Koordlet component. By making full use of the allocated but unused resources, the system not only ensures the resource requirements of high-priority tasks but also improves the resource utilization rate of low-priority tasks, thus optimizing the overall resource management and task processing performance of the cluster.

[0132] From the system objective, module design to the operation process, the method and system for hybrid deployment of deep learning tasks realize the efficient collaborative deployment of online tasks and offline deep learning tasks by extending the native functions of Kubernetes, optimize the resource utilization rate, and ensure the service quality. This architecture provides a theoretical basis and practical support for resource management and task scheduling in data centers.

[0133] II. Optimization strategies for hybrid deployment of online and offline deep learning tasks.

[0134] In the hybrid deployment of online and offline deep learning tasks, the core of optimization is to redesign the resource sharing scheme to achieve the efficient hybrid deployment of different types of tasks. Deep learning tasks are characterized by being computationally and memory-intensive, which makes reasonable scheduling and allocation particularly important in the scenario of shared resources. Therefore, in order to improve the overall utilization rate of resources, it is necessary to design unique sharing strategies for resources in different dimensions such as CPU, memory, and GPU to achieve flexible, elastic, and efficient resource scheduling and hybrid deployment effects. The current completed research has mainly optimized from the perspectives of CPU and memory and verified its effects through experiments, while significant progress has been made in the theoretical construction of the GPU resource sharing scheme. Next, the CPUBurst strategy, memory oversubscription design, GPU time slice sharing strategy, and the hybrid deployment algorithm of online and offline deep learning tasks based on multi-resource awareness will be introduced respectively.

[0135] 2.1 CPU sharing scheme: CPUBurst strategy

[0136] In high-concurrency tasks, the tail latency phenomenon is one of the key issues affecting system performance. Tail latency refers to the phenomenon that in a high-concurrency environment, the response time of some requests is significantly higher than the average response time. Usually, the 99% or 99.9% request latency is concerned, that is, the response time of the slowest 1% or 0.1% of the requests. From the perspective of user experience, the existence of tail latency means that some users will face long waits. Especially in large-scale online services and real-time applications, this long-tail phenomenon may significantly affect the overall user experience and the service quality of the system.

[0137] The present invention designs a CPU Burst strategy to effectively manage the execution of tasks by optimizing the time slice allocation and reduce the tail latency phenomenon. The CPU Burst strategy mainly optimizes resource allocation by dynamically adjusting the length of the time slice and combining task characteristics, thereby improving system performance. It can be specifically divided into the following parts:

[0138] 2.1.1 CPU Burst Time Slice Adjustment Based on Task History:

[0139] To more finely manage the execution time of different tasks, the present invention designs a time slice adjustment strategy based on the CPU Burst history of tasks, as Figure 3 shown. The operating system dynamically adjusts the length of the time slice allocated to a task in the future according to the task's historical CPU Burst time. For tasks with shorter CPU Bursts, this project selects to allocate shorter time slices, so that the tasks can be executed quickly without affecting the system's response efficiency due to excessive waiting time; for tasks with longer CPU Bursts, longer time slices are allocated to avoid the context switching overhead caused by frequent time slice switches.

[0140] In this way, the waiting time of short tasks and the context switching overhead of long tasks can be reduced simultaneously, thereby effectively reducing the tail latency. This strategy helps to reduce the queuing time of short tasks and prevent them from being significantly delayed in response due to waiting for scheduling. At the same time, by reducing the number of context switches, the system's processing overhead is also reduced, enabling more effective utilization of CPU resources.

[0141] 2.1.2 Multilevel Feedback Queue (MLFQ) Scheduling Strategy:

[0142] To further enhance the effect of the CPU Burst strategy, the present invention adopts a multilevel feedback queue (Multilevel Feedback Queue, MLFQ) scheduling method. MLFQ is an adaptable scheduling strategy that can dynamically adjust the priority of tasks according to their CPU Burst characteristics. In MLFQ, different tasks are assigned to different queues, and the queue priorities are continuously adjusted according to the task execution situation:

[0143] Tasks enter the corresponding priority queues according to the initial CPU Burst time and the system scheduling strategy. Shorter tasks are placed in the high-priority queues and use shorter time slices to ensure their rapid completion. Longer tasks are assigned to the low-priority queues and use longer time slices to reduce context switching. Through this multilevel feedback queue mechanism, it can be ensured that shorter tasks are quickly responded to, thereby effectively reducing the overall tail latency; for longer tasks, the overhead caused by excessive switches is avoided by allocating longer time slices.

[0144] In the Kubernetes environment, the multi-level feedback queue (MLFQ) scheduling policy is tightly integrated with the native scheduler and container runtime, forming a set of adaptive task scheduling mechanisms. The entire process starts from task submission. When a user creates a Pod through the Kubernetes native interface, the system not only receives standard resource requests and service level information but also obtains a preliminary CPU Burst estimate through extended task metadata (such as annotations or labels). The scheduler first assigns the Pod to the initial feedback queue based on these estimated information: those tasks with shorter expected CPU Burst are placed in the high-priority queue, while tasks with longer estimated times enter the low-priority queue, thus ensuring that short tasks can obtain a quick response in the initial stage, and long tasks reduce the context switching overhead caused by frequent scheduling through longer time slices.

[0145] When the Pod enters the scheduling queue, the extended scheduler module preferentially selects tasks in the high-priority queue according to the priority order of the multi-level feedback queue. For each selected Pod, the system calculates an appropriate time slice through a custom scheduling algorithm and combines the real-time resource status and health monitoring data of the node to schedule the task to a node with sufficient resources. In this process, the Koordlet component plays a key role. It not only performs specific resource allocation operations but also is responsible for converting the scheduling decision into adjustments to the Cgroups parameters of the underlying Docker containers to ensure that the container obtains CPU resources matching the allocated time slice at the operating system level.

[0146] During the running of the task, its actual CPU usage will be collected and recorded in the historical database in real time. This process depends on the data feedback provided by monitoring tools such as Kubernetes MetricsServer or Prometheus. When the task is completed or reaches the preset time slice upper limit during operation, the system will compare the difference between the actual CPU Burst of the task and the estimated value and use algorithms such as exponential smoothing to update the task's historical record. If the task uses the time slice in its high-priority queue longer than expected, it means that its CPU Burst is longer. At this time, the system will automatically downgrade it to a lower-priority queue; conversely, if the task is quickly completed within the allocated shorter time slice, it may get a priority boost and thus continue to enjoy a higher priority in the next scheduling. Such dynamic adjustment not only ensures the quick response of short tasks but also balances the resource usage of various tasks in the co-location environment and reduces the overall tail latency.

[0147] The entire MLFQ mechanism forms a closed-loop scheduling system through continuous monitoring, feedback, and queue adjustment. In the entire Kubernetes cluster, this mechanism can record the CPU Burst history and current queue status of tasks in the form of structured data with the help of Custom Resource Definitions (CRD), enabling the extended scheduler to read this information in real time and make adjustments. Finally, when tasks optimized by the multi-level feedback queue scheduling strategy run on nodes, they not only obtain resource allocations matching their CPU characteristics but also achieve precise execution of time slice allocation through the underlying resource control of Docker and Cgroups, ensuring the efficiency and flexibility of the overall system scheduling.

[0148] 2.2 Memory sharing scheme: Memory overcommit design.

[0149] In the Kubernetes cluster, each task needs to specify the resources request and resources limit values for memory usage when submitted, representing the minimum memory requirement and the maximum allowed memory amount of the task respectively. The resources request is the main basis for the scheduler to decide which node to schedule the task to, ensuring that the task can obtain the lowest memory resources meeting its requirements after being scheduled to the node. The resources limit is the maximum memory amount that the task is allowed to use during operation and is a means for the system to limit excessive resource usage.

[0150] Although this resource management method based on requests and limits ensures the stable operation of tasks, it will face certain resource waste problems in practice. Specifically, if the memory amount requested by a task is significantly higher than the memory required during actual operation, the cluster scheduler will reserve this part of the memory for the task according to the request value, even if these resources are not actually used. As a result, other tasks will not be able to use these idle memories, ultimately causing resource waste and reducing the overall resource utilization rate of the cluster.

[0151] The present invention designs a memory overcommitment strategy, as Figure 4 shown. Memory overcommit allows the total memory amount allocated to nodes in the cluster to exceed the total physical memory, that is, the so-called "overselling" of memory resources. Through memory overcommit, resources can be managed and optimized more flexibly, maximizing the overall resource utilization rate of the cluster. The implementation of the memory overcommit strategy depends on the flexible configuration of resource request and limit values in Kubernetes and the management of different task priorities in combination with the QoS mechanism of Kubernetes. The specific implementation methods can be divided into the following aspects:

[0152] 2.2.1 Rational Configuration of Resource Requests and Limit Values

[0153] In order to effectively implement memory overcommitment in the cluster, the present invention needs to rationally configure the resources request and resources limit for each task:

[0154] For online deep learning tasks, higher resources request and resources limit should be set to ensure that these tasks can obtain sufficient resources and guarantee the stability and performance of the service. When scheduling these tasks, nodes with relatively abundant resources will be preferentially selected for deployment to reduce potential resource competition.

[0155] For offline deep learning tasks, lower resources request and higher resources limit can be set. These tasks can utilize more memory resources when the cluster resources are abundant, but will be preferentially restricted or evicted when resources are scarce, thus ensuring the normal operation of critical tasks.

[0156] Through reasonable resource configuration, it can be ensured that the cluster can effectively allocate memory under high load conditions, reduce resource waste, and maximize resource utilization.

[0157] 2.2.2 Task Management Based on QoS Categories

[0158] In the Kubernetes cluster, in order to better implement memory overcommitment, custom QoS categories introduced in this project are combined to classify and manage different tasks. The most suitable memory resources can be allocated according to the task characteristics to achieve more flexible and efficient resource management.

[0159] (1) SYSTEM: Resource Management of System Tasks

[0160] In the memory overcommitment design, SYSTEM-type tasks have a higher priority. The system will reserve sufficient memory for these tasks to ensure stable operation, but will limit them to prevent excessive occupation of node resources.

[0161] By setting appropriate resources limit for SYSTEM tasks, the set resource value should be able to support SYSTEM tasks to meet the system's response time, which can ensure that the operation of these services is within control, that is, the response time can meet the requirements of system operation and will not cause other tasks to be unable to run due to resource competition. At the same time, such a design also ensures the reliability of the system's basic services.

[0162] (2) LSE (LatencySensitiveExclusive): Latency-Sensitive Exclusive Task

[0163] Tasks of the LSE (LatencySensitiveExclusive) type are deep learning inference tasks with extremely high latency requirements. Such tasks need to ensure exclusive access to resources, that is, they cannot share resources with other tasks of the same type in the allocated resource pool to ensure resource stability and low-latency performance. Therefore, LSE tasks are generally allocated to independent resource pools for deployment. When applying for resources, LSE tasks require exclusive access to CPU and memory resources and do not allow sharing with other tasks. This resource allocation method ensures task stability and high performance. Especially for real-time inference scenarios, LSE can significantly reduce tail latency.

[0164] Although LSE tasks require exclusive access to resources, since not all allocated memory is usually fully utilized throughout the task execution, the memory oversubscription strategy can appropriately reduce the memory limit of LSE tasks to free up some memory for temporary use by other non-critical tasks.

[0165] (3) LSR (LatencySensitiveReserved): Latency-Sensitive Reserved Task

[0166] Tasks of the LSR (LatencySensitiveReserved) type are mainly for deep learning tasks that require high execution efficiency and resource reservation. These tasks usually use CPU core binding to ensure task execution efficiency. LSR tasks also have very high requirements for memory and CPU resources, especially the need to ensure sufficient resources to obtain higher inference accuracy.

[0167] LSR tasks need to have resources reserved for them. For example, by CPU core pinning to ensure exclusive access to certain cores. This approach can significantly improve task execution efficiency and reduce the performance overhead caused by context switching. At the same time, to ensure that LSR tasks can obtain sufficient memory, LSR tasks are allocated in memory pools with higher priorities, and their memory will not be preempted by other tasks even in the case of memory oversubscription. In addition, the memory usage pattern of LSR tasks is also monitored to avoid excessive reservation during resource tension and affect the overall performance of the cluster.

[0168] (4) LS (LatencySensitive): Latency-Sensitive Task

[0169] LS type (LatencySensitive) tasks also have high latency requirements, but unlike LSE and LSR, the resources of LS tasks can be shared with other tasks within a certain range to achieve better resource elasticity and dynamic adjustment capabilities. This type of task is mainly used for deep learning inference workloads based on microservices. Its goal is to achieve efficient resource sharing and elastic management while ensuring response speed.

[0170] The resource scheduling strategy of LS tasks allows them to obtain higher resource utilization rights when cluster resources are abundant, and to share resources when resources are tight, thereby achieving better elastic management effects. For example, when there are more idle resources in the system, LS tasks can dynamically expand their resource requests to obtain more memory and computing resources. In the design of memory over-issuance, LS tasks can occupy unused memory of other tasks when the load is low. This mechanism provides good elasticity for LS tasks. When the load increases, LS tasks will automatically reduce memory usage and release more memory to LSE or LSR tasks with higher priority, thereby achieving overall memory utilization optimization.

[0171] (5) BE (BestEffort): Best Effort Task

[0172] BE type (BestEffort) tasks correspond to deep learning training tasks or batch processing tasks of some models. These tasks can be killed when resources are tight, and usually have lower requirements for latency and resource guarantees. The goal of BE type tasks is to provide a certain amount of computing throughput at the lowest cost, so they have the lowest priority in the cluster.

[0173] BE tasks do not set specific resource requests and limits, so they are low-priority tasks in memory allocation, and the surplus memory in the memory over-issuance policy will be allocated to BE tasks first. When cluster resources are tight, these tasks will be affected first, including being evicted by the system to ensure that other higher-priority tasks can run stably. Since BE tasks do not have strict requirements on resources, the scheduler can flexibly schedule BE tasks for execution during idle time, thereby improving the overall throughput of the system without affecting critical tasks. These tasks usually use the remaining resources in the cluster at a low cost and can accept being suspended or killed in extreme cases.

[0174] 2.3 GPU Sharing Solution: GPU Time Slice Sharing Strategy

[0175] GPUs are indispensable computing resources in deep learning tasks. Especially in scenarios that require efficient inference and model training, their video memory and computing capabilities directly determine the performance and resource utilization of the tasks. However, due to the fixed video memory resources of GPUs and the large demand for video memory in deep learning tasks, how to efficiently share video memory resources among multiple tasks has become a key challenge in hybrid deployment. To solve this problem, the present invention proposes a GPU video memory sharing strategy based on NVIDIA MIG (Multi-Instance GPU) technology. By combining hardware isolation with Kubernetes scheduling, on-demand allocation of video memory for online inference tasks and offline training tasks is achieved.

[0176] The architecture of the GPU video memory sharing strategy is as Figure 5 shown. The figure shows how to use MIG technology to partition the GPU video memory and register these partitions as independent GPU instances in the Kubernetes cluster. On this basis, the cluster dynamically allocates GPU video memory resources to different tasks through the scheduler. Online inference tasks usually need to obtain video memory resources preferentially to ensure real-time response capabilities; while offline training tasks use the remaining video memory resources for batch computing to improve overall resource utilization. The scheduling system dynamically adjusts the video memory allocation strategy by monitoring the usage of GPU video memory and task loads in real time, thereby achieving efficient sharing of video memory resources among different tasks.

[0177] The MIG technology plays a core role in this strategy. Its principle is to divide the physical GPU video memory into multiple logical partitions, and each partition corresponds to a MIG instance. Each instance has independent computing cores, video memory, caches, and bandwidth resources, ensuring isolation between tasks. Combining with the resource management capabilities of Kubernetes, each MIG instance is exposed as an independent GPU resource type (such as nvidia.com / mig-1g.5gb), allowing users to explicitly request the required video memory partitions in the task description. For example, an online inference task can request a small MIG instance (such as 5GB of video memory), while an offline training task requests a larger MIG instance (such as 20GB of video memory). This video memory management method based on hardware partitioning not only avoids resource contention between tasks but also provides support for fine-grained scheduling of video memory resources.

[0178] In the video memory sharing strategy, the design of the scheduling system is crucial, especially how to dynamically adjust video memory allocation according to task loads. The scheduling system utilizes the extensibility of Kubernetes and introduces a priority queue and a dynamic allocation mechanism to allocate different types of tasks to different queues. For example, online inference tasks are assigned to high-priority queues, while offline training tasks are assigned to low-priority queues. The scheduler dynamically adjusts the allocation of video memory partitions according to the real-time demands of tasks. For example, during the peak period of inference requests, the system preferentially allocates more video memory resources to high-priority queues to ensure low-latency responses for online tasks. During periods with lower request volumes, more video memory resources are allocated to offline training tasks to improve the efficiency of model updates.

[0179] To verify the effectiveness of the video memory sharing strategy, the present invention designs multiple experimental scenarios, including a hybrid deployment scenario of online inference tasks and offline training tasks. In these scenarios, the GPU video memory utilization rate is significantly improved, the response latency of online tasks is reduced, and the training throughput of offline tasks is also guaranteed. Especially in a dynamic load environment, this strategy demonstrates excellent adaptability. By flexibly allocating video memory resources, it effectively balances the system's performance and resource utilization rate.

[0180] The GPU video memory sharing strategy not only provides a new solution for GPU resource management in hybrid deployments but also lays a foundation for multi-tenant GPU resource sharing in cloud computing scenarios. Through the combination of hardware isolation and dynamic scheduling, this strategy achieves a good balance among resource utilization rate, task isolation, and scheduling flexibility, providing strong support for the efficient deployment of deep learning tasks.

[0181] 2.4 Hybrid Deployment Algorithm for Online and Offline Deep Learning Tasks Based on Multi-Resource Awareness.

[0182] Based on the above problems in this chapter, the present invention proposes a hybrid deployment algorithm for online and offline deep learning tasks based on multi- resource awareness. This hybrid deployment algorithm aims to dynamically allocate resources in a Kubernetes cluster to meet the low-latency requirements of online inference tasks and the requirements for efficient completion of offline training tasks while maximizing resource utilization. The algorithm starts with initialization and monitoring, initializing the resource information of all nodes in the cluster, including the capacities of CPU, memory, and GPU. The system will monitor the resource usage of each node in real time and calculate the load status of the nodes to ensure the effective utilization of resources.

[0183] In the task classification and priority assignment phase, the algorithm assigns priorities to tasks according to the task type (online inference task or offline training task), and uses the priority function to sort the tasks to ensure that online tasks have higher priorities in scheduling. Subsequently, the algorithm enters the resource allocation and scheduling phase. For online inference tasks, the system preferentially allocates resources to them and searches for nodes that meet the resource requirements. If a suitable node is found, the task will be scheduled to that node and the resource status of the node will be updated. If the node resources are insufficient, the system will wait temporarily for the resources to be released. For offline training tasks, the system schedules the tasks to suitable nodes under the condition of meeting the resource constraints and updates the resource information of the nodes. When the node resources are insufficient, the algorithm introduces a dynamic adjustment mechanism to reclaim resources from low-priority tasks and reallocate them to online tasks to ensure that the performance requirements of online tasks are preferentially guaranteed.

[0184] The dynamic adjustment and iterative optimization mechanism of the algorithm monitors the resource usage and task completion status of nodes in real time, and adjusts the scheduling strategy in a timely manner when the resource usage is insufficient or the demand changes. This mechanism further optimizes the resource allocation strategy of tasks through elastic resource allocation and capacity borrowing, thereby achieving efficient resource management in a hybrid deployment environment and improving the overall performance of the cluster and the task completion efficiency. The pseudocode of the algorithm is shown below.

[0185] The explanations of some symbols in the following pseudocode are as follows: t j represents the j th task; C i represents the CPU resources of the i th node; M i represents the memory resources of the i th node; G i represents the GPU resources of the i th node; represents the CPU resource requirement of the j th task; represents the memory resources of the j th task; represents the GPU resources of the j th task; n i represents the i th node; A t j represents the node allocation situation of the j th task; ∆ R represents the change in node resources; represents the current resources of the node where the task j is located;​ Represents the task l The current resources of the node where it is located; Represents the task k The current resources of the node where it is located.

[0186]

[0187] According to the algorithm, the time complexity of the algorithm mainly consists of several key steps. First, the operations of initializing and monitoring each node require O(N) time, where O represents the time complexity of the algorithm and N is the total number of nodes. Second, the step of sorting the priorities of all online and offline tasks requires O(TlogT) time, where T is the total number of tasks. The resource allocation and scheduling part involves finding suitable resources for each online task among N nodes, which requires O(T online ×N) time. The scheduling of offline tasks may also include dynamic resource adjustment, which will further increase the time complexity when frequent adjustments are required. Therefore, the overall time complexity is O(TlogT + T online ×N).

[0188] The space complexity mainly consists of two parts. The algorithm needs to store the resource status R i for each node, and store its scheduling status A for each task. The space requirement for this part is O(N + T), where N is the number of nodes and T is the number of tasks, mainly used to store the information of nodes and tasks.

[0189] III. Cluster Deployment Performance Experiment.

[0190] To verify the effectiveness of the CPUBurst policy, memory overcommitment design, and GPU sharing design in a real environment, the present invention conducts a cluster deployment performance experiment on the built Kubernetes platform and realizes scheduling optimization through the above-developed functional components. The main purpose of the experiment is to test the resource scheduling effect of the scheduling policy on different types of deep learning tasks in a hybrid deployment environment, and to evaluate the improvement of the overall performance of the cluster. The performance of the platform is evaluated through the experiment and compared with traditional static scheduling policies and other dynamic scheduling policies to measure its advantages in terms of resource utilization rate, job completion time, scheduling delay, etc.

[0191] In the hybrid cluster deployment performance experiment, the present invention adopts a systematic experimental data collection method to comprehensively evaluate the performance of the cluster and the task processing efficiency. In the experiment, four different types of deep learning tasks are deployed in the cluster, and each task represents a different load situation. To ensure the independence and controllability of the tasks, all tasks are deployed through containerization technology, and Prometheus is used as the cluster monitoring tool for real-time data collection and monitoring.

[0192] First, in terms of resource utilization, the present invention monitors the CPU, memory, storage, and network resource usage of cluster nodes, calculates the resource utilization rate of the cluster under different loads, and evaluates the resource allocation efficiency and load-bearing capacity of the cluster. Second, the task queuing time and average job completion time are recorded through the logs of the cluster scheduling system, specifically including the queuing time from when the task is submitted to when it starts to execute, and the total time required for the task to complete from start to finish. These data help analyze the efficiency of cluster scheduling and the latency of task execution.

[0193] In terms of inference task performance, the present invention evaluates the throughput of the cluster when processing inference tasks by calculating the number of inference requests completed per hour. Finally, to evaluate the service level of the co-located cluster, the present invention conducts a comprehensive analysis of the performance of the cluster under different loads, focusing on key metrics such as the average resource utilization rate, task completion time, throughput per unit time, and task completion rate. These data help the present invention evaluate the stability of the cluster under high loads, the reliability of task scheduling, and the overall performance. All the collected data is statistically analyzed and presented in the form of charts, thus visually showing the performance changes and resource allocation of the cluster under different loads.

[0194] The experiment aims to evaluate the performance of the system designed by the present invention in a hybrid deployment environment and compare it with other existing scheduling strategies. For this purpose, the present invention selects two representative scheduling strategies, namely Lyra and Pollux, as comparison algorithms. The main comparison metrics include the job completion time, GPU utilization rate, job waiting time, etc. The selection of these metrics aims to comprehensively evaluate the performance and advantages of the system proposed by the present invention, especially its flexibility and efficiency in the face of multi-task hybrid deployment and dynamic resource requirements.

[0195] Lyra is an elastic scheduling strategy for deep learning clusters, aiming to improve the throughput and efficiency of the cluster by dynamically adjusting resource allocation. Lyra flexibly adjusts the resource allocation of tasks based on the priority and real-time resource requirements of the tasks. Its core idea is to utilize the elastic resource allocation mechanism to dynamically adjust the GPU allocation according to the computational requirements and progress of the tasks, enabling the cluster resources to be scheduled and utilized more efficiently according to the task requirements. Lyra is mainly oriented towards training tasks and is suitable for computationally intensive deep learning workloads, which can improve the resource utilization rate of the cluster to a certain extent and reduce the task queuing time.

[0196] Pollux is a scheduling strategy that optimizes the throughput of a cluster by jointly adapting the cluster resources and task requirements to maximize the overall performance of the cluster. Pollux combines the flexibility of resource sharing and task scheduling, dynamically allocating resources among multiple tasks to achieve maximum throughput. Different from Lyra, Pollux pays more attention to the overall throughput of the cluster and the long-term efficient utilization of resources, rather than simply optimizing the task completion time. Pollux adopts a mechanism based on task load prediction and dynamic resource allocation, which can further improve the overall efficiency of the system while ensuring the job completion time.

[0197] 3.1 Evaluation of the resource utilization rate of the co-located cluster.

[0198] Analyze the resource utilization performance of the co-located cluster under different load conditions to verify the adaptability and stability of the system of the present invention under different workload conditions. In the experiment, the GPU utilization rate of the co-located cluster in the high-load scenario and the overall resource stability of the cluster were evaluated. These evaluations help to understand whether the scheduling strategy of the co-located system can efficiently allocate and manage resources in a complex hybrid deployment environment. The experiment used multiple evaluation metrics to quantify the resource utilization situation, including the resource utilization rates of CPU, memory, and GPU training tasks, the average system resource utilization rate, and the resource stability under load. The resource utilization rate is defined as the ratio of the resource occupied by tasks over a period of time, which can reflect the efficiency of the scheduling strategy in allocating computing resources. In the experiment, the workloads used included models such as Bert, GNMT-16, ResNet-50, and VGG. These models have different computing characteristics and resource requirements, thus effectively simulating the load situation of the cluster in the production environment. The task deployment strategy of the co-located system was compared with the scheduling strategies of Lyra and Pollux to evaluate their resource utilization performance of nodes under high load. The experimental results are as Figure 6 shown, Figure 6 a) in [figure] is the resource utilization rate of the ResNet50 training task and the co-located system, Figure 6 b) in [figure] is the resource utilization rate of the Bert training task and the co-located system, Figure 6 c) in [figure] is the resource utilization rate of the GNMT training task and the co-located system, Figure 6 d) in [figure] is the resource utilization rate of the VGG training task and the co-located system.

[0199] The results show that, compared with the traditional FIFO baseline strategy, the co-allocation system of the present invention has a 9% increase in GPU utilization rate in high-load scenarios, reaching 81%. In heavy-load and low-load conditions, the co-allocation system of the present invention also exhibits good stability, with the GPU utilization rate maintained between 70% and 75%. Under similar load conditions, the GPU utilization rates of Pollux and Lyra are approximately 69% - 72%. This indicates that the co-allocation system of the present invention can effectively improve the overall resource utilization rate of the cluster, reduce the idle time of resources, and maintain the stable operation of the system through elastic expansion and capacity borrowing strategies.

[0200] 3.2 Performance evaluation of co-allocation cluster tasks.

[0201] A detailed evaluation was conducted on the task performance in the co-allocation cluster, focusing on analyzing the average job completion time (JCT), task queuing time, and latency performance of different task types under different scheduling strategies. To better understand the performance of each scheduling strategy in different task scenarios, the present invention used the median, mean, and 95th percentile as the core statistical metrics. These metrics can comprehensively reflect the performance of the scheduling strategy in terms of average task completion time, tail latency, and system throughput. The experimental results are shown in Table 1.

[0202] Table 1 Queuing time and average job completion time of different-depth tasks in the co-allocation cluster

[0203]

[0204] In the ResNet-50 task, the average JCT of the system proposed in the present invention is 12,875 milliseconds, showing a significant decrease compared with the 20,872 milliseconds of the traditional FIFO baseline. The experiment shows that by using the elastic expansion mechanism and capacity borrowing strategy, the system proposed in the present invention can dynamically adjust the GPU allocation, enabling training tasks to obtain more GPU resources in high-load situations, thus significantly shortening the task completion time. Compared with Lyra and Pollux, the system proposed in the present invention also shows certain advantages. The average JCTs of Lyra and Pollux are 13,653 milliseconds and 12,983 milliseconds respectively, while the average JCT of the system proposed in the present invention is 12,875 milliseconds. Especially in high-concurrency and complex task load scenarios, the system proposed in the present invention can reduce the waiting time and completion time of tasks by adjusting the scheduling strategy in real time and reasonably allocating resources.

[0205] To further analyze the impact of different scheduling strategies on tail latency, the present invention also calculates the 95th percentile of JCT to measure the performance of the system under extreme load conditions. For the ResNet-50 task, the 95th percentile of JCT of the system built by the present invention is 62,093 seconds, which is a significant reduction compared to 85,972 milliseconds of the FIFO baseline. The 95th percentiles of JCT for Lyra and Pollux are 65,392 milliseconds and 63,986 milliseconds respectively, showing the obvious advantage of the system of the present invention in tail latency control. This is because the system of the present invention can flexibly adjust resource allocation according to the priority of the current task and the idle status of system resources during the scheduling process, thus effectively avoiding resource contention under high load conditions. In terms of system throughput, taking the online inference task of the ResNet50 model as an example, the present invention selects the number of online inference tasks completed within 1 hour as the evaluation index, and the method of the present invention shows high stability and throughput capacity under high load. The experimental results are as Figure 7 shown. The overall throughput of the system of the present invention still has a certain advantage compared to Lyra and Pollux in the heavy load scenario. Especially when multiple high-priority tasks arrive simultaneously, the system of the present invention can give priority to ensuring resources for inference tasks, complete jobs quickly, and flexibly borrow resources for subsequent training tasks. This elastic expansion and priority adjustment mechanism enables the system to handle more tasks under high load, thus improving the overall throughput of the system.

[0206] 3.3 Performance evaluation of the service level division of the hybrid cluster.

[0207] This experiment conducted a comprehensive performance evaluation of tasks with four service levels (LSE, LSR, LS, BE) in the hybrid cluster. The experiment used ResNet-50 as a typical deep learning model inference task and combined different model accuracy requirements to distinguish service levels. Specifically, high-precision inference tasks were set as LSE type to meet the requirements of latency sensitivity and resource exclusivity; medium-precision inference tasks were assigned as LSR type to ensure the execution efficiency of tasks through resource reservation strategies; low-precision inference tasks were classified as LS type, allowing resource sharing to enhance elasticity and flexibility. In addition, batch training tasks were used as BE type to execute low-priority computing tasks using the remaining resources in the cluster. Four different load patterns were designed in the experiment to simulate typical usage scenarios in the hybrid cluster. These load patterns include LowLoad, MediumLoad, HighLoad, and BurstLoad. The LowLoad pattern represents a situation where the cluster resources are abundant. At this time, high-priority tasks can be fully allocated resources according to requirements, and low-priority tasks can also efficiently run using idle resources; in the MediumLoad pattern, the resource requirements gradually approach the cluster capacity, and resource competition begins to occur; the HighLoad pattern simulates a scenario where the cluster resources are close to saturation. High-priority tasks still need to ensure performance, while low-priority tasks may face latency or preemption; the BurstLoad pattern focuses on the impact of burst traffic on the system. Sudden high-priority tasks need to quickly obtain resource support, while low-priority tasks may be suspended or interrupted. These four patterns can comprehensively reflect the task performance of the hybrid cluster under different load conditions. The experimental results are as Figure 8 shown, Figure 8 in a) is the average resource utilization rate of ResNet50 in the hybrid under different loads, Figure 8 in b) is the average task latency of ResNet50 in the hybrid under different loads, Figure 8 in c) is the average throughput of ResNet50 in the hybrid under different loads, Figure 8 and in d) is the average task completion rate per unit time of ResNet50 in the hybrid under different loads.

[0208] The experimental results show that tasks with different service levels exhibit significant differences in indicators such as resource utilization, latency, throughput, and task completion rate. In terms of average resource utilization, BE and LS tasks can fully utilize the unallocated idle resources in the cluster, and their utilization rate increases significantly with the increase in load. In contrast, LSE and LSR tasks maintain a relatively stable resource occupancy rate, which reflects the effectiveness of resource exclusivity and reservation strategies for high-priority tasks. In terms of latency performance, the LSE task has the lowest latency in all load scenarios and can effectively meet its strict requirements for real-time performance; the latency of LSR is slightly higher but still stable. In contrast, the latency of LS and BE tasks increases significantly with the increase in load, especially in the scenario of burst traffic. In terms of throughput, LSE and LSR tasks show strong stability in high-load and burst-traffic scenarios, while the throughput of LS and BE decreases significantly with the intensification of resource competition, reflecting their sensitivity to dynamic resource changes. In terms of the task completion rate indicator, LSE and LSR tasks always maintain a completion rate close to 100%, demonstrating their priority status and reliability in the scheduling system. However, the completion rate of BE tasks decreases significantly with the increase in load and can be as low as 60% in high-load and burst-traffic scenarios. Overall, the experimental results verify the effectiveness of the redesigned service level division strategy, which can not only guarantee the performance requirements of high-priority tasks but also improve the overall efficiency of the hybrid deployment cluster by fully utilizing system resources. This result provides strong support for task scheduling optimization in hybrid deployment scenarios.

[0209] The main advantages of the present invention include the following aspects:

[0210] 1) Improve resource utilization:

[0211] By improving the resource sharing and scheduling mechanisms of CPU, memory, and GPU, the idle rate of resources such as CPU, memory, and GPU is significantly reduced, and the overall computing efficiency of the cluster is improved.

[0212] 2) Optimize task execution performance:

[0213] By introducing a load-aware dynamic scheduling mechanism, the task queuing time and completion time are effectively shortened, the system throughput per unit time is increased, and at the same time, the task response time is ensured to meet the service requirements.

[0214] The present invention not only provides theoretical and practical support for the intelligent hybrid deployment of deep learning tasks but also provides important reference value for the sustainable development and operation cost optimization of modern data centers.

[0215] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A hybrid deployment method for deep learning tasks, characterized in that: include: S1. The user submits a task through the Kubernetes native interface; S2. The task scheduling optimizer analyzes the tasks submitted by users, extracts key information such as task type, resource requirements, priority, service level, and data locality, and evaluates the resource consumption and execution time of the tasks based on historical data and the current system status. It uses a multi-dimensional scoring model to evaluate each node based on the real-time resource usage and load status of the nodes, and finally determines one or more target nodes suitable for task operation. S3. The SLO manager tracks the real-time data of various performance indicators of the system and compares the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimal scheduling strategy; S4. After determining the node, the scheduling strategy is sent to the Koordlet component. The Koordlet component performs specific resource allocation operations on the target node, continuously monitors the node resource usage, and feeds back the node resource usage and load status after the task is executed to the task scheduling optimizer; S5. The traffic security monitor monitors the data flow between nodes in real time. By comparing with the historical normal status data, when abnormal traffic patterns or sudden traffic surges are identified, the detected abnormal information will be fed back to the SLO manager in real time; In step S2, the service level quality type of the task includes: SYSTEM: In the system process, strict resource constraints are designed based on the required resource usage corresponding to the minimum latency of the SYSTEM response, to ensure that the system process obtains the necessary resources to ensure the latency response while not occupying too many resources; LSE: Develop dedicated resource pools and resource reservation machines to ensure that these tasks can be executed stably and with low latency in a low-interference environment; LSR: Introduces stricter resource isolation policies and priority scheduling. The resource isolation policy reserves necessary resources for these LSR tasks, and the priority scheduling ensures that the priority of LSR tasks is always in a higher priority sequence. LS: Adopts elastic resource sharing mode. When system resources are sufficient, more resources are allocated to tasks. When resource contention between tasks is detected or user request traffic suddenly surges, the occupied resources of tasks are reduced. BE: Use a best-effort strategy. The best-effort strategy is to make full use of idle resources in the cluster while ensuring that the resource requirements of high-priority tasks are not affected.

2. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: In step S2, the task scheduling optimizer execution process also includes a rescheduling component optimizing the task distribution process according to the node load status: During the task execution process, the rescheduling component dynamically evaluates the current task distribution based on the node load status and task performance indicators collected in real time. If it is found that a node is overloaded or the task is running inefficiently, the rescheduling component triggers the rescheduling process and migrates some tasks to nodes with sufficient resources and lower load.

3. The deep learning task hybrid deployment method according to claim 1, characterized in that: In step S2, service levels are divided for different tasks, and an association between service levels and task scheduling priorities is established. Online reasoning tasks are assigned high service levels, and offline training tasks are assigned low service levels. When there is competition for resources in the task scheduling optimizer, online reasoning tasks with high service levels are executed first.

4. The deep learning task hybrid deployment method according to claim 1, characterized in that: The optimized scheduling strategy includes a CPUBurst strategy, which specifically includes: a1. CPUBurst time slice adjustment based on task history: Dynamically adjust the time slice length allocated to the task in the future according to the historical CPUBurst time of the task. For tasks with shorter CPUBurst, a shorter time slice is allocated, and for tasks with longer CPUBurst, a longer time slice is allocated; a2. Multi-level feedback queue scheduling strategy: Different tasks are assigned to different queues, and the queue priority is continuously adjusted according to the task execution status. Tasks enter the corresponding priority queue according to the initial CPUBurst time and scheduling strategy. Shorter tasks are placed in the high-priority queue and use shorter time slices, while longer tasks are assigned to the low-priority queue and use longer time slices.

5. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: The optimization scheduling strategy includes a memory over-issuance design strategy, and the memory over-issuance design strategy specifically includes: b1. Reasonable configuration of resource request and limit values: Configure resourcesrequest and resourceslimit for each task. For online deep learning tasks, set higher resourcesrequest and resourceslimit. When scheduling, online deep learning tasks will give priority to nodes with abundant resources for deployment. For offline deep learning tasks, set lower resourcesrequest and higher resourceslimit. When cluster resources are abundant, offline deep learning tasks will use more memory resources and will be restricted or evicted first when resources are scarce. b2. Task management based on QoS categories: Different tasks are classified and managed in the Kubernetes cluster based on QoS categories: b21. Resource management of system tasks SYSTEM: Set appropriate resourceslimit for SYSTEM type tasks. The set resource value supports SYSTEM tasks to meet the system response time to ensure that SYSTEM type tasks run within the control range; b22. Delay-sensitive exclusive tasks LSE: Reduce the memory limit of LSE type tasks to release some memory for temporary use by other non-critical tasks; b23. Latency-sensitive reserved tasks (LSR): LSR-type tasks are bound to CPU cores, assigned to memory pools with higher priorities, and the memory usage patterns of LSR-type tasks are monitored; b24. Latency-sensitive tasks LS: Allow LS-type tasks to obtain higher resource utilization rights when cluster resources are abundant, and share resources when resources are scarce; occupy memory not used by other tasks when the load is low, and automatically reduce memory usage when the load increases; b25. Best-effort task BE: No specific resource request and limit values ​​are set.

6. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: The optimization scheduling strategy includes a GPU time slice sharing strategy, and the GPU time slice sharing strategy specifically includes: c1. Use MIG technology to partition the GPU and register each partition as an independent device to the cluster through DevicePlugin; c2. When a user submits a GPU task, the scheduler makes a decision based on the task requirements and node resource status. When the Pod is started, container-toolkit is responsible for mapping the specified MIG instance to the container and setting the corresponding environment to ensure that the container can only access the predetermined GPU resources. c3. Through continuous monitoring and closed-loop feedback mechanisms, the graphics memory allocation is dynamically adjusted according to the real-time load, and GPU resources are efficiently shared in hybrid deployment scenarios.

7. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: The optimization scheduling strategy includes a multi-resource-aware scheduling strategy, which specifically includes: d1. Input online reasoning task set T online , offline training task set T office , node resource status R i ( i =1,…, N ), task priority function P ( t j );in, t j Indicates j tasks; d2. Initialization and monitoring: Initialize all nodes R i =( C i , M i , G i ), monitor the resource usage of all nodes in real time and calculate the load status; C i Indicates i The CPU resources of each node, M i Indicates i The memory resources of each node, G i Indicates i GPU resources of each node; d3.Task classification and priority allocation: Combine tasks T online ∪ T office ,according to P ( t j ) sorted by priority; d4. Resource allocation and scheduling: For online reasoning tasks t j ∈ T online , find available nodes i , meeting resource requirements , , If a suitable node is found, the task will be scheduled to the node and the resource status of the node will be updated. A [ t j ]= n i , the task t j Assign to Node n i , update node resources , if the node resources are insufficient, wait for the resources to be released; among them, Indicates j The CPU resource requirements of each task; Indicates j Memory resources for each task; Indicates j GPU resources for each task; n i Indicates i nodes; A [ t j ] indicates the j Node allocation of tasks; For offline training tasks, t j ∈ T office , if the node has enough resources, schedule the task t j Go to the node that meets the constraints and update the node resources When node resources are insufficient, a dynamic adjustment mechanism is introduced to calculate Reclaim resources from low-priority tasks, update , , reallocate resources to online reasoning tasks; where, Indicates changes in node resources. Indicates the task j The current resources of the node. Indicates the task l The current resources of the node. Indicates the task k The current resources of the node; d5. Dynamic adjustment and iterative optimization: monitor resource utilization and task completion status in real time, dynamically adjust resource allocation according to node load, optimize scheduling strategy, and finally output task-to-node deployment plan mapping A as the task scheduling plan.

8. A deep learning task hybrid deployment system, characterized in that: include: The task scheduling optimizer is used to parse the tasks submitted by users, extract key information such as task type, resource requirements, priority, service level, and data locality, and evaluate the resource consumption and execution time of the task based on historical data and current system status. Combined with the real-time resource usage and load status of the node, a multi-dimensional scoring model is used to evaluate each node, and finally one or more target nodes suitable for task operation are determined; The SLO manager is used to track the real-time data of various performance indicators of the system and compare the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimized scheduling strategy; Koordlet component is used to execute specific resource allocation operations issued by the scheduling policy on the target node, continuously monitor the node resource usage, and feed back the node resource usage and load status after the task is executed to the task scheduling optimizer; Traffic security monitor, which is used to monitor the data flow between nodes in real time. By comparing with the historical normal status data, when abnormal traffic patterns or sudden traffic surges are identified, the detected abnormal information will be fed back to the SLO manager in real time; In the task scheduling optimizer, the service level quality types of tasks include: SYSTEM: In the system process, strict resource constraints are designed based on the required resource usage corresponding to the minimum latency of the SYSTEM response, to ensure that the system process obtains the necessary resources to ensure the latency response while not occupying too many resources; LSE: Develop dedicated resource pools and resource reservation machines to ensure that these tasks can be executed stably and with low latency in a low-interference environment; LSR: Introduces stricter resource isolation policies and priority scheduling. The resource isolation policy reserves necessary resources for these LSR tasks, and the priority scheduling ensures that the priority of LSR tasks is always in a higher priority sequence. LS: Adopts elastic resource sharing mode. When system resources are sufficient, more resources are allocated to tasks. When resource contention between tasks is detected or user request traffic suddenly surges, the occupied resources of tasks are reduced. BE: adopts a best-effort strategy. The best-effort strategy is to make full use of idle resources in the cluster while ensuring that the resource requirements of high-priority tasks are not affected.

9. The deep learning task hybrid deployment system according to claim 8, characterized in that: The SLO manager includes: The colocation configuration component is used by users to perform simple colocation configuration. Resource over-issuance management component, used to dynamically adjust the resource over-issuance ratio according to the real-time status of the node; Workload statistics component, used to analyze and predict resource requirements using histograms and optimize load distribution; The task scheduling optimizer comprises: The colocation scheduler component is used to balance node loads based on load-aware policies, provide fine-grained resource isolation policies for different tasks, and support dynamic allocation of heterogeneous resources. The rescheduling component is used to dynamically evaluate the current task distribution based on the node load status and task performance indicators collected in real time to readjust the resource allocation strategy.

Citation Information

Patent Citations

  • Online training-oriented computing power resource elastic distribution system

    CN119166278A

  • High throughput cloud computing resource recovery system

    WO2023015787A1