Deep learning task hybrid deployment method and system

By introducing a multi-dimensional scoring model and dynamic resource scheduling mechanism in the deep learning task hybrid deployment system, the problems of low resource utilization and high task delay in the existing technology are solved, and efficient coordinated deployment and intelligent decision-making of online and offline tasks are realized.

CN119987974AActive Publication Date: 2025-05-13HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202510445993.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-13
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

When faced with complex task loads and diversified hardware environments, existing deep learning task hybrid deployment technology is difficult to achieve refined scheduling, energy consumption optimization and intelligent decision-making, resulting in low resource utilization and high task delays and failure rates.

Method used

A hybrid deployment method for deep learning tasks is proposed. Tasks are submitted through Kubernetes native interface, and components such as task scheduling optimizer, SLO manager, Koordlet component and traffic security monitor are used to realize the evaluation of nodes and dynamic resource scheduling by multi-dimensional scoring models to ensure the improvement of service level and resource utilization of online tasks.

Benefits of technology

Through dynamic resource scheduling and intelligent decision-making, dynamic resource sharing between online and offline tasks is realized, ensuring that the system can still operate stably and efficiently under high load conditions, and improving resource utilization and task throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987974A_ABST
    Figure CN119987974A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer deep learning, in particular to a deep learning task hybrid deployment method and system. The method comprises the steps that S1, a user submits a task through a Kubernetes native interface; s2, a task scheduling optimizer analyzes tasks according to resource requirements and service levels and allocates the tasks to proper nodes; s3, the SLO manager tracks real-time data of various performance indexes of the system and compares the real-time data with a preset service grade target; s4, the Koordlet component executes a specific resource allocation operation on the target node; and S5, the traffic safety monitor monitors the data streams among the nodes in real time. According to the method, a mixed deployment strategy for different types of resources is provided by analyzing the periodical rule of deep learning task resource use, dynamic resource sharing of online and offline tasks is achieved, and meanwhile it is ensured that the system can still run stably and efficiently under the high-load condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer deep learning technology, and in particular to a method and system for hybrid deployment of deep learning tasks. Background Art

[0002] With the rapid development of deep learning technology, its wide application in image recognition, natural language processing, and recommendation systems continues to increase the demand for computing resources. Faced with the significant difference in resource requirements between online real-time reasoning and offline batch training deep learning tasks, hybrid deployment methods have become a key research direction for optimizing cloud data center resource utilization and improving task performance.

[0003] Current hybrid deployment technologies mainly focus on the following aspects: 1) Task characteristic analysis and classification: By analyzing the low latency requirements of online tasks and the high throughput requirements of offline tasks, existing methods classify tasks based on computing intensity, communication mode, and latency sensitivity to provide a scheduling basis for hybrid deployment. For example, the real-time requirements of online tasks are prioritized while processing offline tasks using resource gaps.

[0004] 2) Resource Scheduling and Allocation: Resource scheduling is the core of hybrid deployment. Static resource partitioning methods have limited efficiency, while dynamic scheduling algorithms adjust resource allocation in real time according to task load, improving resource utilization. Intelligent scheduling methods based on reinforcement learning and deep learning optimize decisions through task characteristics and historical data, further improving system performance.

[0005] 3) Optimization of online and offline task coexistence: Optimize resource competition between tasks through task priority adjustment and resource isolation mechanism. For example, containerization technology implements task resource isolation, while dynamic resource concession strategy balances task performance and system throughput.

[0006] 4) Collaborative scheduling of heterogeneous hardware: Combine task characteristics with hardware characteristics such as GPU and TPU to match and allocate tasks and improve computing efficiency. Lightweight tasks are prioritized to low-power hardware, and computationally intensive tasks rely more on GPUs, achieving dual optimization of performance and energy consumption.

[0007] The current research on hybrid deployment methods for deep learning tasks mainly focuses on four directions: task classification, resource scheduling, coexistence optimization, and heterogeneous hardware collaboration. These technologies have promoted the further popularization of deep learning in large-scale applications by optimizing the resource utilization and task performance of cloud data centers from multiple angles. However, in the face of increasingly complex task load scenarios and diverse hardware environments in the future, hybrid deployment technology still needs to be further explored in terms of refined scheduling, energy consumption optimization, and intelligent decision-making.

[0008] The current hybrid deployment technology has the following technical defects: 1) Hybrid deployment data center resource allocation and scheduling solution: Although the existing resource allocation schemes for deep learning task deployment have improved resource utilization to a certain extent, they also have many shortcomings. Although the static allocation method is simple in structure, it lacks flexibility and is difficult to cope with fluctuations in dynamic task loads, which can easily lead to resource waste or increased task delays. The dynamic allocation method improves efficiency by adjusting resource configuration in real time, but its computational complexity is high and it is easy to cause system performance fluctuations due to frequent resource reallocation. In addition, the existing dynamic scheduling algorithm lacks comprehensive consideration of task priority and resource requirements in multi-task mixed load scenarios, making it difficult to guarantee the service level objective (SLO) of online tasks. Although more complex scheduling strategies (such as reinforcement learning-driven solutions) have excellent theoretical performance, they have high deployment costs in actual production environments and face the problem of high sensitivity to tuning parameters.

[0009] 2) Task hybrid deployment system: Although the mixed task deployment system has achieved certain results in practical applications, there are still key challenges. First, the current system generally lacks the ability to deeply perceive different types of tasks, and only allocates resources through simple scoring rules or static strategies, which makes it difficult to adapt to complex dynamic environments. In addition, many systems have insufficient trade-offs between task isolation and resource sharing. Over-emphasis on isolation may lead to resource waste, while excessive sharing will cause resource competition and affect the stability of online tasks. Even for more advanced systems (such as Google Borg and Tencent Caelus), task scheduling still has problems with increased latency and decreased throughput under large-scale task load environments. Another prominent problem is the insufficient overhead control of containerization technology, especially in scenarios with high-frequency task migration.

[0010] 3) Deep learning task service system optimization strategy: The optimization of deep learning task service systems has achieved remarkable results in improving resource utilization and performance, but it still faces many limitations. Existing methods mostly focus on latency optimization and throughput improvement, but it is difficult to take both into account at the same time in multi-task concurrent scenarios. For example, although the InferLine and INFaaS systems perform well in latency control, their scalability in large-scale model concurrent deployment is weak. In addition, the efficient use of heterogeneous hardware resources such as GPUs remains a difficulty. Existing methods mostly rely on preemption and concurrency control, but may cause bottlenecks due to high utilization of hardware resources. Although the dynamic expansion and resource allocation schemes of model services can improve performance, it is difficult to achieve stability under burst traffic, resulting in a significant increase in task delays and failure rates. At the same time, in response to the service needs of diverse deep learning models, existing optimization strategies lack sufficient flexibility and generalization capabilities, and it is difficult to meet the complex requirements of different task loads and hardware environments.

[0011] Although existing technologies have made significant progress in resource allocation, task scheduling, and system optimization, there is still room for improvement in resource utilization efficiency, task isolation and sharing balance, system scalability, and dynamic environment adaptability. Future research needs to further combine intelligent scheduling algorithms, refined resource management, and load-aware technology to improve the performance and stability of hybrid deployment methods for deep learning tasks.

[0012] Currently, cloud data centers are widely used for the deployment and execution of deep learning tasks, but the independent deployment mode of online and offline tasks leads to low resource utilization. In particular, resource allocation and scheduling issues are particularly prominent in the context of deep learning tasks being highly dependent on GPU resources. To address this issue, cloud service providers have begun to try a hybrid deployment strategy of online and offline tasks, using the idle resources of online tasks through offline tasks to improve overall resource utilization. However, this hybrid model also brings the risk of violating the service level agreement of online tasks. At the same time, when facing burst traffic and complex computing dependencies, it is still difficult to balance system stability and resource utilization efficiency. Summary of the invention

[0013] The present invention provides a method and system for hybrid deployment of deep learning tasks, aiming to optimize resource scheduling and task management, and to improve data center resource utilization and task throughput while ensuring the service level of online tasks.

[0014] The present invention provides a deep learning task hybrid deployment method, comprising: S1. The user submits a task through the Kubernetes native interface; S2. The task scheduling optimizer analyzes the tasks submitted by users, extracts key information such as task type, resource requirements, priority, service level, and data locality, and evaluates the resource consumption and execution time of the tasks based on historical data and the current system status. It uses a multi-dimensional scoring model to evaluate each node based on the real-time resource usage and load status of the nodes, and finally determines one or more target nodes suitable for task operation. S3. The SLO manager tracks the real-time data of various performance indicators of the system and compares the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimal scheduling strategy; S4. After determining the node, the scheduling strategy is sent to the Koordlet component. The Koordlet component performs specific resource allocation operations on the target node, continuously monitors the node resource usage, and feeds back the node resource usage and load status after the task is executed to the task scheduling optimizer; S5. The traffic security monitor monitors the data flow between nodes in real time. By comparing it with the historical normal status data, when an abnormal traffic pattern or sudden traffic surge is identified, the detected abnormal information will be fed back to the SLO manager in real time.

[0015] As a further improvement of the present invention, in step S2, the task scheduling optimizer execution process also includes a rescheduling component optimizing the task distribution process according to the node load status: During the task execution process, the rescheduling component dynamically evaluates the current task distribution based on the node load status and task performance indicators collected in real time. If it is found that a node is overloaded or the task is running inefficiently, the rescheduling component triggers the rescheduling process and migrates some tasks to nodes with sufficient resources and lower load.

[0016] As a further improvement of the present invention, in step S2, service levels are divided for different tasks, an association between service levels and task scheduling priorities is established, online reasoning tasks are assigned high service levels, and offline training tasks are assigned low service levels. When there is competition for resources in the task scheduling optimizer, online reasoning tasks with high service levels are executed first.

[0017] As a further improvement of the present invention, in step S2, the service level quality type of the task includes: SYSTEM: In the system process, strict resource constraints are designed based on the required resource usage corresponding to the minimum latency of the SYSTEM response, to ensure that the system process obtains the necessary resources to ensure the latency response while not occupying too many resources; LSE: Develop dedicated resource pools and resource reservation machines to ensure that these tasks can be executed stably and with low latency in a low-interference environment; LSR: Introduces stricter resource isolation policies and priority scheduling. The resource isolation policy reserves necessary resources for these LSR tasks, and the priority scheduling ensures that the priority of LSR tasks is always in a higher priority sequence. LS: Adopts elastic resource sharing mode. When system resources are sufficient, more resources are allocated to tasks. When resource contention between tasks is detected or user request traffic suddenly surges, the occupied resources of tasks are reduced. BE: adopts a best-effort strategy. The best-effort strategy is to make full use of idle resources in the cluster while ensuring that the resource requirements of high-priority tasks are not affected.

[0018] As a further improvement of the present invention, the optimized scheduling strategy includes a CPUBurst strategy, and the CPUBurst strategy specifically includes: a1. CPUBurst time slice adjustment based on task history: Dynamically adjust the time slice length allocated to the task in the future according to the historical CPUBurst time of the task. For tasks with shorter CPUBurst, a shorter time slice is allocated, and for tasks with longer CPUBurst, a longer time slice is allocated; a2. Multi-level feedback queue scheduling strategy: Different tasks are assigned to different queues, and the queue priority is continuously adjusted according to the task execution status. Tasks enter the corresponding priority queue according to the initial CPUBurst time and scheduling strategy. Shorter tasks are placed in the high-priority queue and use shorter time slices, while longer tasks are assigned to the low-priority queue and use longer time slices.

[0019] As a further improvement of the present invention, the optimization scheduling strategy includes a memory over-issuance design strategy, and the memory over-issuance design strategy specifically includes: b1. Reasonable configuration of resource request and limit values: Configure resourcesrequest and resourceslimit for each task. For online deep learning tasks, set higher resourcesrequest and resourceslimit. When scheduling, online deep learning tasks will give priority to nodes with abundant resources for deployment. For offline deep learning tasks, set lower resourcesrequest and higher resourceslimit. When cluster resources are abundant, offline deep learning tasks will use more memory resources and will be restricted or evicted first when resources are scarce. b2. Task management based on QoS categories: Different tasks are classified and managed in the Kubernetes cluster based on QoS categories: b21. Resource management of system tasks SYSTEM: Set appropriate resourceslimit for SYSTEM type tasks to ensure that SYSTEM type tasks run within the control range; b22. Delay-sensitive exclusive tasks LSE: Reduce the memory limit of LSE type tasks to release some memory for temporary use by other non-critical tasks; b23. Latency-sensitive reserved tasks (LSR): LSR-type tasks are bound to CPU cores, assigned to memory pools with higher priorities, and the memory usage patterns of LSR-type tasks are monitored; b24. Latency-sensitive tasks LS: Allow LS-type tasks to obtain higher resource utilization rights when cluster resources are abundant, and share resources when resources are scarce; occupy memory not used by other tasks when the load is low, and automatically reduce memory usage when the load increases; b25. Best-effort task BE: No specific resource request and limit values ​​are set.

[0020] As a further improvement of the present invention, the optimization scheduling strategy includes a GPU time slice sharing strategy, and the GPU time slice sharing strategy specifically includes: c1. Use MIG technology to partition the GPU and register each partition as an independent device to the cluster through DevicePlugin; c2. When a user submits a GPU task, the scheduler makes a decision based on the task requirements and node resource status. When the Pod is started, container-toolkit is responsible for mapping the specified MIG instance to the container and setting the corresponding environment to ensure that the container can only access the predetermined GPU resources. c3. Through continuous monitoring and closed-loop feedback mechanisms, the graphics memory allocation is dynamically adjusted according to the real-time load, and GPU resources are efficiently shared in hybrid deployment scenarios.

[0021] As a further improvement of the present invention, the optimization scheduling strategy includes a multi-resource-aware scheduling strategy, and the multi-resource-aware scheduling strategy specifically includes: d1. Input online reasoning task set T online , offline training task set T office , node resource status R i ( i =1,…, N ), task priority function P (t j );in, t j Indicates j tasks; d2. Initialization and monitoring: Initialize all nodes R i =( C i , M i , G i ), monitor the resource usage of all nodes in real time and calculate the load status; C i Indicates i The CPU resources of each node, M i Indicates i The memory resources of each node, G i Indicates i GPU resources of each node; d3.Task classification and priority allocation: Combine tasks T online ∪ T office ,according to P ( t j ) sorted by priority; d4. Resource allocation and scheduling: For online reasoning tasks t j ∈ T online , find available nodes i , meeting resource requirements , , If a suitable node is found, the task will be scheduled to the node and the resource status of the node will be updated. A [ t j ]= n i , the task t j Assign to Node n i , update node resources , if the node resources are insufficient, wait for the resources to be released; among them, Indicates j The CPU resource requirements of each task; Indicates j Memory resources for each task; Indicatesj GPU resources for each task; n i Indicates i nodes; A [ t j ] indicates the j Node allocation of tasks; For offline training tasks, t j ∈ T office , if the node has enough resources, schedule the task t j Go to the node that meets the constraints and update the node resources When node resources are insufficient, a dynamic adjustment mechanism is introduced to calculate ∆ R Reclaim resources from low-priority tasks, update , , reallocate resources to online reasoning tasks; where ∆ R Indicates changes in node resources. Representation Task j The current resources of the node. Representation Task l The current resources of the node. Representation Task k The current resources of the node; d5. Dynamic adjustment and iterative optimization: monitor resource utilization and task completion status in real time, dynamically adjust resource allocation according to node load, optimize scheduling strategy, and finally output task-to-node deployment plan mapping A as the task scheduling plan.

[0022] The present invention also provides a deep learning task hybrid deployment system comprising: The task scheduling optimizer is used to parse the tasks submitted by users, extract key information such as task type, resource requirements, priority, service level, and data locality, and evaluate the resource consumption and execution time of the task based on historical data and current system status. Combined with the real-time resource usage and load status of the node, a multi-dimensional scoring model is used to evaluate each node, and finally one or more target nodes suitable for task operation are determined; The SLO manager is used to track the real-time data of various performance indicators of the system and compare the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimized scheduling strategy; Koordlet component is used to execute specific resource allocation operations issued by the scheduling policy on the target node, continuously monitor the node resource usage, and feed back the node resource usage and load status after the task is executed to the task scheduling optimizer; The traffic security monitor is used to monitor the data flow between nodes in real time. By comparing with the historical normal status data, when abnormal traffic patterns or sudden traffic surges are identified, the detected abnormal information will be fed back to the SLO manager in real time.

[0023] As a further improvement of the present invention, the SLO manager includes: a colocation configuration component, which is used by users to perform simple colocation configuration; a resource over-issuance management component, which is used to dynamically adjust the resource over-issuance ratio according to the real-time status of the node; a workload statistics component, which is used to use histograms to analyze and predict resource requirements and optimize load distribution; The task scheduling optimizer includes: a co-location scheduler component, which is used to balance node loads according to a load-aware strategy, provide a sophisticated resource isolation strategy for different tasks, and support dynamic allocation of heterogeneous resources; a rescheduling component, which is used to dynamically evaluate the current task distribution according to the node load status and task performance indicators collected in real time to readjust the resource allocation strategy.

[0024] The beneficial effects of the present invention are: by analyzing the periodic laws of resource usage of deep learning tasks, a strategy for mixed deployment of different types of resources is proposed to achieve dynamic resource sharing of online and offline tasks, while ensuring that the system can still operate stably and efficiently under high load conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is an architectural diagram of the deep learning task hybrid deployment system of the present invention; Figure 2 It is a total resource demand division strategy diagram of the present invention; Figure 3 It is the CPUBurst strategy principle diagram of the present invention; Figure 4 It is a schematic diagram of the memory over-issuance allocation scheme of the present invention; Figure 5 is a schematic diagram of the GPU sharing solution of the present invention; Figure 6 1 is an experimental result diagram of the utilization rate of various system resources during different types of deep learning training tasks and colocation tasks of the present invention; Figure 7 This is the throughput result graph of the ResNet50 inference task within one hour of the present invention; Figure 8 This is a performance diagram of the mixed-part performance of the ResNet50 task under different loads of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0027] 1. Hybrid deployment method and system for deep learning tasks based on Kubernetes.

[0028] 1.1 Deep learning task hybrid deployment method and system description: like Figure 1 As shown, a deep learning task hybrid deployment system of the present invention aims to achieve efficient collaborative deployment of online tasks and offline deep learning tasks by extending the native functions of Kubernetes to maximize resource utilization and service quality.

[0029] The core objectives of the system include the following aspects: High reusability: Make use of Kubernetes' native functions as much as possible to avoid extensive modifications to the existing platform; Scalability: modular design supports flexible expansion and maintenance; Resource optimization: Optimize the use of key resources such as CPU, GPU and memory through dynamic scheduling strategies; Service quality assurance: Through the management of service level objectives (SLO), the performance requirements of online tasks are ensured, and offline tasks are effectively isolated to avoid resource competition.

[0030] The system architecture consists of native components and extended components of Kubernetes. Native components are responsible for the basic functions of the cluster, such as Kubernetes-APIServer on the master node is responsible for cluster node status management, Kubernetes scheduler is responsible for task scheduling control, and Kubelet on the worker node is responsible for managing the worker node status; extended components include node center control components (corresponding to SLO manager), traffic security monitor, Koordlet of the worker node is responsible for the status management of the worker node, and other custom components to implement advanced functions of the system.

[0031] Specifically, the deep learning task hybrid deployment system includes: The task scheduling optimizer is used to parse the tasks submitted by users, extract key information such as task type, resource requirements, priority, service level, and data locality, and evaluate the resource consumption and execution time of the task based on historical data and current system status. Combined with the real-time resource usage and load status of the node, a multi-dimensional scoring model is used to evaluate each node, and finally one or more target nodes suitable for task operation are determined; The SLO manager is used to track the real-time data of various performance indicators of the system and compare the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimized scheduling strategy; Koordlet component is used to execute specific resource allocation operations issued by the scheduling policy on the target node, continuously monitor the node resource usage, and feed back the node resource usage and load status after the task is executed to the task scheduling optimizer; Traffic security monitor, which is used to monitor the data flow between nodes in real time. By comparing with the historical normal status data, when abnormal traffic patterns or sudden traffic surges are identified, the detected abnormal information will be fed back to the SLO manager in real time; Task scheduling optimizer: designs refined scheduling plans based on task characteristics and dynamically adjusts task allocation strategies; traffic security monitor: monitors system traffic status in real time, detects abnormal traffic patterns and prevents potential attacks; SLO manager: dynamically adjusts resource allocation strategies based on the service level requirements of tasks; container runtime trigger in the Koordlet component, namely the resource isolation module: ensures the coordinated operation of online and offline tasks through fine-grained resource isolation technology.

[0032] The node-centric control component is deployed in the cluster in the form of a Deployment, and consists of a leader and a backup instance. Its main functions include: Colocation configuration component: supports colocation configuration, and users can achieve efficient task integration through simple configuration; resource over-issuance management component: dynamically adjusts the resource over-issuance ratio according to the real-time status of the node to maintain service stability; workload statistics component: uses histograms to analyze and predict resource requirements and optimize load distribution.

[0033] The hybrid scheduler component is used to schedule tasks and isolate resources on Kubernetes. It enhances the functions of the Kubernetes native scheduler and supports the following capabilities: QoS-aware scheduling: balancing node loads based on load-aware policies; differentiated SLO: providing fine-grained resource isolation policies for different tasks; elastic resource management: supporting dynamic allocation of heterogeneous resources to improve system throughput. The hybrid scheduler also provides resource reservation and node defragmentation functions to further improve scheduling efficiency.

[0034] The rescheduling component optimizes task distribution through a load-aware scheduling framework to avoid node hot spots. This component can dynamically adjust the task distribution of nodes to improve system stability and performance.

[0035] The Koordlet component is deployed as a DaemonSet to support resource management in hybrid deployments, including the following features: Container runtime trigger: real-time estimation of Pod resource usage, support for resource over-issuance and lifecycle monitoring; QoS management controller: dynamically adjust node resource allocation strategies to suppress interference that may affect service quality; Node storage and resource optimization: provides optimized management for CPU, memory, and video memory.

[0036] Load balancing is achieved through the collaboration of the task scheduling optimizer and the Koordlet component. The scheduler optimizes task allocation globally, while the Koordlet component executes resource allocation strategies within the node, ensuring efficient operation of the system through a multi-level scheduling mechanism.

[0037] Based on the deep learning task hybrid deployment system, the present invention also provides a deep learning task hybrid deployment method, including: S1. The user submits a task through the Kubernetes native interface; S2. The task scheduling optimizer analyzes the tasks submitted by users, extracts key information such as task type, resource requirements, priority, service level, and data locality, and evaluates the resource consumption and execution time of the tasks based on historical data and the current system status. It uses a multi-dimensional scoring model to evaluate each node based on the real-time resource usage and load status of the nodes, and finally determines one or more target nodes suitable for task operation. The rescheduling component optimizes the task distribution process according to the node load status: During the task execution process, the rescheduling component dynamically evaluates the current task distribution based on the node load status and task performance indicators collected in real time. If it is found that a node is overloaded or the task is running inefficiently, the rescheduling component triggers the rescheduling process and migrates some tasks to nodes with sufficient resources and lower loads. S3. The SLO manager tracks the real-time data of various performance indicators of the system and compares the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimal scheduling strategy; S4. After determining the node, the scheduling strategy is sent to the Koordlet component. The Koordlet component performs specific resource allocation operations on the target node, continuously monitors the node resource usage, and feeds back the node resource usage and load status after the task is executed to the task scheduling optimizer; S5. The traffic security monitor monitors the data flow between nodes in real time. By comparing it with the historical normal status data, when an abnormal traffic pattern or sudden traffic surge is identified, the detected abnormal information will be fed back to the SLO manager in real time.

[0038] The task scheduling optimizer is dedicated to achieving efficient scheduling and optimizing resource utilization in a hybrid deployment environment for offline deep learning tasks. First, the tasks submitted by users are parsed to extract key information such as task type, resource requirements, priority, service level requirements, and data locality, and to establish a task feature vector. By using historical task data, the system can estimate the execution time and resource consumption of tasks, thereby providing a basis for task classification and stratification, while identifying the dependencies between tasks to ensure a reasonable scheduling order and avoid resource conflicts.

[0039] On this basis, the system collects the resource information and health status of each node in real time, obtains the CPU, memory, GPU, storage and network bandwidth usage of each node through monitoring tools such as KubernetesAPI and Prometheus, and marks or temporarily excludes abnormal nodes in combination with the load level and historical stability of the node. The resource capacity model of each node not only records the total resources and idle resources, but also includes data locality information, so that tasks can be scheduled to nodes close to the location of the required data as much as possible to reduce data transmission delays.

[0040] Task scheduling decisions are made by constructing a matching matrix between task requirements and node resource capabilities. First, nodes that meet the minimum resource requirements are screened out, and then these nodes are comprehensively evaluated using a multi-dimensional scoring mechanism. This scoring mechanism takes into account multiple factors such as resource matching, node load, data locality, and SLO satisfaction. Each indicator is quantified by setting a weighted scoring function, and the node with the highest score is finally selected as a candidate target. Depending on the priority and urgency of the task, the system will also adopt reservation, preemption, or task migration strategies in scenarios with tight resources or high concurrency to ensure that the optimal scheduling decision can be implemented on specific nodes.

[0041] In terms of resource reservation, the scheduling optimizer reserves the required resources after determining the target node through Kubernetes' scheduling extension or custom scheduler interface to prevent resources from being occupied by other tasks during the scheduling process. After the task scheduling decision result is passed to Kubernetes, the system can successfully start the task on the reserved resources and monitor the startup and execution of the task in real time. After the task is executed, the system collects performance data including execution time, resource utilization, and task success rate, providing feedback for the optimization of subsequent scheduling strategies and resource allocation models.

[0042] The solution also builds a closed-loop feedback mechanism, which feeds back the data after task execution to the scheduling optimizer to achieve dynamic adjustment of future scheduling strategies. For long-running tasks or sudden load changes, the system can evaluate the task status in real time and trigger task migration or resource reallocation when necessary to ensure that service level objectives are achieved. Furthermore, the system introduces a machine learning model to predict task execution time, resource consumption, and node load through historical scheduling data, thereby optimizing scheduling decisions in advance. At the same time, it uses a reinforcement learning algorithm to continuously adjust itself during the scheduling process to form an adaptive intelligent optimization mechanism. In order to take into account resource utilization, task execution efficiency, and SLO satisfaction, the system also builds a multi-objective optimization model, and combines heuristic search, greedy algorithms, or genetic algorithms to quickly solve the optimal or near-optimal scheduling solution under resource constraints.

[0043] The traffic security monitor and SLO manager play the dual roles of real-time security protection and performance assurance in the system, and the two work together to dynamically adjust resource allocation strategies. After the system is started, the traffic security monitor begins to monitor the data flow between nodes in real time, and identifies any abnormal traffic patterns or sudden traffic surges by comparing with historical normal status data. This process relies on preset security rules and advanced data analysis algorithms. Once potential attack risks or abnormal behaviors are detected, the monitor will immediately trigger resource adjustment strategies, such as reducing resource supply to affected nodes, or directing part of the traffic to a dedicated isolated environment for in-depth detection and processing, thereby effectively reducing the spread of security risks.

[0044] At the same time, the SLO manager continuously tracks various performance indicators of the system, such as response time, throughput, and latency, and compares these real-time data with the pre-set service level objectives. If the performance indicators of a service or task deviate, the SLO manager will immediately start the dynamic resource scheduling process and reallocate resources by calling the scheduling optimization module inside the system. For example, when it is detected that the response time of a task exceeds the standard, the system may automatically increase CPU or memory resources for the task, or adjust the priority of the task so that critical tasks are given priority in resource competition. This process not only relies on real-time monitoring data, but also uses historical operation data and prediction models to predict resource requirements, thereby achieving accurate scheduling.

[0045] The entire adjustment mechanism forms a closed-loop feedback system: the abnormal information detected by the traffic security monitor will be fed back to the SLO manager in real time, which will start the resource adjustment strategy based on the performance deviation data and notify the scheduling optimizer to reconfigure the tasks and nodes. After the adjustment is completed, the system will continue to monitor resource allocation and task operation to ensure that security protection measures and service quality requirements are always in a balanced state, and further iterate and optimize when necessary. Through this specific implementation plan, the system can respond quickly to sudden security incidents and performance bottlenecks, dynamically balance security and performance requirements, and ultimately ensure the stable and efficient operation of the overall service.

[0046] During the operation of the entire system, the user first submits a task through the Kubernetes native interface. At this time, the task contains resource requirements, service level objectives, and other task metadata. After receiving the task, the task scheduling optimizer will first parse and preprocess the task, extract key attributes, and evaluate the resource consumption and priority of the task based on historical data and current system status. Based on this information, the scheduler combines the real-time resource status and load status of the node, and uses a multi-dimensional scoring model to evaluate each node, and finally determines one or more target nodes that are most suitable for the task to run.

[0047] After the node is determined, the scheduling decision is sent to the Koordlet component. Koordlet is responsible for performing specific resource allocation operations on the target node. It reserves the required resources in advance, starts the task container, and ensures that the task runs smoothly in the allocated resource environment through deep integration with the Kubernetes API. At the same time, Koordlet also continuously monitors the resource usage of the node, such as CPU, memory, network bandwidth, and storage, so as to promptly detect resource bottlenecks or abnormal loads during the task operation.

[0048] At the same time, the traffic security monitor inside the system monitors the network traffic and data exchange status of the task in real time with millisecond-level accuracy. By comparing the preset security rules with the historical traffic model, once abnormal traffic or potential signs of attack are detected, an alarm signal will be quickly issued, and it will work with the SLO management module to adjust the traffic of the task or node, such as limiting the traffic, isolating abnormal nodes, or directing the traffic to the security inspection area, so as to ensure the overall security and stability of the system.

[0049] During the task execution process, the system also has a built-in rescheduling component, which dynamically evaluates the current task distribution based on the node load data and task performance indicators collected in real time. If some nodes are found to be overloaded or the task operation efficiency is low, the rescheduling component will trigger the rescheduling process and migrate some tasks to nodes with sufficient resources and lower load. This dynamic rescheduling mechanism can not only relieve the pressure on a single node in a timely manner, but also ensure that the system is always in the best operating state and ensure that the service level objectives are continuously met.

[0050] The entire process forms a closed-loop feedback system, in which tasks are continuously iterated and optimized during submission, scheduling, resource allocation, real-time monitoring, and dynamic rescheduling, ensuring that the system can respond quickly to sudden traffic changes and resource bottlenecks, and continuously improve overall performance and security.

[0051] 1.2 Service level classification in the colocation system: In a hybrid deployment system, the resource demand analysis and service level division of offline deep learning tasks are the basis for efficient scheduling and optimal resource utilization. In this environment, online reasoning tasks and offline training tasks often need to run at the same time, but their demand characteristics and service requirements for system resources are very different. Online reasoning tasks are usually highly sensitive to latency, so their low latency and high response speed must be prioritized to meet the performance requirements of users; while offline training tasks focus more on overall computing efficiency and resource utilization, and can tolerate a certain degree of scheduling delay and resource competition.

[0052] For hybrid deployment systems, it is necessary to divide service levels for different tasks, and to achieve reasonable allocation and use of resources by associating service levels with task scheduling priorities. Online reasoning tasks are usually assigned higher service levels because their latency directly affects user experience and the overall performance of the system. Offline training tasks may be at a lower service level so that they can obtain more resources when cluster resources are sufficient, and actively release resources when resources are tight to give priority to meeting the needs of online reasoning tasks. By defining the service level of tasks, it can be ensured that the resource scheduler gives priority to meeting the needs of critical tasks when facing resource competition, thereby improving the QoS (Quality of Service) of the overall system. At the same time, the division of service levels also provides clear priority guidance for scheduling strategies, allowing the scheduler to dynamically adjust according to the importance and urgency of different tasks.

[0053] Through a detailed discussion of the resource requirements and service level division of offline deep learning tasks, theoretical support is provided for the scheduling mechanism of the entire hybrid deployment platform. Such analysis will provide a data basis for formulating efficient resource scheduling solutions, so that the response time and latency of online reasoning tasks can be effectively controlled, while also ensuring that offline training tasks obtain as many resources as possible without affecting online services, achieving the optimal configuration of overall performance. This in-depth understanding of resource requirements and service levels can help the hybrid deployment platform maximize resource utilization and fine-tune task scheduling while ensuring service quality.

[0054] In the present invention, the total amount of resources is divided into four indicator lines according to the proportion: limit, usage, short-term reservation, and long-term reservation. The basic idea is to use those allocated but unused resources to run low-priority pods, such as Figure 2 shown.

[0055] Limit in the figure: The gray line indicates the amount of resources requested by the high-priority Pod, that is, the requested amount of resources for the highest-priority inference deep learning task, which corresponds to the native Burstable type Pod request of Kubernetes.

[0056] Usage: The red line indicates the actual amount of resources used by the Pod. The horizontal axis is the timeline, and the red line is the fluctuation curve of the Pod load over time.

[0057] Short-term reservation: The dark blue line is an estimate of resource usage in the future based on resource usage in the past (short) period. The difference between reservation and limit is that the allocated unused (resources that will not be used in the future) can be used to run batch Pods for short-term execution.

[0058] Long-termreservation: The light blue line is similar to short-termreservation, but the estimated historical usage period is longer. Resources from reserved to limited can be used for Pods with longer life cycles, and compared with short-term forecast values, there are fewer available resources, but they are more stable.

[0059] At the same time, since the granularity of Kubernetes' native service level quality (Quality of Service, QoS) division is too large, it cannot meet the fine-grained task priority differentiation in the co-location scenario. The present invention redesigns five new QoS types supported by the scheduling system in the co-location system.

[0060] (1) SYSTEM: SYSTEM-type tasks are system processes, and their resource usage is restricted. For example, system services such as DaemonSets need to ensure the latency of system services, but they also need to limit the resource usage of these system service containers on the node to ensure that they do not occupy too many resources. For SYSTEM-type tasks, a strict resource restriction design is adopted, mainly in the system process based on the required resource usage corresponding to the minimum latency time of its response. This ensures that system processes (such as DaemonSets, etc.) do not occupy too many resources while obtaining the necessary resources to ensure latency response, thereby ensuring the stable operation of the entire system.

[0061] (2) LSE (LatencySensitiveExclusive): LSE tasks require the reservation of resources and the organization of pods with the same QoS to share resources. They are generally used for deep learning inference tasks with special latency requirements. They are used in independent resource pools and correspond to some of the native Guaranteed tasks of Kubernetes. The core of the design of LSE tasks is resource exclusivity. Through dedicated resource pools and resource reservation mechanisms (such as CPU binding and exclusive memory allocation), these tasks can be executed stably and with low latency in a low-interference environment.

[0062] (3) LSR (Latency Sensitive Reserved): LSR tasks need to reserve resources to obtain better accuracy, and adopt strategies such as CPU core binding to ensure execution efficiency. This corresponds to some of the Guaranteed type tasks native to Kubernetes. LSR tasks further emphasize the accuracy and stability of task execution on this basis. They not only reserve necessary resources, but also introduce stricter resource isolation strategies and priority scheduling in the scheduling process. The stricter resource isolation strategy means that the system will reserve some resources for these LSR tasks, and these resources will not be occupied by other tasks, so that when LSR tasks need them, they can use these resources quickly; the stricter priority scheduling means that the priority of these LSR tasks will always be in a higher priority sequence to ensure their priority execution, thereby achieving dual guarantees for task execution quality while maintaining low latency response.

[0063] (4) LS (Latency Sensitive): LS tasks need to share resources to ensure better resilience to burst traffic. This corresponds to deep learning inference workloads based on microservices, thereby achieving better resource elasticity and more flexible resource adjustment capabilities. It corresponds to some of the native Guaranteed and Burstable tasks of Kubernetes. For LS tasks that share resources, an elastic resource sharing mode is designed. The elastic resource sharing mode allocates more resources to tasks when system resources are sufficient; when resource contention is detected between tasks, resulting in performance degradation, or when user request traffic suddenly surges, the resources occupied by tasks will be reduced to reduce resource contention.

[0064] In this mode, LS tasks can share resources with other tasks during normal operation, but when the system detects burst traffic or performance degradation, the scheduling system can quickly adjust resource allocation and dynamically expand the resource supply of LS tasks to meet temporary high concurrency requirements. This design relies on real-time monitoring data and a closed-loop feedback mechanism to achieve precise resource control in colocation scenarios.

[0065] (5) BE (BestEffort): BE tasks have limited resource operation quality and may even be killed in extreme cases. They have the typical QoS level of batch jobs, stable computing throughput over a certain period of time, and low-cost resources, corresponding to the native BestEffort type tasks of Kubernetes. As for BE tasks, a best-effort strategy is adopted to make full use of idle resources in the cluster while ensuring that the resource requirements of high-priority tasks are not affected. Through dynamic detection and automatic recycling mechanisms, the system allows BE tasks to run when resources are abundant, but gives priority to reservation and migration when resources are tight, and even allows them to be terminated in extreme cases to maintain the balance of overall resource scheduling.

[0066] Overall, this design builds a fine-grained, multi-dimensional resource management model, which closely revolves around the characteristics of different QoS types, from task metadata expansion, scheduling strategy adjustment to real-time monitoring and dynamic feedback. Through deep integration with Koordlet components, the system can flexibly adopt dedicated, shared or reserved resource allocation strategies according to the different requirements of each QoS type during task scheduling, which not only fully utilizes idle resources in the cluster, but also ensures resource guarantees for high-priority tasks, thereby achieving high efficiency and flexibility in resource scheduling in a mixed deployment environment.

[0067] The present invention redesigns the service level quality (QoS) type and resource partitioning strategy, and effectively realizes the coordinated scheduling of tasks of different priorities in a mixed deployment environment through the Koordlet component. By making full use of the allocated but unused resources, the system not only ensures the resource requirements of high-priority tasks, but also improves the resource utilization of low-priority tasks, thereby optimizing the overall resource management and task processing performance of the cluster.

[0068] From system goals and module design to operation process, the system extends the native functions of Kubernetes to achieve efficient collaborative deployment of online and offline deep learning tasks, optimize resource utilization and ensure service quality. This architecture provides theoretical basis and practical support for resource management and task scheduling in data centers.

[0069] 2. Hybrid deployment optimization strategy for offline deep learning tasks.

[0070] In the hybrid deployment of offline deep learning tasks, the core of optimization is to redesign the resource sharing scheme to achieve efficient hybrid deployment of different types of tasks. Deep learning tasks are computationally and memory intensive, which makes reasonable scheduling and allocation particularly important in the scenario of shared resources. Therefore, in order to improve the overall resource utilization, it is necessary to design unique sharing strategies for resources of different dimensions such as CPU, memory and GPU to achieve flexible, elastic and efficient resource scheduling and hybrid deployment effects. The current completed research mainly optimizes from the perspective of CPU and memory, and verifies its effect through experiments, while the sharing scheme of GPU resources has made significant progress in theoretical construction. Next, we will introduce the CPUBurst strategy, memory over-issuance design, GPU time slice sharing strategy, and the hybrid deployment algorithm of offline deep learning tasks based on multi-resource perception.

[0071] 2.1 CPU Sharing Solution: CPUBurst Strategy In high-concurrency tasks, tail latency is one of the key issues that affect system performance. Tail latency refers to the phenomenon that the response time of some requests is significantly higher than the average response time in a high-concurrency environment. Usually, the focus is on the 99th or 99.9th percentile request latency, that is, the response time of the slowest 1% or 0.1% of requests. From the perspective of user experience, the existence of tail latency means that some users will face long waits, especially in large-scale online services and real-time applications. This long-tail phenomenon may significantly affect the overall user experience and system service quality.

[0072] The present invention designs a CPUBurst strategy to effectively manage the execution of tasks by optimizing time slice allocation and reducing tail delay. The CPUBurst strategy mainly optimizes resource allocation by dynamically adjusting the length of time slices and combining task characteristics to improve system performance. It can be specifically divided into the following parts: 2.1.1 CPUBurst time slice adjustment based on task history: In order to manage the execution time of different tasks more finely, the present invention designs a time slice adjustment strategy based on the task CPUBurst history record, such as Figure 3 As shown. The operating system dynamically adjusts the length of the time slice allocated to the task in the future according to the historical CPUBurst time of the task. For tasks with shorter CPUBurst, this project chooses to allocate shorter time slices, so that the task can be completed quickly without affecting the response efficiency of the system due to long waiting time; for tasks with longer CPUBurst, longer time slices are allocated to avoid the context switching overhead caused by frequent time slice switching.

[0073] In this way, the waiting time of short tasks and the context switching overhead of long tasks can be reduced at the same time, thereby effectively reducing tail latency. This strategy helps to reduce the queuing time of short tasks and prevents them from significantly delaying responses due to waiting for scheduling. At the same time, by reducing the number of context switches, the system's processing overhead is also reduced, allowing CPU resources to be used more efficiently.

[0074] 2.1.2 Multi-Level Feedback Queue (MLFQ) Scheduling Strategy: In order to further enhance the effect of the CPUBurst strategy, the present invention adopts a multi-level feedback queue (MLFQ) scheduling method. MLFQ is a highly adaptable scheduling strategy that can dynamically adjust the priority of a task according to its CPUBurst characteristics. In MLFQ, different tasks are assigned to different queues, and the queue priority is continuously adjusted according to the task execution status: Tasks enter the corresponding priority queue according to the initial CPUBurst time and the system scheduling policy. Shorter tasks are placed in the high-priority queue and use shorter time slices to ensure that they are completed quickly. Longer tasks are assigned to the low-priority queue and use longer time slices to reduce context switching. Through this multi-level feedback queue mechanism, it can ensure that shorter tasks are responded to quickly, thereby effectively reducing the overall tail delay; for longer tasks, longer time slices are allocated to avoid the overhead caused by excessive switching.

[0075] In the Kubernetes environment, the multi-level feedback queue (MLFQ) scheduling strategy is closely integrated with the native scheduler and container runtime to form an adaptive task scheduling mechanism. The entire process starts with task submission. When the user creates a Pod through the Kubernetes native interface, the system not only receives standard resource requests and service level information, but also obtains preliminary CPUBurst estimates through extended task metadata (such as annotations or labels). The scheduler first assigns the Pod to the initial feedback queue based on this estimated information: tasks with shorter expected CPUBurst are placed in a high-priority queue, while tasks with longer estimated times enter a low-priority queue, thereby ensuring that short tasks can get a quick response in the initial stage, and long tasks use longer time slices to reduce the context switching overhead caused by frequent scheduling.

[0076] When a Pod enters the scheduling queue, the extended scheduler module will give priority to tasks in the high-priority queue according to the priority order of the multi-level feedback queue. For each selected Pod, the system calculates the appropriate time slice through a custom scheduling algorithm, and schedules the task to a node with sufficient resources based on the real-time resource status and health monitoring data of the node. In this process, the Koordlet component plays a key role. It not only performs specific resource allocation operations, but is also responsible for converting scheduling decisions into Cgroups parameter adjustments for the underlying Docker container, ensuring that the container obtains CPU resources that match the allocated time slice at the operating system level.

[0077] During the running of a task, its actual CPU usage will be collected in real time and recorded in the historical database. This process relies on data feedback provided by monitoring tools such as KubernetesMetricsServer or Prometheus. When a task is completed or reaches the preset time slice upper limit during operation, the system will compare the difference between the actual CPUBurst of the task and the estimated value, and use algorithms such as exponential smoothing to update the task's historical record. If a task uses a time slice in its high-priority queue longer than expected, it means that its CPUBurst is longer. At this time, the system will automatically downgrade it to a lower priority queue; conversely, if the task is completed quickly within the allocated shorter time slice, it may be given a priority upgrade, so that it will continue to enjoy a higher priority in the next scheduling. Such dynamic adjustments not only ensure the fast response of short tasks, but also balance the resource usage of various tasks in a mixed deployment environment and reduce overall tail latency.

[0078] The entire MLFQ mechanism forms a closed-loop scheduling system through continuous monitoring, feedback, and queue adjustment. In the entire Kubernetes cluster, this mechanism can use Custom Resource Definitions (CRD) to record the CPUBurst history and current queue status of the task in the form of structured data, so that the extended scheduler can read this information in real time and make adjustments. Finally, when the task optimized by the multi-level feedback queue scheduling strategy runs on the node, it not only obtains resource allocation that matches its CPU characteristics, but also realizes the precise execution of time slice allocation through the underlying resource control of Docker and Cgroups, ensuring the efficiency and flexibility of the overall scheduling of the system.

[0079] 2.2 Memory sharing solution: memory over-issuance design.

[0080] In a Kubernetes cluster, each task needs to specify the resourcesrequest and resourceslimit values ​​for memory usage when submitting, which represent the minimum memory requirement and maximum allowed memory of the task respectively. Resourcesrequest is the main basis for the scheduler to decide which node to schedule the task to, ensuring that the task can obtain the minimum memory resources that meet its needs after being scheduled to the node. Resourceslimit is the maximum amount of memory allowed to be used by the task during operation, and is a means used by the system to limit excessive use of resources.

[0081] Although this resource management method based on requests and limits ensures the stable operation of tasks, it will face certain resource waste problems in practice. Specifically, if the amount of memory requested by a task is significantly higher than the memory required for actual runtime, the cluster scheduler will reserve this part of memory for the task based on the request value, even if these resources are not actually used. This will cause other tasks to be unable to use these idle memories, ultimately resulting in resource waste and reducing the overall resource utilization of the cluster.

[0082] The present invention designs a memory overcommitment strategy, such as Figure 4 As shown. Memory oversubscription allows the total amount of memory allocated to nodes in the cluster to exceed the total amount of their physical memory, which is the so-called "overselling" of memory resources. Through memory oversubscription, resource allocation can be managed and optimized more flexibly, maximizing the overall resource utilization of the cluster. The implementation of the memory oversubscription strategy depends on the flexible configuration of resource requests and limit values ​​in Kubernetes, as well as the management of different task priorities in combination with the Kubernetes QoS mechanism. The specific implementation methods can be divided into the following aspects: 2.2.1 Reasonable configuration of resource requests and limit values In order to effectively implement memory over-issuance in the cluster, the present invention needs to reasonably configure the resourcesrequest and resourceslimit of each task: For online deep learning tasks, higher resourcesrequest and resourceslimit should be set to ensure that these tasks can obtain sufficient resources to ensure service stability and performance. When scheduling these tasks, nodes with more abundant resources will be prioritized for deployment to reduce potential resource competition.

[0083] For offline deep learning tasks, you can set lower resourcesrequest and higher resourceslimit. These tasks can use more memory resources when cluster resources are abundant, but will be restricted or evicted first when resources are tight, so as to ensure the normal operation of key tasks.

[0084] Through reasonable resource configuration, we can ensure that the cluster can effectively allocate memory under high load conditions, reduce resource waste, and maximize resource utilization.

[0085] 2.2.2 Task Management Based on QoS Category In order to better achieve memory over-issuance in the Kubernetes cluster, the custom QoS categories introduced in this topic are combined to classify and manage different tasks. The most appropriate memory resources can be allocated to the tasks according to their characteristics, achieving more flexible and efficient resource management.

[0086] (1) SYSTEM: Resource management of system tasks In the memory over-issuance design, SYSTEM type tasks have higher priority. The system will reserve enough memory for these tasks to ensure stable operation, but will limit them to prevent them from occupying too many node resources.

[0087] By setting appropriate resourceslimit for SYSTEM tasks, the set resource value should be able to support SYSTEM tasks to meet the system response time, which can ensure that the operation of these services is within the control range, that is, the response time can meet the system operation requirements, and other tasks will not be unable to run due to resource competition. At the same time, such a design also ensures the reliability of the system's basic services.

[0088] (2) LSE (LatencySensitiveExclusive): Latency-sensitive exclusive tasks LSE (LatencySensitiveExclusive) tasks are deep learning inference tasks that have extremely high latency requirements. This type of task needs to ensure resource exclusivity, that is, it cannot be shared with other similar tasks in the allocated resource pool to ensure resource stability and low latency performance. Therefore, LSE tasks are generally allocated to independent resource pools for deployment. When applying for resources, LSE tasks require exclusive CPU and memory resources and are not allowed to share with other tasks. This resource allocation method ensures task stability and high performance, especially for real-time inference scenarios, LSE can significantly reduce tail latency.

[0089] Although LSE tasks require exclusive resources, since usually not all allocated memory is fully used during the entire task execution, the memory over-issuance strategy can appropriately reduce the memory limit of LSE tasks to release part of the memory for temporary use by other non-critical tasks.

[0090] (3) LSR (LatencySensitiveReserved): Latency-sensitive reserved tasks LSR (LatencySensitiveReserved) tasks are mainly for deep learning tasks that require high execution efficiency and resource reservation. These tasks usually use the CPU core binding method to ensure the execution efficiency of the task. LSR tasks also have very high requirements for memory and CPU resources, especially the need to ensure sufficient resources to obtain higher inference accuracy.

[0091] LSR tasks need to have resources reserved for them, such as CPU core binding (CPUPinning) to ensure that they exclusively occupy certain cores. This approach can significantly improve the execution efficiency of tasks and reduce the performance overhead caused by context switching. At the same time, in order to ensure that LSR tasks can obtain sufficient memory, LSR tasks are allocated in a memory pool with a higher priority to ensure that their memory will not be preempted by other tasks even in the case of memory over-issuance. In addition, the memory usage pattern of LSR tasks will also be monitored to avoid excessive retention when resources are tight, which will affect the overall performance of the cluster.

[0092] (4) LS (LatencySensitive): delay-sensitive tasks LS type (LatencySensitive) tasks also have high latency requirements, but unlike LSE and LSR, the resources of LS tasks can be shared with other tasks within a certain range to achieve better resource elasticity and dynamic adjustment capabilities. This type of task is mainly used for deep learning inference workloads based on microservices. Its goal is to achieve efficient resource sharing and elastic management while ensuring response speed.

[0093] The resource scheduling strategy of LS tasks allows them to obtain higher resource utilization rights when cluster resources are abundant, and to share resources when resources are tight, thereby achieving better elastic management effects. For example, when there are more idle resources in the system, LS tasks can dynamically expand their resource requests to obtain more memory and computing resources. In the design of memory over-issuance, LS tasks can occupy unused memory of other tasks when the load is low. This mechanism provides good elasticity for LS tasks. When the load increases, LS tasks will automatically reduce memory usage and release more memory to LSE or LSR tasks with higher priority, thereby achieving overall memory utilization optimization.

[0094] (5) BE (BestEffort): Best Effort Task BE type (BestEffort) tasks correspond to deep learning training tasks or batch processing tasks of some models. These tasks can be killed when resources are tight, and usually have lower requirements for latency and resource guarantees. The goal of BE type tasks is to provide a certain amount of computing throughput at the lowest cost, so they have the lowest priority in the cluster.

[0095] BE tasks do not set specific resource requests and limits, so they are low-priority tasks in memory allocation, and the surplus memory in the memory over-issuance policy will be allocated to BE tasks first. When cluster resources are tight, these tasks will be affected first, including being evicted by the system to ensure that other higher-priority tasks can run stably. Since BE tasks do not have strict requirements on resources, the scheduler can flexibly schedule BE tasks for execution during idle time, thereby improving the overall throughput of the system without affecting critical tasks. These tasks usually use the remaining resources in the cluster at a low cost and can accept being suspended or killed in extreme cases.

[0096] 2.3 GPU Sharing Solution: GPU Time Slice Sharing Strategy GPU is an indispensable computing resource in deep learning tasks, especially in scenarios that require efficient reasoning and model training, where its video memory and computing power directly determine the performance and resource utilization of the task. However, since the video memory resources of GPU are fixed and deep learning tasks have a large demand for video memory, how to efficiently share video memory resources among multiple tasks becomes a key challenge for hybrid deployment. In order to solve this problem, the present invention proposes a GPU video memory sharing strategy based on NVIDIA MIG (Multi-Instance GPU) technology, which realizes on-demand allocation of video memory for online reasoning tasks and offline training tasks through the combination of hardware isolation and Kubernetes scheduling.

[0097] The architecture of the GPU memory sharing strategy is as follows Figure 5As shown in the figure, it shows how to use MIG technology to partition the GPU memory and register these partitions as independent GPU instances in the Kubernetes cluster. On this basis, the cluster dynamically allocates GPU memory resources to different tasks through the scheduler. Online inference tasks usually need to obtain memory resources first to ensure real-time responsiveness; while offline training tasks use the remaining memory resources for batch calculations to improve overall resource utilization. The scheduling system monitors the GPU memory usage and task load in real time, and dynamically adjusts the memory allocation strategy to achieve efficient sharing of memory resources among different tasks.

[0098] MIG technology plays a core role in this strategy. Its principle is to divide the physical GPU memory into multiple logical partitions, each partition corresponds to a MIG instance. Each instance has independent computing cores, memory, cache, and bandwidth resources to ensure isolation between tasks. Combined with Kubernetes' resource management capabilities, each MIG instance is exposed as an independent GPU resource type (such as nvidia.com / mig-1g.5gb), allowing users to explicitly request the required memory partition in the task description. For example, an online inference task can request a small MIG instance (such as 5GB of memory), while an offline training task requests a larger MIG instance (such as 20GB of memory). This memory management method based on hardware partitioning not only avoids resource contention between tasks, but also provides support for the refined scheduling of memory resources.

[0099] In the memory sharing strategy, the design of the scheduling system is the key, especially how to dynamically adjust the memory allocation according to the task load. The scheduling system uses the scalability of Kubernetes to introduce a priority queue and dynamic allocation mechanism to assign different types of tasks to different queues. For example, online inference tasks are assigned to high-priority queues, while offline training tasks are assigned to low-priority queues. The scheduler dynamically adjusts the allocation of memory partitions according to the real-time needs of the tasks. For example, during the peak period of inference requests, the system prioritizes allocating more memory resources to high-priority queues to ensure low-latency response of online tasks. During periods of low request volume, more memory resources are allocated to offline training tasks to improve the efficiency of model updates.

[0100] In order to verify the effectiveness of the video memory sharing strategy, the present invention designed a variety of experimental scenarios, including mixed deployment scenarios of online reasoning tasks and offline training tasks. In these scenarios, the GPU video memory utilization rate is significantly improved, the response delay of online tasks is reduced, and the training throughput of offline tasks is also guaranteed. Especially in dynamic load environments, the strategy shows excellent adaptability, and effectively balances the system performance and resource utilization by flexibly allocating video memory resources.

[0101] The GPU memory sharing strategy not only provides a new solution for GPU resource management in hybrid deployment, but also lays the foundation for multi-tenant GPU resource sharing in cloud computing scenarios. Through the combination of hardware isolation and dynamic scheduling, this strategy achieves a good balance between resource utilization, task isolation, and scheduling flexibility, providing strong support for the efficient deployment of deep learning tasks.

[0102] 2.4 Hybrid deployment algorithm for offline deep learning tasks based on multi-resource awareness.

[0103] Based on the above problems in this chapter, this paper proposes a hybrid deployment algorithm for offline deep learning tasks based on multi-resource perception. The hybrid deployment algorithm aims to dynamically allocate resources in the Kubernetes cluster to meet the low latency requirements of online reasoning tasks and the efficient completion requirements of offline training tasks, while maximizing resource utilization. The algorithm starts with initialization and monitoring, initializing the resource information of all nodes in the cluster, including the capacity of CPU, memory, and GPU. The system monitors the resource usage of each node in real time and calculates the load status of the node to ensure the effective use of resources.

[0104] In the task classification and priority assignment stage, the algorithm assigns priorities to tasks according to the task type (online reasoning task or offline training task), and sorts the tasks using the priority function to ensure that online tasks have a higher priority in scheduling. Subsequently, the algorithm enters the resource allocation and scheduling stage. For online reasoning tasks, the system prioritizes allocating resources to them and searches for nodes that meet resource requirements. If a suitable node is found, the task will be scheduled to the node and the resource status of the node will be updated. If the node resources are insufficient, the system will temporarily wait for the resources to be released. For offline training tasks, the system schedules the task to a suitable node while meeting the resource constraints and updates the node's resource information. When node resources are insufficient, the algorithm introduces a dynamic adjustment mechanism to recycle resources from low-priority tasks and reallocate them to online tasks to ensure that the performance requirements of online tasks are prioritized.

[0105] The algorithm's dynamic adjustment and iterative optimization mechanism monitors the node's resource usage and task completion status in real time, and adjusts the scheduling strategy in a timely manner when resources are insufficient or demand changes. This mechanism further optimizes the task's resource allocation strategy through elastic resource allocation and capacity borrowing, thereby achieving efficient resource management in a hybrid deployment environment and improving the overall performance of the cluster and task completion efficiency. The pseudo code of the algorithm is shown below.

[0106] Some symbols in the following pseudo code are explained as follows: t j Indicates j tasks; C iIndicates i CPU resources of each node; M i Indicates i Memory resources of each node; G i Indicates i GPU resources of each node; Indicates j The CPU resource requirements of each task; Indicates j Memory resources for each task; Indicates j GPU resources for each task; n i Indicates i nodes; A [ t j ] indicates the j Node allocation of tasks; R Indicates changes in node resources; Representation Task j The current resources of the node; Representation Task l The current resources of the node; Representation Task k The current resources of the node.

[0107]

[0108] According to the algorithm, the time complexity of the algorithm is mainly composed of several key steps. First, the operation of initializing and monitoring each node takes O(N) time, where O represents the time complexity of the algorithm and N is the total number of nodes. Secondly, the step of prioritizing all online and offline tasks takes O(TlogT) time, where T is the total number of tasks. The resource allocation and scheduling part involves finding suitable resources for each online task in N nodes, which takes O(TlogT). online ×N) time. The scheduling of offline tasks may also include dynamic resource adjustment, which will further increase the time complexity when frequent adjustments are required. Therefore, the overall time complexity is O(TlogT+T online ×N).

[0109] The space complexity is mainly composed of two parts. The algorithm needs to store the resource status R for each node i , and store the scheduling status A for each task. The space requirement of this part is O(N+T), where N is the number of nodes and T is the number of tasks, and it is mainly used to store the information of nodes and tasks.

[0110] 3. Cluster deployment performance experiment.

[0111] In order to verify the effectiveness of the CPUBurst strategy, memory over-issuance design, and GPU sharing design in a real environment, the present invention conducted a cluster deployment performance experiment on the built Kubernetes platform, and implemented scheduling optimization through the functional components developed above. The main purpose of the experiment is to test the resource scheduling effect of the scheduling strategy on different types of deep learning tasks in a mixed deployment environment, and to evaluate the improvement of the overall performance of the cluster. The performance of the platform is evaluated experimentally and compared with traditional static scheduling strategies and other dynamic scheduling strategies to measure its advantages in resource utilization, job completion time, scheduling delay, etc.

[0112] In the colocation cluster deployment performance experiment, the present invention adopts a systematic experimental data collection method to comprehensively evaluate the performance of the cluster and the task processing efficiency. In the experiment, four different types of deep learning tasks were deployed in the cluster, and each task represented a different load situation. To ensure the independence and controllability of the tasks, all tasks were deployed through containerization technology, and Prometheus was used as a cluster monitoring tool for real-time data collection and monitoring.

[0113] First, in terms of resource utilization, the present invention monitors the CPU, memory, storage and network resource usage of cluster nodes, and calculates the resource utilization of the cluster under different loads to evaluate the resource allocation efficiency and load carrying capacity of the cluster. Secondly, the task queuing time and the average job completion time are recorded through the log of the cluster scheduling system, including the queuing time from task submission to start of execution, and the total time required for the task from start to completion. These data are helpful for analyzing the efficiency of cluster scheduling and the delay of task execution.

[0114] In terms of reasoning task performance, the present invention evaluates the throughput of the cluster when processing reasoning tasks by calculating the number of reasoning requests completed per hour. Finally, in order to evaluate the service level of the mixed cluster, the present invention conducts a comprehensive analysis of the performance of the cluster under different loads, focusing on key indicators such as the average resource utilization, task completion time, throughput per unit time, and task completion rate. These data help the present invention evaluate the stability of the cluster under high load, the reliability of task scheduling, and the overall performance. All collected data are statistically analyzed and displayed in the form of charts, so as to intuitively present the performance changes and resource allocation of the cluster under different loads.

[0115] The experiment aims to evaluate the performance of the system designed by the present invention in a hybrid deployment environment and compare it with other existing scheduling strategies. To this end, the present invention selects two representative scheduling strategies, Lyra and Pollux, as comparison algorithms. The main indicators of comparison include job completion time, GPU utilization, job waiting time, etc. The selection of these indicators is intended to comprehensively evaluate the performance and advantages of the system proposed by the present invention, especially its flexibility and efficiency in the face of multi-task hybrid deployment and dynamic resource requirements.

[0116] Lyra is an elastic scheduling strategy for deep learning clusters, which aims to improve the throughput and efficiency of the cluster by dynamically adjusting resource allocation. Lyra flexibly adjusts the resource allocation of tasks based on the priority of the task and the real-time resource requirements. The core idea is to use the elastic resource allocation mechanism to dynamically adjust the GPU allocation according to the computing requirements and progress of the task, so that cluster resources can be more efficiently scheduled and utilized according to the needs of the task. Lyra is mainly aimed at training tasks and is suitable for compute-intensive deep learning workloads. It can improve the resource utilization of the cluster to a certain extent and reduce the queuing time of tasks.

[0117] Pollux is a scheduling strategy that optimizes cluster throughput, maximizing the overall performance of the cluster by jointly adapting cluster resources and task requirements. Pollux combines resource sharing with the flexibility of task scheduling, dynamically allocating resources between multiple tasks to maximize throughput. Unlike Lyra, Pollux focuses more on the overall throughput of the cluster and the long-term efficient use of resources, rather than simply optimizing the completion time of tasks. Pollux uses a mechanism based on task load prediction and dynamic resource allocation, which can further improve the overall efficiency of the system while ensuring the completion time of jobs.

[0118] 3.1 Evaluation of resource utilization of colocation clusters

[0119] The resource utilization performance of the hybrid cluster under different load conditions is analyzed to verify the adaptability and stability of the system of the present invention under different workload conditions. The experiment evaluated the GPU utilization of the hybrid cluster under high-load scenarios and the overall resource stability of the cluster. These evaluations help to understand whether the scheduling strategy of the hybrid system can efficiently allocate and manage resources in a complex hybrid deployment environment. The experiment uses multiple evaluation indicators to quantify the resource utilization, including CPU, memory and GPU training task resource utilization, system average resource utilization, and resource stability under load. Resource utilization is defined as the ratio of the resource occupied by tasks over a period of time, which can reflect the efficiency of the scheduling strategy in allocating computing resources. In the experiment, the workloads used include Bert, GNMT-16, ResNet-50, VGG and other models. These models have different computing characteristics and resource requirements, thereby effectively simulating the load of the cluster in a production environment. The task deployment strategy of the hybrid system is compared with the scheduling strategies of Lyra and Pollux to evaluate their node resource utilization performance under high load. The experimental results are shown in the figure. Figure 6 As shown, Figure 6 a) in the figure is the ResNet50 training task and the system resource utilization under hybrid deployment. Figure 6 b) in the figure is the utilization rate of system resources under Bert training task and hybrid deployment. Figure 6 c) in the figure is the system resource utilization rate under GNMT training task and hybrid operation. Figure 6 d) in the figure is the system resource utilization under VGG training tasks and hybrid deployment.

[0120] The results show that the hybrid system of the present invention has a GPU utilization rate of 9% higher than that of the traditional FIFO baseline strategy in high-load scenarios, reaching 81%; under heavy and low load conditions, the hybrid system of the present invention also shows good stability, with GPU utilization maintained between 70-75%, while Pollux and Lyra have GPU utilization rates of about 69-72% under similar load conditions. This shows that the hybrid system of the present invention can effectively improve the overall resource utilization of the cluster and reduce the idle time of resources through elastic expansion and capacity borrowing strategies, while maintaining stable system operation.

[0121] 3.2 Performance Evaluation of Colocation Cluster Tasks

[0122] A detailed evaluation of the task performance in the colocation cluster was conducted, focusing on the average job completion time (JobCompletionTime, JCT), task queuing time, and delay performance of different task types under different scheduling strategies. In order to better understand the performance of each scheduling strategy in different task scenarios, the present invention uses the median, mean, and 95% quantile as core statistical indicators. These indicators can fully reflect the performance of the scheduling strategy in terms of average task completion time, tail latency, and system throughput. The experimental results are shown in Table 1.

[0123] Table 1. Task queuing time and average job completion time at different depths in a colocation cluster

[0124] In the ResNet-50 task, the average JCT of the system proposed in the present invention is 12875 milliseconds, which is a significant decrease compared to the 20872 milliseconds of the traditional FIFO baseline. Experiments show that by using elastic expansion mechanism and capacity borrowing strategy, the system proposed in the present invention can dynamically adjust GPU allocation, so that under high load conditions, training tasks can obtain more GPU resources, thereby significantly shortening the task completion time. Compared with Lyra and Pollux, the system proposed in the present invention also shows certain advantages. The average JCT of Lyra and Pollux are 13653 milliseconds and 12983 milliseconds respectively, while the average JCT of the system proposed in the present invention is 12875 milliseconds. Especially in high concurrency and complex task load scenarios, the system proposed in the present invention can reasonably allocate resources by adjusting the scheduling strategy in real time to reduce the waiting time and completion time of tasks.

[0125] In order to further analyze the impact of different scheduling strategies on tail latency, the present invention also calculated the 95% quantile of JCT to measure the performance of the system under extreme load conditions. For the ResNet-50 task, the JCT95% quantile of the system built by the present invention is 62093 seconds, which is a significant reduction compared to the 85972 milliseconds of the FIFO baseline. The JCT95% quantiles of Lyra and Pollux are 65392 milliseconds and 63986 milliseconds, respectively, showing the obvious advantages of the system of the present invention in tail latency control. This is because the system of the present invention can flexibly adjust resource allocation according to the priority of the current task and the idle status of system resources during the scheduling process, thereby effectively avoiding resource contention under high load conditions. In terms of system throughput, taking the online inference task of the ResNet50 model as an example, the present invention selects the number of online inference tasks completed within 1 hour as the evaluation indicator. The method of the present invention shows higher stability and throughput under high load. The experimental results are as follows. Figure 7As shown in the figure, the overall throughput of the system of the present invention under heavy load scenarios is still superior to that of Lyra and Pollux, especially when multiple high-priority tasks arrive at the same time, the system of the present invention can prioritize the resources of the reasoning task, quickly complete the job, and flexibly lend the resources to subsequent training tasks. This elastic expansion and priority adjustment mechanism enables the system to process more tasks under high load, thereby improving the overall throughput of the system.

[0126] 3.3 Performance evaluation of service level division in colocation clusters.

[0127] This experiment conducted a comprehensive performance evaluation of four service level tasks (LSE, LSR, LS, BE) in a colocation cluster. The experiment used ResNet-50 as a typical deep learning model inference task, and distinguished service levels in combination with different model accuracy requirements. Specifically, high-precision inference tasks were set to the LSE type to meet the requirements of latency sensitivity and resource exclusivity; medium-precision inference tasks were assigned to the LSR type, and the execution efficiency of tasks was guaranteed through resource reservation strategies; low-precision inference tasks were classified as the LS type, allowing resource sharing to enhance elasticity and flexibility. In addition, batch training tasks were used as the BE type to use the remaining resources in the cluster to perform low-priority computing tasks. Four different load modes were designed in the experiment to simulate typical usage scenarios in a colocation cluster. These load modes include low load (LowLoad), medium load (MediumLoad), high load (HighLoad), and burst load (BurstLoad). The low-load mode represents a situation where cluster resources are abundant. At this time, high-priority tasks can allocate resources completely according to demand, and low-priority tasks can also use idle resources to run efficiently. In the medium-load mode, resource demand gradually approaches the cluster capacity, and resource competition begins to appear. The high-load mode simulates a scenario where cluster resources are close to saturation. High-priority tasks still need to guarantee performance, while low-priority tasks may face delays or preemption. The burst load mode focuses on the impact of burst traffic on the system. Burst high-priority tasks need to quickly obtain resource support, while low-priority tasks may be suspended or interrupted. These four modes can fully reflect the task performance of the colocation cluster under different load conditions. The experimental results are as follows: Figure 8 As shown, Figure 8 a) in the figure is the average resource utilization of ResNet50 under different loads. Figure 8 b) in the figure is the average task delay of ResNet50 under different loads. Figure 8 c) in the figure is the average throughput of ResNet50 under different loads. Figure 8 d) in the figure is the average completion rate of mixed tasks per unit time of ResNet50 under different loads.

[0128] The experimental results show that tasks of different service levels show significant differences in indicators such as resource utilization, latency, throughput and task completion rate. In terms of average resource utilization, BE and LS tasks can make full use of unallocated idle resources in the cluster, and their utilization rate increases significantly with the increase of load, while LSE and LSR tasks maintain a relatively stable resource occupancy rate, which reflects the effect of resource exclusivity and reservation strategies in high-priority tasks. In terms of latency performance, LSE tasks have the lowest latency in all load scenarios and can effectively meet their strict real-time requirements; LSR has a slightly higher latency but is still stable; in contrast, the latency of LS and BE tasks increases significantly with the increase of load, especially in burst traffic scenarios. In terms of throughput, LSE and LSR tasks show strong stability in high load and burst traffic scenarios, while the throughput of LS and BE decreases significantly with the intensification of resource competition, reflecting their sensitivity to dynamic changes in resources. In terms of task completion rate, LSE and LSR tasks always maintain a completion rate close to 100%, showing their priority and reliability in the scheduling system; however, the completion rate of BE tasks drops significantly as the load increases, and can be as low as 60% under high load and burst traffic scenarios. Overall, the experimental results verify the effectiveness of the redesigned service level division strategy, which can not only guarantee the performance requirements of high-priority tasks, but also improve the overall efficiency of the co-location cluster by making full use of system resources. This result provides strong support for task scheduling optimization in hybrid deployment scenarios.

[0129] The main advantages of the present invention include the following aspects: 1) Improve resource utilization: By improving the resource sharing and scheduling mechanism of CPU, memory and GPU, the idle rate of resources such as CPU, memory and GPU can be significantly reduced, and the overall computing efficiency of the cluster can be improved.

[0130] 2) Optimize task execution performance: By introducing a load-aware dynamic scheduling mechanism, the task queuing time and completion time can be effectively shortened, the system throughput per unit time can be improved, and at the same time, the task response time can be ensured to meet service requirements.

[0131] This invention not only provides theoretical and practical support for the intelligent hybrid deployment of deep learning tasks, but also provides important reference value for the sustainable development and operation cost optimization of modern data centers.

[0132] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A hybrid deployment method for deep learning tasks, characterized in that: include: S1. The user submits a task through the Kubernetes native interface; S2. The task scheduling optimizer analyzes the tasks submitted by users, extracts key information such as task type, resource requirements, priority, service level, and data locality, and evaluates the resource consumption and execution time of the tasks based on historical data and the current system status. It uses a multi-dimensional scoring model to evaluate each node based on the real-time resource usage and load status of the nodes, and finally determines one or more target nodes suitable for task operation. S3. The SLO manager tracks the real-time data of various performance indicators of the system and compares the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimal scheduling strategy; S4. After determining the node, the scheduling strategy is sent to the Koordlet component. The Koordlet component performs specific resource allocation operations on the target node, continuously monitors the node resource usage, and feeds back the node resource usage and load status after the task is executed to the task scheduling optimizer; S5. The traffic security monitor monitors the data flow between nodes in real time. By comparing it with the historical normal status data, when an abnormal traffic pattern or sudden traffic surge is identified, the detected abnormal information will be fed back to the SLO manager in real time.

2. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: In step S2, the task scheduling optimizer execution process also includes a rescheduling component optimizing the task distribution process according to the node load status: During the task execution process, the rescheduling component dynamically evaluates the current task distribution based on the node load status and task performance indicators collected in real time. If it is found that a node is overloaded or the task is running inefficiently, the rescheduling component triggers the rescheduling process and migrates some tasks to nodes with sufficient resources and lower load.

3. The deep learning task hybrid deployment method according to claim 1, characterized in that: In step S2, service levels are divided for different tasks, and an association between service levels and task scheduling priorities is established. Online reasoning tasks are assigned high service levels, and offline training tasks are assigned low service levels. When there is competition for resources in the task scheduling optimizer, online reasoning tasks with high service levels are executed first.

4. The deep learning task hybrid deployment method according to claim 1, characterized in that: In step S2, the service level quality type of the task includes: SYSTEM: In the system process, strict resource constraints are designed based on the required resource usage corresponding to the minimum latency of the SYSTEM response, to ensure that the system process obtains the necessary resources to ensure the latency response while not occupying too many resources; LSE: Develop dedicated resource pools and resource reservation machines to ensure that these tasks can be executed stably and with low latency in a low-interference environment; LSR: Introduces stricter resource isolation policies and priority scheduling. The resource isolation policy reserves necessary resources for these LSR tasks, and the priority scheduling ensures that the priority of LSR tasks is always in a higher priority sequence. LS: Adopts elastic resource sharing mode. When system resources are sufficient, more resources are allocated to tasks. When resource contention between tasks is detected or user request traffic suddenly surges, the occupied resources of tasks are reduced. BE: adopts a best-effort strategy. The best-effort strategy is to make full use of idle resources in the cluster while ensuring that the resource requirements of high-priority tasks are not affected.

5. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: The optimized scheduling strategy includes a CPUBurst strategy, which specifically includes: a1. CPUBurst time slice adjustment based on task history: Dynamically adjust the time slice length allocated to the task in the future according to the historical CPUBurst time of the task. For tasks with shorter CPUBurst, a shorter time slice is allocated, and for tasks with longer CPUBurst, a longer time slice is allocated; a2. Multi-level feedback queue scheduling strategy: Different tasks are assigned to different queues, and the queue priority is continuously adjusted according to the task execution status. Tasks enter the corresponding priority queue according to the initial CPUBurst time and scheduling strategy. Shorter tasks are placed in the high-priority queue and use shorter time slices, while longer tasks are assigned to the low-priority queue and use longer time slices.

6. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: The optimization scheduling strategy includes a memory over-issuance design strategy, and the memory over-issuance design strategy specifically includes: b1. Reasonable configuration of resource request and limit values: Configure resourcesrequest and resourceslimit for each task. For online deep learning tasks, set higher resourcesrequest and resourceslimit. When scheduling, online deep learning tasks will give priority to nodes with abundant resources for deployment. For offline deep learning tasks, set lower resourcesrequest and higher resourceslimit. When cluster resources are abundant, offline deep learning tasks will use more memory resources and will be restricted or evicted first when resources are scarce. b2. Task management based on QoS categories: Different tasks are classified and managed in the Kubernetes cluster based on QoS categories: b21. Resource management of system tasks SYSTEM: Set appropriate resourceslimit for SYSTEM type tasks. The set resource value supports SYSTEM tasks to meet the system response time to ensure that SYSTEM type tasks run within the control range; b22. Delay-sensitive exclusive tasks LSE: Reduce the memory limit of LSE type tasks to release some memory for temporary use by other non-critical tasks; b23. Latency-sensitive reserved tasks (LSR): LSR-type tasks are bound to CPU cores, assigned to memory pools with higher priorities, and the memory usage patterns of LSR-type tasks are monitored; b24. Latency-sensitive tasks LS: Allow LS-type tasks to obtain higher resource utilization rights when cluster resources are abundant, and share resources when resources are scarce; occupy memory not used by other tasks when the load is low, and automatically reduce memory usage when the load increases; b25. Best-effort task BE: No specific resource request and limit values ​​are set.

7. The method for hybrid deployment of deep learning tasks according to claim 1, characterized in that: The optimization scheduling strategy includes a GPU time slice sharing strategy, and the GPU time slice sharing strategy specifically includes: c1. Use MIG technology to partition the GPU and register each partition as an independent device to the cluster through DevicePlugin; c2. When a user submits a GPU task, the scheduler makes a decision based on the task requirements and node resource status. When the Pod is started, container-toolkit is responsible for mapping the specified MIG instance to the container and setting the corresponding environment to ensure that the container can only access the predetermined GPU resources. c3. Through continuous monitoring and closed-loop feedback mechanisms, the graphics memory allocation is dynamically adjusted according to the real-time load, and GPU resources are efficiently shared in hybrid deployment scenarios.

8. The deep learning task hybrid deployment method according to claim 1, characterized in that: The optimization scheduling strategy includes a multi-resource-aware scheduling strategy, which specifically includes: d1. Input online reasoning task set T online , offline training task set T office , node resource status R i ( i =1,…, N ), task priority function P ( t j );in, t j Indicates j tasks; d2. Initialization and monitoring: Initialize all nodes R i =( C i , M i , G i ), monitor the resource usage of all nodes in real time and calculate the load status; C i Indicates i The CPU resources of each node, M i Indicates i The memory resources of each node, G i Indicates i GPU resources of each node; d3.Task classification and priority allocation: Combine tasks T online ∪ T office ,according to P ( t j ) sorted by priority; d4. Resource allocation and scheduling: For online reasoning tasks t j ∈ T online , find available nodes i , meeting resource requirements , , If a suitable node is found, the task will be scheduled to the node and the resource status of the node will be updated. A [ t j ]= n i , the task t j Assign to Node n i , update node resources , if the node resources are insufficient, wait for the resources to be released; among them, Indicates j The CPU resource requirements of each task; Indicates j Memory resources for each task; Indicates j GPU resources for each task; n i Indicates i nodes; A [ t j ] indicates the j Node allocation of tasks; For offline training tasks, t j ∈ T office , if the node has enough resources, schedule the task t j Go to the node that meets the constraints and update the node resources When node resources are insufficient, a dynamic adjustment mechanism is introduced to calculate ∆ R Reclaim resources from low-priority tasks, update , , reallocate resources to online reasoning tasks; where ∆ R Indicates changes in node resources. Indicates the task j The current resources of the node. Indicates the task l The current resources of the node. Indicates the task k The current resources of the node; d5. Dynamic adjustment and iterative optimization: monitor resource utilization and task completion status in real time, dynamically adjust resource allocation according to node load, optimize scheduling strategy, and finally output task-to-node deployment plan mapping A as the task scheduling plan.

9. A deep learning task hybrid deployment system, characterized in that: include: The task scheduling optimizer is used to parse the tasks submitted by users, extract key information such as task type, resource requirements, priority, service level, and data locality, and evaluate the resource consumption and execution time of the task based on historical data and current system status. Combined with the real-time resource usage and load status of the node, a multi-dimensional scoring model is used to evaluate each node, and finally one or more target nodes suitable for task operation are determined; The SLO manager is used to track the real-time data of various performance indicators of the system and compare the real-time data with the pre-set service level objectives. If the real-time data of the performance indicator of a task deviates, the dynamic resource scheduling process is started to reallocate resources by calling the optimized scheduling strategy; Koordlet component is used to execute specific resource allocation operations issued by the scheduling policy on the target node, continuously monitor the node resource usage, and feed back the node resource usage and load status after the task is executed to the task scheduling optimizer; The traffic security monitor is used to monitor the data flow between nodes in real time. By comparing with the historical normal status data, when abnormal traffic patterns or sudden traffic surges are identified, the detected abnormal information will be fed back to the SLO manager in real time.

10. The deep learning task hybrid deployment system according to claim 9, characterized in that: The SLO manager includes: The colocation configuration component is used by users to perform simple colocation configuration. Resource over-issuance management component, used to dynamically adjust the resource over-issuance ratio according to the real-time status of the node; Workload statistics component, used to analyze and predict resource requirements using histograms and optimize load distribution; The task scheduling optimizer comprises: The colocation scheduler component is used to balance node loads based on load-aware policies, provide fine-grained resource isolation policies for different tasks, and support dynamic allocation of heterogeneous resources. The rescheduling component is used to dynamically evaluate the current task distribution based on the node load status and task performance indicators collected in real time to readjust the resource allocation strategy.

Citation Information

Patent Citations

  • Online training-oriented computing power resource elastic distribution system

    CN119166278A

  • High throughput cloud computing resource recovery system

    WO2023015787A1

Cited By

  • Burst load-oriented on-line and off-line task hybrid deployment method

    CN117234688A

  • Off-line task hybrid deployment method for burst load

    CN117234688B

  • Heterogeneous GPU resource pooling method and system combining Ray and Volcano

    CN120276872A

  • Data processing resource dynamic allocation method for data-in-data station

    CN120371544A

  • Inference service system

    CN120450053A