Multi-cloud GPU (Graphics Processing Unit) task scheduling method for resource-limited scene
By using a two-layer queue structure and real-time resource monitoring method in multi-cloud GPU task scheduling, the queue length and task allocation are dynamically adjusted, and the problems of unbalanced resource utilization and poor system stability are solved, and the resource utilization optimization and system stability are achieved.
Patent Information
- Application Number
- CN202510120642.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multi-cloud GPU task scheduling solutions are difficult to achieve optimal scheduling effects in the unbalanced resource utilization and poor system stability, especially in resource-constrained scenarios.
The scheduling method of a two-layer queue structure is adopted to monitor the resource status of each cloud platform in real time, and dynamically adjust the queue length to achieve unified management and flexible distribution of tasks. This method includes real-time resource monitoring mechanism, task execution statistics, optimal queue length calculation, task distribution control, resource protection mechanism and task priority processing.
The optimization of resource utilization is achieved, the problem of resource overload is avoided, the stability and reliability of the system are improved, and the differentiated needs of different tasks are met.
Smart Images

Figure BDA0005258847930000051 
Figure BDA0005258847930000061 
Figure BDA0005258847930000062
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing resource scheduling, and particularly to a GPU computing resource scheduling method for a multi-cloud platform. Background Art
[0002] With the rapid development of artificial intelligence and deep learning technologies, the demand for GPU computing resources in scenarios such as model training and inference is increasing day by day. To meet the growing computing requirements, enterprises usually choose to use GPU resources provided by multiple cloud platforms. In practical applications, how to efficiently manage and schedule GPU resources of multi-cloud platforms to achieve the maximum utilization of resources is of great significance in the field of artificial intelligence applications.
[0003] Currently, two main solutions are adopted to manage GPU resources of multi-cloud platforms: one is to independently configure a task scheduling system for each cloud platform to manage the GPU resources and task queues of its own platform respectively; the other is to adopt a centralized scheduling system to uniformly manage the GPU resources of all cloud platforms. The former is simple to operate and has low implementation difficulty; the latter can achieve unified allocation of resources and improve the overall utilization efficiency.
[0004] However, these existing solutions have obvious deficiencies in practical applications. The independent scheduling solution leads to scattered resources, making it difficult to achieve cross-platform load balancing, and there are significant differences in resource utilization rates among platforms; although the centralized scheduling solution can uniformly manage resources, it fails to fully consider the limiting factors of GPU resources of each cloud platform, easily causing resource overload on some platforms and affecting system stability. Especially in resource-constrained situations, these problems are more prominent.
[0005] To solve the above problems, some studies have proposed a dynamic scheduling solution based on resource prediction, attempting to optimize task allocation by predicting the resource availability of each platform. However, such solutions often rely too much on prediction accuracy and fail to effectively handle sudden changes in resource requirements, still facing great challenges in practical applications. At the same time, most existing scheduling algorithms focus on the optimization of a single metric and lack comprehensive consideration of complex constraint conditions in a multi-cloud environment.
[0006] Therefore, there is an urgent need for a multi-cloud GPU task scheduling method that can adapt to resource-constrained scenarios and balance resource utilization rate and system stability. This method should be able to dynamically adjust the task allocation strategy according to the real-time resource status of each cloud platform to ensure the optimal scheduling effect can still be achieved under resource constraints. Summary of the Invention
[0007] The object of the present invention is to solve the problems of unbalanced resource utilization and poor system stability existing in the existing multi-cloud GPU task scheduling solutions. Specifically, in view of the problems of resource overload or idleness caused by limited GPU resources and unreasonable task allocation in each cloud platform, an adaptive scheduling solution is provided. In addition, the present invention also aims to achieve unified management and flexible scheduling of GPU resources in multi-cloud platforms, and improve the overall resource utilization rate.
[0008] To achieve the above object, the present invention provides a multi-cloud GPU task scheduling method for resource-constrained scenarios, which uses a two-layer queue structure to achieve unified management and distribution of tasks. This method dynamically adjusts the queue length by real-time monitoring the resource status of each cloud platform to ensure optimal scheduling under resource constraints.
[0009] Specifically, the present invention constructs a two-layer queue structure of a total task queue and local queues of each cloud platform to achieve centralized management and decentralized execution of tasks. The total queue is responsible for unified reception and priority management of global tasks, while the local queue is responsible for execution control of specific tasks.
[0010] Furthermore, the present invention implements a real-time resource monitoring mechanism to continuously collect key metrics such as GPU utilization rate, task execution duration, and queue status of each cloud platform. Through these data, the system can timely grasp the resource status and processing capabilities of each platform. At the same time, the present invention also includes a task execution situation statistics mechanism to real-time statistics the task completion rate and failure rate of each cloud platform, and these data are used as important reference metrics for evaluating the platform processing capabilities and adjusting the queue length.
[0011] Preferably, the present invention dynamically calculates the optimal queue length of each cloud platform based on the monitoring data. This calculation process comprehensively considers multiple factors such as GPU resource utilization rate, historical task processing capabilities, and preset thresholds to ensure the rationality of the queue length. By analyzing historical data and current status, the system can accurately evaluate the actual processing capabilities of each platform, thereby optimizing task allocation.
[0012] Optionally, the present invention designs a task distribution control mechanism based on the optimal queue length. The system periodically checks the status of each local queue, and triggers task distribution when the queue length is less than the optimal value, and the distribution quantity is equal to the difference between the optimal queue length and the current queue length. This mechanism ensures precise control of task distribution and full utilization of resources.
[0013] In some embodiments, the present invention also includes a resource protection mechanism. When it is detected that the GPU utilization rate exceeds the warning threshold, the system will automatically reduce the upper limit of the queue length of the corresponding platform and can pause task distribution if necessary to prevent resource overload. This mechanism can effectively prevent the occurrence of system instability.
[0014] In addition, the present invention implements a task priority processing mechanism, which can set priorities for different types of tasks. For high-priority tasks, they are allowed to break through the conventional queue length limit under specific conditions to ensure the timely processing of important tasks. At the same time, the system maintains a task priority queue to ensure the rationality of resource scheduling.
[0015] In a preferred embodiment, the system sets a fixed monitoring period (such as 1 minute) to regularly collect and update resource status data. The calculation of the optimal queue length takes into account the recent resource usage trends and sets an appropriate adjustment step size to avoid drastic fluctuations in the queue length. At the same time, the system dynamically adjusts the resource allocation strategy for each platform based on the statistical data of task completion rate and failure rate.
[0016] By adopting the above solutions, the present invention has the following beneficial effects: 1. Through the double-layer queue structure, the unified management and flexible distribution of tasks are realized; 2. The dynamic scheduling mechanism based on real-time monitoring and statistical data ensures the optimization of resource utilization; 3. Through the control of the optimal queue length, the problem of resource overload is effectively avoided; 4. The introduction of the resource protection mechanism improves the stability and reliability of the system; 5. The priority processing mechanism meets the different needs of different tasks.
[0017] In summary, the scheduling method provided by the present invention can effectively solve the task scheduling problem in the scenario of limited GPU resources in multi-cloud platforms, and realizes the unity of resource utilization rate and system stability. By comprehensively applying various technical mechanisms, the present invention significantly improves the scheduling efficiency and reliability of GPU resources in multi-cloud platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0019] Figure 1 It is a system infrastructure diagram. This diagram shows the core components of the entire scheduling system, including the interaction relationship between the total task queue, the local queue cluster, and four key functional modules (queue management module, resource monitoring module, scheduling control module, protection processing module). Through the hierarchical structure design, the overall architecture of the system and the data flow process are clearly expressed.
[0020] Figure 2It is a flowchart of the resource monitoring mechanism. This figure uses the form of a sequence diagram to show the complete process of resource monitoring, including the interaction sequence among the cloud platform monitoring, data processor, time series database, anomaly monitor, and task monitor. It highlights the data collection process with a 1-minute cycle and the continuous task status monitoring process.
[0021] Figure 3 It is a flowchart for calculating the optimal queue length. This figure uses the form of a state diagram to detail each step of the queue length calculation, from parameter processing, GPU data smoothing, historical data analysis, to the complete calculation process of basic length calculation, dynamic adjustment, step size control, and fluctuation detection. Through nested states, it shows the internal details of some steps.
[0022] Figure 4 It is a state machine diagram for task distribution control. This figure describes the state transition relationships during the task distribution process, including states such as idle, check trigger, distribution evaluation, resource check, task distribution, execution monitoring, exception handling, recovery handling, etc., as well as the transition conditions and trigger events between each state.
[0023] Figure 5 It is an architecture diagram of the priority processing mechanism. This figure shows the three main parts of the priority processing system: priority management, resource reservation, and task processing, as well as their associated relationships. Through sub-diagrams, it clearly shows the internal components of each functional module and their interaction methods. Specific implementation manners
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0025] In the description of the present invention, it should be noted that the orientation or positional relationships indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0026] Embodiment 1: Basic system architecture
[0027] Such as Figure 1As shown in the figure, this embodiment provides an infrastructure for a multi-cloud GPU task scheduling system for resource-constrained scenarios. This architecture adopts an innovative double-layer queue structure to achieve unified management and precise distribution of tasks.
[0028] In the double-layer queue structure, the total queue serves as the global task management unit, which is implemented using a priority queue and is used to store all tasks to be processed. Each task item contains attribute information such as task ID, task type, resource requirements, and priority. The total queue realizes fast reading and writing through memory mapping, supports concurrent access, and ensures performance stability under high load conditions.
[0029] The local queues of each cloud platform are implemented using circular buffers, which have a fixed capacity limit and are used to store the tasks to be executed allocated to that platform. The local queues maintain task status information, including statuses such as waiting to execute, executing, and execution completed. Each local queue is equipped with independent read and write pointers to ensure concurrent safety through atomic operations.
[0030] The queue management module is the core component of the system and is responsible for coordinating the data flow between the total queue and the local queues. This module realizes atomic operations of tasks to ensure data consistency during task migration. The queue management module responds to queue status changes through an event-driven mechanism and adjusts the task allocation strategy in real time.
[0031] The resource monitoring module obtains the GPU resource usage by calling the API interfaces provided by each cloud platform. The monitoring data includes metrics such as GPU utilization rate, video memory usage rate, and computing load. This module adopts an asynchronous acquisition method to avoid the impact of monitoring operations on system performance. The collected data is preprocessed and stored in a time series database for subsequent analysis and decision-making.
[0032] The scheduling control module makes task distribution decisions based on the resource monitoring data. This module realizes an adaptive scheduling algorithm based on resource status and dynamically adjusts the task allocation ratio according to the real-time load conditions of each platform. The scheduling control module also maintains a scheduling state machine to handle various state transitions in the task life cycle.
[0033] The protection processing module is responsible for the exception detection and handling of the system. This module realizes a multi-level protection mechanism, including functions such as resource overload protection, task timeout processing, and fault recovery. When an abnormal situation is detected, the protection processing module will trigger the corresponding processing process to ensure the stable operation of the system.
[0034] Each functional module communicates through standardized interfaces and realizes decoupling using the message queue mechanism. The data interaction between modules follows strict protocol specifications to ensure the reliability and consistency of data transmission. The system also realizes a complete logging mechanism to support real-time monitoring of the running state and problem tracking.
[0035] In actual deployment, each component of the system can be independently scaled to support horizontal expansion. System parameters can be flexibly adjusted through configuration files to adapt to different operating environments and business requirements. The system adopts a distributed architecture, and each component maintains a connection through a heartbeat mechanism to ensure the high availability of the system.
[0036] The system architecture design in this embodiment fully considers the special requirements in a multi-cloud environment. Through an innovative two-layer queue structure and modular design, efficient task scheduling in resource-constrained scenarios is achieved. This architecture has good scalability and adaptability, laying a foundation for the subsequent implementation of specific functions.
[0037] Embodiment 2: Resource Monitoring Mechanism
[0038] Based on Embodiment 1, this embodiment details the specific implementation scheme of the resource monitoring mechanism. As Figure 2 shown, the resource monitoring mechanism adopts a hierarchical design to achieve comprehensive monitoring of GPU resource usage and task execution status.
[0039] GPU resource monitoring adopts a periodic sampling method, and the sampling period is set to 1 minute. The monitoring program obtains GPU-related metrics, including data such as the utilization rate of computing units, the utilization rate of video memory, and temperature, by calling the standard API interfaces provided by each cloud platform. To ensure the accuracy of the data, multiple data are obtained at each sampling point and the average value is calculated to eliminate the influence of instantaneous fluctuations. The collected raw data is stored in a time series database after preprocessing, and the preprocessing process includes steps such as outlier filtering and data standardization.
[0040] The calculation of the resource utilization rate adopts the following formula:
[0041] where α and β are weight coefficients, ComputeUtil represents the utilization rate of computing units, and MemoryUtil represents the utilization rate of video memory. The weight coefficients can be dynamically adjusted according to the actual application scenario, and the default settings are α = 0.7 and β = 0.3.
[0042] The task execution monitoring module realizes the tracking of the entire life cycle of tasks. For each task, the system records the following key time points: task submission time, start execution time, and execution completion time. Based on these time points, the actual execution duration and queuing duration of the task are calculated. The calculation of the execution duration takes into account the interruption and recovery of the task to ensure the accuracy of the statistical data.
[0043] The statistics of task completion rate and failure rate adopt a sliding time window mechanism, and the window size is configurable, with the default setting of 1 hour. In each statistical period, the system calculates the following metrics:
[0044] Among them, CompletedTasks represents the number of tasks successfully completed, FailedTasks represents the number of tasks that failed to execute, and TotalTasks represents the total number of tasks.
[0045] To handle outliers in the monitoring data, the system implements an anomaly detection mechanism based on statistical methods. For any data point that exceeds the normal range, the system will mark and record the anomaly, and at the same time trigger the corresponding processing flow. The anomaly detection uses the following judgment condition: |x - μ| > kσ(4)
[0046] Among them, x is the current data point, μ is the historical data mean, σ is the standard deviation, and k is a configurable threshold coefficient with a default value of 3.
[0047] The storage of monitoring data adopts a hierarchical architecture. Real-time data is stored in an in-memory database, and historical data is periodically archived to persistent storage. The data storage format adopts a time series format, and each data point contains a timestamp, a metric value, and related tag information. The system implements a data compression mechanism to downsample and store historical data, which not only ensures the availability of data but also reduces the storage overhead.
[0048] The resource monitoring mechanism of this embodiment provides reliable data support for task scheduling decisions through accurate data collection and processing processes. This mechanism has good scalability and can add new monitoring metrics or adjust statistical methods according to actual needs. The monitoring results are provided to the scheduling control module through a standard interface to support subsequent queue length calculation and task distribution control.
[0049] Embodiment 3: Optimal Queue Length Calculation
[0050] Based on the monitoring data obtained in Embodiment 2, this embodiment details the calculation method of the optimal queue length. As Figure 3 shown, this calculation process uses a multi-factor comprehensive evaluation model to achieve precise control of the queue length.
[0051] The calculation of the optimal queue length first requires processing the input parameters. The system performs time series analysis on the GPU utilization rate data collected in Embodiment 2 and eliminates the influence of short-term fluctuations through the exponential smoothing method. The processed GPU utilization rate is expressed as: GPU smooth (t) = α · GPU raw (t) + (1 - α) · GPU smooth (t - 1)(5)
[0052] Among them, GPU raw (t) is the original usage rate of the current sampling point, GPU smooth (t - 1) is the smoothed value at the previous moment, α is the smoothing coefficient, and its value range is [0, 1], with the default setting being 0.3.
[0053] The historical task processing ability evaluation adopts a time - weighted statistical method. The system calculates the average task processing rate:
[0054] Among them, w i is the time weight, and the more recent the data, the greater the weight. CompletedTasks i is the number of tasks completed in the i - th time window, TimeWindow i is the time window length.
[0055] Based on the processed parameters, the system calculates the base queue length: BaseLength=ProcessRate·AverageExecutionTime·CapacityFactor(7)
[0056] Among them, AverageExecutionTime is the average task execution time, and CapacityFactor is the capacity coefficient, which is used to reserve processing capacity margin.
[0057] The dynamic calculation of the optimal queue length adopts an adaptive adjustment mechanism:
[0058] Among them, MaxGPUUsage is the maximum GPU usage rate threshold, and StabilityFactor is the stability factor, which is used to smooth the change of the queue length.
[0059] To avoid violent fluctuations in the queue length, the system implements a step - size control mechanism: ΔQueue=min(|MaxQueue - CurrentQueue|, MaxStepSize)(9)
[0060] The new queue length reaches the target value through gradual adjustment: NewQueue=CurrentQueue+sign(MaxQueue - CurrentQueue)·ΔQueue(10)
[0061] The system also implements a queue length fluctuation suppression mechanism. When it is detected that the queue length changes frequently within a short period of time, the adjustment speed is slowed down by increasing the value of the stability factor: StabilityFactor = BaseStability·(1 + OscillationIndex)(11)
[0062] Among them, OscillationIndex represents the fluctuation index of the queue length, which is calculated from the recent adjustment frequency.
[0063] To adapt to different operating environments, the system provides an automatic parameter optimization mechanism. By analyzing historical operation data, the system can automatically adjust various parameters, including the smoothing coefficient, capacity coefficient, maximum step size, etc., to make the queue length calculation more in line with actual requirements.
[0064] The optimal queue length calculation method in this embodiment realizes the precise control of the queue length by comprehensively considering multiple influencing factors. This method has good adaptability and can dynamically adjust the queue length according to the resource status and task characteristics, providing a reliable decision-making basis for task distribution.
[0065] Embodiment 4: Task Distribution Control
[0066] Based on the optimal queue length calculated in Embodiment 3, this embodiment details the specific implementation scheme of task distribution control. As Figure 4 shown, the task distribution control mechanism adopts a state machine design to achieve precise task distribution and exception handling.
[0067] The trigger of distribution control adopts a dual mechanism. The periodic trigger is implemented through a timer, and the check period is configurable, with a default setting of 30 seconds. The event trigger is caused by a change in the queue status and is immediately triggered for inspection when the local queue length is lower than the set threshold. The judgment of the trigger condition uses the following formula: TriggerCondition = (CurrentTime - LastCheckTime > CheckInterval) ∨ (QueueLength < Threshol (12)
[0068] Among them, ThresholdRatio is the trigger threshold ratio, with a default setting of 0.5.
[0069] The calculation of the task distribution quantity takes into account multiple constraints. First, calculate the theoretical dispatchable quantity: TheoreticalDispatch = MaxQueue - CurrentQueue(13)
[0070] The actual dispatchable quantity also needs to consider the current processing capacity of the system: ActualDispatch = min(TheoreticalDispatch, AvailableCapacity)(14)
[0071] Among them, AvailableCapacity is dynamically calculated based on the current resource status:
[0072] The resource protection mechanism implements a multi-level protection strategy. When the GPU utilization rate exceeds the warning threshold, the system automatically reduces the allowed queue length:
[0073] Among them, λ is the protection intensity coefficient, and the default value is 2.
[0074] Task anomaly detection uses multi-dimensional metrics for monitoring. The system defines the criteria for determining the task anomaly status:
[0075] When a task anomaly is detected, the system executes the following processing flow: 1. Pause the distribution of related task types 2. Record the anomaly details 3. Trigger an alarm notification 4. Execute the recovery strategy
[0076] The recovery mechanism adopts a progressive scheme and controls the recovery speed through the following formula:
[0077] Among them, NormalRunningTime is the duration for the system to resume normal operation, and RecoveryTimeConstant is the recovery time constant.
[0078] The execution of task distribution uses atomic operations to ensure consistency. Each distribution operation generates a unique operation ID to record the complete distribution process: OperationLog = {OpID, Timestamp, TaskList, TargetQueue, Status}(19)
[0079] The system also implements a rollback mechanism for distribution operations. When an anomaly occurs during the distribution process, the system can be restored to the state before the distribution. The rollback operation is implemented through transaction logs to ensure data consistency.
[0080] The task distribution control mechanism of this embodiment realizes safe and reliable task distribution through precise trigger judgment and multiple protection strategies. This mechanism can effectively handle various abnormal situations, ensure the stable operation of the system, and at the same time has good scalability, and can adjust various parameters and strategies according to actual needs.
[0081] Example 5: Priority Processing Mechanism
[0082] Based on the foregoing embodiments, this embodiment details the specific implementation scheme of the priority processing mechanism. As Figure 5 shown, this mechanism realizes precise control of high-priority tasks through dynamic priority management and resource reservation strategies.
[0083] The priority level is designed using a logarithmic scale to ensure the rationality of priority differences: PriorityLevel = log 2 (BasePriority + 1)·PriorityScale(20)
[0084] Among them, BasePriority is the base priority value, and PriorityScale is the scaling factor. The larger the priority value, the higher the priority.
[0085] The dynamic adjustment of task priority is based on a comprehensive evaluation of multiple factors. The system defines a priority adjustment function: PriorityAdjustment = W t ·TimeFactor + W r ·ResourceFactor + W s ·StrategyFactor (21)
[0086] Among them, TimeFactor reflects the impact of the waiting time of the task, ResourceFactor represents the impact of resource requirements, StrategyFactor represents the strategy adjustment factor, and W t 、W r 、W s are the corresponding weight coefficients.
[0087] To process high-priority tasks, the system implements a resource reservation mechanism. The reserved resource amount is dynamically calculated according to the priority distribution:
[0088] Among them, α i is the reservation coefficient for different priority levels, and HighPriorityRatio i is the proportion of tasks at this level in the total tasks.
[0089] The condition judgment for breaking through the queue length limit adopts a priority threshold mechanism:
[0090] Among them, ThresholdPriority is the priority threshold for breaking through the limit, and SafetyRatio is the safety ratio coefficient.
[0091] To ensure system stability, the execution of high-priority tasks is subject to resource constraints:
[0092] Among them, HighPriorityRatio is the maximum proportion of high-priority tasks, and FlexibilityFactor is the flexibility coefficient.
[0093] The maintenance of the priority queue is implemented using a heap structure and supports the following operations: PriorityQueue = {Enqueue(task), Dequeue(), UpdatePriority(task), GetHighestPriority()} (25)
[0094] The time complexity of queue operations remains at the O(logn) level to ensure efficient task management.
[0095] The system also implements a priority inheritance mechanism to handle dependencies between tasks:
[0096] To prevent priority inversion, the system adopts a priority boosting strategy: BoostPriority = OriginalPriority · (1 + BoostFactor · WaitingTime)(27)
[0097] Among them, BoostFactor is the boosting coefficient, and WaitingTime is the task waiting time.
[0098] The priority processing mechanism of this embodiment realizes differential processing of tasks with different priorities through precise priority calculation and dynamic adjustment strategies. This mechanism can meet the special needs of high-priority tasks while ensuring system stability, and at the same time has good scalability and adaptability.
[0099] It should be understood that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and all of them should be covered by the scope of the claims of the present invention.
Claims
1. A multi-cloud GPU task scheduling method for resource-constrained scenarios, characterized in that: include: (a) Construct a two-layer queue structure consisting of a general task queue and local queues on each cloud platform; (b) Real-time monitoring of GPU resource usage, task execution, and queue status of each cloud platform; (c) Dynamically calculate the optimal queue length of each cloud platform based on monitoring data; (d) Control the distribution of tasks from the total queue to the local queue based on the optimal queue length.
2. The scheduling method according to claim 1, characterized in that: The monitoring includes: (a) Periodically collect GPU resource usage data; (b) Record task execution time and queue time; (c) Statistics on task completion rate and failure rate.
3. The scheduling method according to claim 1, characterized in that: The calculation of the optimal queue length takes into account the following factors: (a) Current utilization of GPU resources; (b) historical task processing capabilities; (c) Preset resource usage threshold.
4. The scheduling method according to claim 1, characterized in that: The task distribution includes: (a) Periodically check the local queue length; (b) triggering distribution when the local queue length is less than the optimal queue length; (c) The distribution quantity is equal to the difference between the optimal queue length and the current queue length.
5. The scheduling method according to claim 1, characterized in that: Also includes resource protection mechanisms: (a) When the GPU usage exceeds the threshold, the queue length limit is automatically reduced; (b) When a task anomaly is detected, the task distribution on the corresponding cloud platform is suspended.
6. The scheduling method according to claim 1, characterized in that: Priority processing includes: (a) Set priorities for different types of tasks; (b) High priority tasks can exceed the normal queue length limit.
7. A device for implementing the scheduling method according to any one of claims 1 to 6, characterized in that: include: (a) a queue management module, used to manage a double-layer queue structure; (b) a resource monitoring module, used to monitor resource status; (c) a scheduling control module, used to execute task distribution; (d) Protection processing module, used to execute resource protection strategy.