Large-scale concurrent task dynamic scheduling method and system of computing power host management platform

By monitoring the task load in real time and adjusting resource allocation dynamically, combining priority scheduling and current limiting mechanisms, the problems of resource competition and waste under high concurrent tasks are solved, and the efficiency of computing resources and system stability are improved.

CN120179398APending Publication Date: 2025-06-20INSPUR COMM TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510280916.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When the existing computing power host management platform handles high-concurrent computing tasks, resource competition is serious, low-priority tasks affect high-priority tasks, resource waste is serious, and the system is prone to overload, affecting stability.

Method used

Through real-time task load monitoring, automatic resource allocation, priority scheduling and current limiting mechanisms, the resource occupation of tasks is dynamically adjusted to ensure that high-load tasks prioritize resources acquisition, downgrade low-load tasks, limit the rate of new tasks submission, and prevent system overload.

Benefits of technology

It improves the efficiency of computing resources, ensures the timely execution of high-priority tasks, reduces resource waste, enhances system stability, and adapts to the needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179398A_ABST
    Figure CN120179398A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a large-scale concurrent task dynamic scheduling method and system for a computing power host management platform, and the method comprises the following steps: real-time task load monitoring, automatic resource allocation, priority scheduling and flow limiting. The method has the beneficial effects that the use efficiency of computing resources is improved by dynamically adjusting CPU, memory and GPU resource occupation of tasks; a priority scheduling and task queue mechanism is adopted, so that the influence of low-priority tasks on high-priority tasks is avoided, and the fairness of the calculation tasks is improved; in combination with a current limiting strategy, computing resources are effectively prevented from being exhausted by overload tasks under the high-concurrency condition, and normal operation of the system is ensured; resources are automatically released after tasks are completed, so that computing resources are reasonably distributed, and waste caused by long-term occupation of low-load tasks is avoided; the method is suitable for scenes needing efficient computing power distribution such as AI training, high-performance computing and cloud computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a method and system for dynamically scheduling large-scale concurrent tasks of a computing power host management platform. Background Art

[0002] With the wide application of technologies such as artificial intelligence, big data analysis, and high-performance computing (HPC), the computing power host management platform needs to process a large number of concurrent computing tasks. The current computing resource management methods mainly rely on static or simple resource allocation strategies, which cannot effectively adapt to the dynamic changes of computing tasks, resulting in the following problems:

[0003] 1. Severe resource competition: The scramble for CPU, memory, and GPU resources by high-concurrency tasks leads to unbalanced system load and affects task execution efficiency.

[0004] 2. Low-priority tasks affect high-priority tasks: Lack of a reasonable scheduling mechanism, important tasks may be delayed due to insufficient resources.

[0005] 3. Resource waste: Some low-load tasks occupy computing power resources for a long time, while high-load tasks cannot obtain more resources in time, resulting in a decrease in overall resource utilization.

[0006] 4. System overload risk: In the case of a sharp increase in concurrent computing tasks, lack of effective task queue management and flow control mechanisms may lead to exhaustion of computing power platform resources and affect system stability. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for dynamically scheduling large-scale concurrent tasks of a computing power host management platform to solve the problems raised in the above background art.

[0008] To achieve the above purpose, the present invention provides the following technical solution: A method for dynamically scheduling large-scale concurrent tasks of a computing power host management platform, the method comprising the following steps:

[0009] a) Real-time task load monitoring: Monitor the CPU, memory, and GPU resource occupancy of computing tasks, and combine historical data to analyze the task load trend, and use an adaptive learning algorithm to predict the change in the computing power demand of tasks;

[0010] b) Automatic resource allocation: Adjust the resource allocation ratio in real time according to the computing load, preferentially allocate additional computing power resources to high-load tasks, automatically degrade the computing power occupancy of low-load tasks, and automatically release the computing resources they occupy after the tasks are completed or timed out;

[0011] c) Priority scheduling and rate limiting: Sort tasks based on the real-time nature, resource requirements, and business priorities of the tasks to ensure that important tasks are executed first, and limit the submission rate of new tasks when computing resources are scarce.

[0012] Preferably, the real-time task load monitoring step further includes:

[0013] a1) Monitor the resource usage of each computing task in real time, including but not limited to CPU usage rate, memory occupancy rate, and GPU occupancy rate;

[0014] a2) Establish a task load trend model through historical data analysis to predict the change in resource requirements of tasks in the future for a period of time.

[0015] Preferably, the resource automatic allocation step further includes:

[0016] b1) Set a resource allocation threshold. When the task load exceeds the set threshold, automatically allocate more resources to it;

[0017] b2) When the task load is below the set threshold for a certain period of time, automatically reduce its resource allocation and release the excess resources for other tasks to use.

[0018] Preferably, the priority scheduling and rate limiting step further includes:

[0019] c1) Set task priority sorting rules according to the real-time nature, resource requirements, and business priorities of the tasks;

[0020] c2) When the system resources are scarce, start the rate limiting mechanism to limit the submission rate of new tasks to ensure the stable operation of the system;

[0021] c3) Monitor the length of the task queue. When the queue length exceeds the set threshold, further adjust the rate limiting strategy.

[0022] Preferably, the method is applicable to various application scenarios that require efficient computing power allocation, such as AI training, high-performance computing, and cloud computing. By dynamically adjusting the resource occupancy of tasks, it improves the utilization efficiency of computing resources and ensures the stable operation of the system in a high-concurrency task environment.

[0023] A large-scale concurrent task dynamic scheduling system for a computing power host management platform, which is applied to a large-scale concurrent task dynamic scheduling method for a computing power host management platform. The system includes:

[0024] A real-time task load monitoring module, which is used to monitor the CPU, memory, and GPU resource occupancy of computing tasks, analyze the task load trend in combination with historical data, and use an adaptive learning algorithm to predict the change in computing power requirements of tasks;

[0025] The resource automatic allocation module dynamically adjusts the allocation ratios of CPU, memory, and GPU resources during task execution according to the data provided by the real-time task load monitoring module. It preferentially allocates additional computing power resources to high-load tasks, automatically reduces the computing power occupancy of low-load tasks, and automatically releases the computing resources they occupy after the tasks are completed or timed out.

[0026] The priority scheduling and rate limiting module sorts tasks based on the real-time nature, resource requirements, and business priorities of the tasks to ensure that important tasks are executed first. When computing resources are scarce, it limits the submission rate of new tasks to prevent system overload.

[0027] Preferably, the real-time task load monitoring module further includes:

[0028] The resource occupancy monitoring unit is used to monitor the CPU usage rate, memory occupancy rate, and GPU occupancy rate of each computing task in real time.

[0029] The trend analysis and prediction unit establishes a task load trend model based on historical data and uses an adaptive learning algorithm to predict the change in resource requirements of tasks in the future for a period of time.

[0030] Preferably, the resource automatic allocation module further includes:

[0031] The resource allocation strategy unit sets resource allocation thresholds and dynamic adjustment rules, and adjusts the resource allocation ratio in real time according to the task load.

[0032] The resource release management unit automatically triggers the resource release mechanism after the task is completed or timed out to ensure that the computing resources are recycled and reallocated in a timely manner.

[0033] Preferably, the priority scheduling and rate limiting module further includes:

[0034] The task sorting unit sets sorting rules according to the real-time nature, resource requirements, and business priorities of the tasks to ensure that important tasks are executed first.

[0035] The rate limiting control unit activates the rate limiting mechanism when system resources are scarce, prevents system overload by controlling the submission rate of new tasks, and at the same time monitors the length of the task queue and dynamically adjusts the rate limiting strategy to adapt to changes in task load.

[0036] Preferably, the system is applicable to various application scenarios that require efficient computing power allocation, such as AI training, high-performance computing, and cloud computing. By monitoring task load in real time, dynamically adjusting resource allocation, priority scheduling, and rate limiting control, it realizes the efficient utilization of computing resources and the stable operation of the system, improves the computing power utilization rate, reduces resource waste, enhances system stability, and adapts to the requirements of different application scenarios.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] The dynamic scheduling method and system for large-scale concurrent tasks of the computing power host management platform proposed by the present invention improve the utilization efficiency of computing resources by dynamically adjusting the CPU, memory, and GPU resource occupancy of tasks; adopt a priority scheduling and task queue mechanism to avoid low-priority tasks from affecting high-priority tasks and improve the fairness of computing tasks; combine a flow control strategy to effectively prevent computing resources from being exhausted by overloaded tasks in high-concurrency situations and ensure the normal operation of the system; automatically release resources after tasks are completed to reasonably allocate computing resources and avoid waste caused by long-term occupation by low-load tasks; and are applicable to scenarios such as AI training, high-performance computing, and cloud computing that require efficient computing power allocation. Brief Description of the Drawings

[0039] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0040] In order to clearly and completely describe the objectives, technical solutions of the present invention and make the advantages more clear, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, not all of the embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0041] Embodiment 1, please refer to Figure 1 , the present invention provides a technical solution: a dynamic scheduling method for large-scale concurrent tasks of a computing power host management platform, the method includes the following steps:

[0042] 1. Real-time task load monitoring

[0043] Monitor the CPU, memory, and GPU resource occupancy of computing tasks, and analyze the task load trend in combination with historical data.

[0044] Adopt an adaptive learning algorithm to predict the change of computing power demand of tasks and provide data support.

[0045] 2. Automatic resource allocation

[0046] During the execution of tasks, adjust the resource allocation ratio in real time according to the computing load.

[0047] High-load tasks are preferentially allocated additional computing power resources to ensure the efficient completion of computing tasks.

[0048] Low-load tasks automatically degrade the computing power occupancy and release unnecessary resources for other tasks to use.

[0049] After the task is completed or times out, the computing resources it occupies are automatically released to prevent the resources from being occupied by inefficient tasks for a long time and improve the computing power reuse rate.

[0050] 3. Priority Scheduling and Rate Limiting

[0051] Tasks are sorted according to the real-time nature, resource requirements, and business priorities of the tasks to ensure that important tasks are executed first and to prevent the computing platform from being overloaded by instantaneous high-concurrency tasks.

[0052] When computing resources are scarce, the submission rate of new tasks is restricted to ensure the stable operation of the system.

[0053] In the second embodiment, based on the first embodiment, a dynamic scheduling system for large-scale concurrent tasks of a computing power host management platform is proposed, which is applied to a dynamic scheduling method for large-scale concurrent tasks of a computing power host management platform. The system includes:

[0054] A real-time task load monitoring module for monitoring the CPU, memory, and GPU resource occupancy of computing tasks, analyzing the task load trend in combination with historical data, and predicting the change in the computing power requirements of tasks using an adaptive learning algorithm; it also includes:

[0055] A resource occupancy monitoring unit for real-time monitoring of the CPU usage rate, memory occupancy rate, and GPU occupancy rate of each computing task;

[0056] A trend analysis and prediction unit for establishing a task load trend model based on historical data and predicting the change in resource requirements of tasks in the future for a period of time using an adaptive learning algorithm.

[0057] A resource automatic allocation module that dynamically adjusts the allocation ratio of CPU, memory, and GPU resources during the task execution process according to the data provided by the real-time task load monitoring module, preferentially allocates additional computing power resources to high-load tasks, automatically reduces the computing power occupancy of low-load tasks, and automatically releases the computing resources they occupy after the task is completed or times out; it also includes:

[0058] A resource allocation strategy unit that sets resource allocation thresholds and dynamic adjustment rules and adjusts the resource allocation ratio in real time according to the task load;

[0059] A resource release management unit that automatically triggers a resource release mechanism after the task is completed or times out to ensure the timely recovery and reallocation of computing resources.

[0060] A priority scheduling and rate limiting module that sorts tasks according to the real-time nature, resource requirements, and business priorities of the tasks to ensure that important tasks are executed first and restricts the submission rate of new tasks when computing resources are scarce to prevent system overload; it also includes:

[0061] The task sorting unit sets sorting rules according to the real-time nature, resource requirements, and business priorities of tasks to ensure that important tasks are executed first;

[0062] The flow-limiting control unit activates the flow-limiting mechanism when system resources are scarce. By controlling the submission rate of new tasks, it prevents system overload. At the same time, it monitors the length of the task queue and dynamically adjusts the flow-limiting strategy to adapt to changes in task load.

[0063] The system is applicable to various application scenarios that require efficient computing power allocation, such as AI training, high-performance computing, and cloud computing. By real-time monitoring task load, dynamically adjusting resource allocation, priority scheduling, and flow-limiting control, it realizes the efficient utilization of computing resources and the stable operation of the system, improves the computing power utilization rate, reduces resource waste, enhances system stability, and adapts to the requirements of different application scenarios.

[0064] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large-scale concurrent task dynamic scheduling method for a computing power host management platform, characterized by: The method comprises the following steps: a) Real-time task load monitoring: monitor the CPU, memory, and GPU resource usage of computing tasks, analyze task load trends based on historical data, and use adaptive learning algorithms to predict changes in task computing power requirements; b) Automatic resource allocation: Adjust resource allocation ratio in real time according to computing load, prioritize additional computing resources for high-load tasks, automatically downgrade computing power usage of low-load tasks, and automatically release computing resources occupied by tasks after completion or timeout; c) Priority scheduling and rate limiting: Sort tasks based on their real-time nature, resource requirements, and business priorities to ensure that important tasks are executed first, and limit the rate at which new tasks are submitted when computing resources are tight.

2. According to claim 1, a large-scale concurrent task dynamic scheduling method for a computing power host management platform is characterized in that: The real-time task load monitoring steps also include: a1) Real-time monitoring of resource usage of each computing task, including but not limited to CPU usage, memory usage, and GPU usage; a2) Establish a task load trend model through historical data analysis to predict changes in task demand for resources in the future.

3. According to claim 1, a large-scale concurrent task dynamic scheduling method for a computing power host management platform is characterized in that: The automatic resource allocation step also includes: b1) Set a resource allocation threshold. When the task load exceeds the set threshold, more resources are automatically allocated to it; b2) When the task load is lower than the set threshold for a certain period of time, its resource allocation is automatically reduced and excess resources are released for use by other tasks.

4. The method for dynamic scheduling of large-scale concurrent tasks of a computing power host management platform according to claim 1, characterized in that: Priority scheduling and current limiting steps also include: c1) Set task priority sorting rules based on the real-time nature of the task, resource requirements, and business priorities; c2) When system resources are tight, the current limiting mechanism is activated to limit the submission rate of new tasks to ensure stable operation of the system; c3) Monitor the task queue length and further adjust the flow control strategy when the queue length exceeds the set threshold.

5. The method for dynamic scheduling of large-scale concurrent tasks of a computing power host management platform according to claim 1, characterized in that: The method is applicable to various application scenarios such as AI training, high-performance computing, and cloud computing that require efficient computing power allocation. By dynamically adjusting task resource occupancy, the efficiency of computing resource utilization is improved, ensuring the stable operation of the system in a high-concurrency task environment.

6. A large-scale concurrent task dynamic scheduling system for a computing power host management platform, applied to a large-scale concurrent task dynamic scheduling method for a computing power host management platform as described in any one of claims 1 to 5, characterized in that: The system comprises: Real-time task load monitoring module, which is used to monitor the CPU, memory, and GPU resource usage of computing tasks, analyze task load trends based on historical data, and use adaptive learning algorithms to predict changes in task computing power requirements; The automatic resource allocation module dynamically adjusts the allocation ratio of CPU, memory, and GPU resources during task execution based on the data provided by the real-time task load monitoring module, prioritizes additional computing resources for high-load tasks, automatically downgrades the computing power usage of low-load tasks, and automatically releases the computing resources occupied by the task after it is completed or timed out; The priority scheduling and current limiting module sorts tasks according to their real-time nature, resource requirements, and business priorities to ensure that important tasks are executed first, and limits the submission rate of new tasks when computing resources are tight to prevent system overload.

7. The large-scale concurrent task dynamic scheduling system of a computing power host management platform according to claim 6 is characterized in that: The real-time task load monitoring module also includes: Resource occupancy monitoring unit, used to monitor the CPU usage, memory usage, and GPU usage of each computing task in real time; The trend analysis and prediction unit establishes a task load trend model based on historical data and uses an adaptive learning algorithm to predict changes in task demand for resources in the future.

8. The large-scale concurrent task dynamic scheduling system of a computing power host management platform according to claim 6, characterized in that: The automatic resource allocation module also includes: Resource allocation strategy unit, which sets resource allocation thresholds and dynamic adjustment rules, and adjusts resource allocation ratios in real time according to task load; The resource release management unit automatically triggers the resource release mechanism after the task is completed or timed out, ensuring that computing resources are recovered and reallocated in a timely manner.

9. The large-scale concurrent task dynamic scheduling system of a computing power host management platform according to claim 6, characterized in that: The priority scheduling and current limiting module also includes: The task sorting unit sets sorting rules based on the real-time nature of the tasks, resource requirements, and business priorities to ensure that important tasks are executed first; The current limiting control unit starts the current limiting mechanism when system resources are tight, prevents system overload by controlling the submission rate of new tasks, monitors the length of the task queue, and dynamically adjusts the current limiting strategy to adapt to changes in task load.

10. The large-scale concurrent task dynamic scheduling system of a computing power host management platform according to claim 6, characterized in that: The system is suitable for various application scenarios that require efficient computing power allocation, such as AI training, high-performance computing, and cloud computing. By real-time monitoring of task loads, dynamic adjustment of resource allocation, priority scheduling, and current limiting control, it can achieve efficient utilization of computing resources and stable system operation, improve computing power utilization, reduce resource waste, enhance system stability, and adapt to the needs of different application scenarios.

Citation Information

Cited By

  • Method and system for intelligent capacity planning in hybrid cloud environment

    CN120353610A

  • Hard disk resource adjusting method and device, server, storage medium and product

    CN120723470A

  • Hard disk resource adjustment method and device, server, storage medium and product

    CN120723470B

  • Auxiliary driving device and system based on video image processing

    CN120766237A