AI accelerator card resource scheduling method based on optimization model

Through real-time monitoring and dynamic adjustment of computing core allocation, combined with machine learning algorithms and task dependency graphs, the AI accelerated card resource scheduling is solved, and efficient and flexible task execution and response speed improvement is achieved.

CN120407161APending Publication Date: 2025-08-01GUIZHOU BLUESKY INNOVATIVE SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510452363.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing AI accelerator card resource scheduling method lacks the comprehensive collection and initialization of task priorities and resource data in the initial stage. The model construction process is unclear, resulting in insufficient application and accuracy, too much focus on genetic algorithm optimization, neglecting real-time monitoring and dynamic adjustment strategies, and lacking forward-looking and accurate.

Method used

By monitoring the calculation core usage and execution progress of tasks in real time, combining dynamic adjustment strategies and machine learning algorithms, optimizing calculation core allocation, dynamically adjusting task priority and resource allocation, and building a task dependency graph to identify affected tasks and perform local updates.

Benefits of technology

It improves the resource utilization rate of AI accelerator card, ensures efficient task execution, enhances predictability and flexibility, optimizes task scheduling strategies, and improves overall performance and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407161A_ABST
    Figure CN120407161A_ABST
Patent Text Reader

Abstract

The invention discloses an AI accelerator card resource scheduling method based on an optimization model, and the method comprises the steps: firstly, obtaining and recording the priorities and resource data of all tasks in an AI accelerator card; thirdly, a task resource allocation optimization model is built based on the data, and relevant parameters are initialized; in the task execution process, the utilization rate and the execution progress of the calculation cores of the task are monitored in real time, and distribution of the calculation cores is adjusted in time according to a dynamic adjustment strategy in combination with historical execution data; meanwhile, a machine learning algorithm is used for deeply analyzing the change trend of task resources, and a calculation core allocation strategy is planned and adjusted in advance; according to the method, the utilization rate of AI accelerator card resources is effectively improved, and efficient and orderly execution of tasks is ensured; through a mechanism of dynamically adjusting the task priority, the low-priority task can be timely processed when the resource is allowed, the task scheduling strategy is further optimized, and the overall performance of the AI acceleration card and the response speed of the AI acceleration card to the task are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing and resource scheduling, and particularly relates to an AI accelerator card resource scheduling method based on an optimization model. Background Art

[0002] An AI accelerator card, also known as a deep learning card or a GPU accelerator card, is a hardware component designed specifically to accelerate artificial intelligence (AI) and machine learning (ML) tasks. It integrates high-performance computing cores and a large amount of memory, aiming to accelerate the computing process of algorithms such as machine learning and deep learning. During the training process of a deep learning model, the AI accelerator card can utilize its powerful parallel computing ability to accelerate the processing and calculation of large-scale data, thereby shortening the model training time; after the model is deployed, the AI accelerator card can also accelerate the inference process to achieve fast and accurate prediction and response.

[0003] The invention patent with the patent application number 202410881509.5 discloses an optimal resource scheduling general model and a corresponding objective function for an AI accelerator card resource scheduling based on a dual optimization model; simplifies the optimal resource scheduling general model into a dual optimization model of global-level multi-task execution sequence arrangement and local-level accelerator card core allocation; respectively uses an improved adaptive genetic algorithm with an elite strategy and the Lagrange multiplier method for optimization and solution to obtain the solution result, that is, the accelerator card resource scheduling scheme. The method combines the characteristics of the AI accelerator card itself with the requirements, priority order, and required computing resources of multi-task processing, quickly obtains the scheduling scheme, improves the overall computing efficiency of a limited number of accelerator cards for multi-tasks, and realizes the intelligent, dynamic, and efficient management of AI computing resources.

[0004] However, there is a lack of comprehensive collection and initialization of task priorities and resource data in the initial stage of resource scheduling; the description of the model construction process is not detailed enough, affecting the applicability and accuracy; it overly focuses on genetic algorithm optimization, ignoring real-time monitoring and dynamic adjustment strategies; it does not consider using machine learning algorithms to analyze the resource change trend, lacking foresight and accuracy. Summary of the Invention

[0005] This application provides an AI accelerator card resource scheduling method based on an optimization model. By combining real-time monitoring, dynamic adjustment, and machine learning prediction, it optimizes the resource allocation of the AI accelerator card, improves the resource scheduling speed, and can manage AI computing resources more intelligently, dynamically, and efficiently.

[0006] This application provides an AI accelerator card resource scheduling method based on an optimization model, including: S1, obtaining and recording the priority information and task resource data of all tasks in the AI accelerator card; S2. Construct an optimization model for task resource allocation according to the task priority and task resource data, and initialize the parameters of the optimization model; S3. Monitor the computing core utilization rate and execution progress of tasks in real time, and adjust the computing core allocation in a timely manner according to the dynamic adjustment strategy in combination with the historical task execution data; S4. Analyze the changing trend of task resources using machine learning algorithms, and adjust the computing core allocation strategy in advance.

[0007] Preferably, in S3, the dynamic adjustment strategy specifically includes: S3a. For each task, compare its current computing core utilization rate with the set threshold, and judge whether the task needs to adjust the number of computing cores according to the comparison result; S3b. If the computing core utilization rate of the task is lower than the threshold and the execution progress is lower than 1 / 3 of the task, trigger the dynamic adjustment mechanism and dynamically increase the number of computing cores of the task; S3c. If the computing core utilization rate is higher than the threshold and the execution progress is higher than 2 / 3 of the task, evaluate the impact of reducing the number of computing cores on the execution speed of the task S3d. Dynamically adjust the number of computing cores of the task according to the evaluation result, and release resources to other tasks that require computing resources; S3e. Update the status information such as the number of computing cores and execution progress of tasks in the database, and continue to monitor in real time.

[0008] Preferably, the evaluation method specifically includes: simulating the task execution speed after reducing the number of computing cores, and comparing it with the current execution speed. According to the evaluation result, determine whether to reduce the number of computing cores and the reduced quantity; assume that n computing cores are reduced, then simulate the task execution time T_new after reduction, and compare it with the current execution time T_current; if T_new is within an acceptable range, then it can be considered to reduce n computing cores.

[0009] Preferably, S4 specifically includes: collecting historical task execution data, including task type, execution time, required number of computing cores, and execution duration; after cleaning and normalizing the data, extract the task type, average task resource volume during the execution period, and the dependency relationship characteristics between tasks from the historical data, and construct a feature vector as the input of the machine learning model; use the trained model to predict the task resources in a future period of time, and output the prediction results, including the predicted task resources and changing trends; according to the prediction results, plan the allocation of computing cores in advance.

[0010] Preferably, S3 further includes: S31. Based on the type of task, task resources, and the status of AI acceleration card resources, perform a simple classification of several categories of tasks; S32. Set a "rated" number of computing cores for each category of tasks, and allocate computing core resources according to priority within the same category of tasks; where "rated" refers to the number of computing cores that should be allocated to this category of tasks under normal circumstances; S33. Set a waiting counter for the tasks under each classification to record the waiting time of the tasks in the queue; S34. As the waiting time increases, increase the actual scheduling priority of the tasks.

[0011] Preferably, S33 includes: setting a waiting counter W for each task to record the waiting time of the task in the queue; initially, W = 0; whenever the AI acceleration card performs a resource allocation or task status update, for each task in the waiting queue, increment its waiting counter by 1.

[0012] Preferably, in S34, the actual scheduling priority includes: the actual scheduling priority adjustment formula is: , is the initial priority of the task; is the waiting time of the task (in seconds), is the priority adjustment coefficient, used to control the influence degree of the waiting time on the priority; is to take the natural logarithm of the waiting time to avoid the priority increasing too fast linearly with the waiting time, update the actual scheduling priorities of all tasks , and re - sort.

[0013] Preferably, S3 further includes: S35. Construct a task dependency graph for the classified tasks; S36. When the priority of a task changes, identify the affected tasks; S37. Perform local updates and re - allocate computing cores for the affected tasks.

[0014] Preferably, in S37, re - allocating computing cores further includes: S37a. For the affected tasks, recalculate the actual scheduling priority of the currently affected tasks according to the formula in step S34; S37b. According to the new actual scheduling priority, re - sort the affected tasks to determine the new task execution order and dependency relationship; S37c. For the tasks with increased priority, try to reclaim computing core resources from the tasks with decreased priority; for the tasks with decreased priority, promptly release the excess computing core resources they occupy for other tasks to use; S37d. During the resource reallocation process, keep the total amount of resources allocated in the same type of tasks unchanged, that is, reallocate resources within the same type of tasks.

[0015] Preferably, for S37, perform local updates on the affected tasks, specifically including: S371. Divide the task dependencies into dependency levels and assign a level value to each type of dependency; S372. When the dependency changes and there are cross-category tasks in the successor tasks, identify the cross-category tasks; S373. According to the new dependencies and task priorities, reclassify the cross-category tasks and correct the simple classification of the tasks; S374. According to the reclassified tasks, recalculate the task priorities using the actual scheduling priority adjustment calculation formula; S375. According to the recalculated priorities, sort and schedule the tasks and add them to the machine learning algorithm to analyze the changing trend of task resources, and pre-adjust the allocation of computing cores.

[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: By monitoring in real time and dynamically adjusting the computing core allocation, the utilization rate of AI acceleration card resources is effectively improved, ensuring the efficient execution of tasks. At the same time, using the machine learning algorithm to predict the task resource trend and plan resource allocation in advance enhances the predictability and flexibility of the AI acceleration card. In addition, the dynamic adjustment of task priority mechanism ensures that low-priority tasks can also be processed in a timely manner when resources are abundant, optimizing the task scheduling strategy and improving the overall performance and response speed of the AI acceleration card.

[0017] Based on the type of tasks, task resources, and the resource status of the AI acceleration card, a "rated" number of computing cores is customized for each type of task, ensuring the rationality of resource allocation. At the same time, by introducing a waiting counter and a dynamic priority adjustment mechanism, tasks can dynamically increase their priorities according to their waiting time during the waiting process, effectively avoiding the problem that low-priority tasks cannot obtain computing core resources for a long time. This design not only improves the execution efficiency of tasks but also enhances the response speed and fairness of the AI acceleration card. In addition, the adjustment formula for the actual scheduling priority fully considers the waiting time and initial priority of tasks, making resource allocation more accurate and efficient, and further improving the overall performance and stability of the AI acceleration card.

[0018] By constructing a task dependency graph and clarifying the direct and indirect dependency relationships between tasks, the task management AI acceleration card can more accurately grasp the logical relationships between tasks. When the task priority changes, the AI acceleration card can quickly trigger the dependency analysis, comprehensively identify the affected tasks, and perform local updates and reallocate computing cores for these tasks. This solution not only improves the flexibility and response speed of task scheduling, but also optimizes resource allocation, ensuring the efficiency and stability of task execution. At the same time, by reallocating resources within the same type of tasks, the balance of the total resource allocation is maintained, avoiding resource waste and conflicts. In addition, this solution also considers the coping strategies in case of resource shortage, ensuring that tasks can be successfully executed under the limit of the rated number of computing cores, further improving the efficiency of task completion.

[0019] By carefully dividing the task dependency levels and monitoring the changes in dependency relationships in real time, the flexibility and accuracy of task scheduling are significantly improved. When the dependency relationship changes, especially when there are cross-category tasks, this solution can quickly identify and reclassify the affected tasks, ensuring the rationality and consistency of task classification. At the same time, the actual scheduling priority adjustment calculation formula is used to recalculate the task priority, making the task sorting and scheduling more in line with the actual business requirements and resource conditions. Brief Description of the Drawings

[0020] Figure 1 It is a flowchart showing a resource scheduling method for an AI acceleration card based on an optimization model according to an embodiment of the present invention. Detailed Embodiments

[0021] To facilitate the understanding of the present invention, the present application will be described more comprehensively with reference to the relevant drawings; the drawings show preferred embodiments of the present invention, however, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs; the terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0023] Embodiment 1: Figure 1 It is a flowchart showing a resource scheduling method for an AI acceleration card based on an optimization model according to an embodiment of the present invention.

[0024] As Figure 1As shown in the figure, an AI acceleration card resource scheduling method based on an optimization model includes the following steps: S1. Obtain and record the priority information and task resource data of all tasks in the AI acceleration card.

[0025] Among them, all tasks in this application refer to all training tasks and inference tasks to be scheduled on the AI acceleration card.

[0026] Specifically, traverse and collect all tasks to be scheduled in the running state of the AI acceleration card, and record the priority information such as the expected completion time, urgency, and task type for each task to be scheduled; record the task resources of each task in detail, including but not limited to the computing core requirements (such as the number of cores of the AI acceleration card required, the number of GPUs, or the type and number of specific AI acceleration cards), the video memory size requirements, real-time requirements (such as latency limits, throughput requirements), etc., to ensure that the collected data can accurately reflect the resource consumption during the actual execution of the task.

[0027] S2. According to the task priority and task resource data, construct a task resource allocation optimization model and initialize the optimization model parameters.

[0028] Specifically, clarify that the goal of the task resource allocation optimization model is to maximize resource utilization and minimize task completion time. Considering the computing power, solution speed, accuracy, and adaptability to problem description of the model, select a linear programming model, a non-linear programming model, or a reinforcement learning model, determine the decision variables in the model (such as task allocation variables, resource allocation variables, etc.), and set constraint conditions for each decision variable. According to the AI acceleration card resources and task requirements, establish the constraint conditions of the model (total resource limit, task execution order limit, and task execution time limit, etc.) to ensure that the constraint conditions can accurately reflect the actual situation and limitations of the AI acceleration card operation. Combine the objective function, decision variables, and constraint conditions into a complete model expression.

[0029] According to the AI acceleration card resources and task requirements, set an initial task resource allocation strategy, and set priority weights for different tasks to reflect the relative importance of tasks in the AI acceleration card. The priority weights should be determined comprehensively according to factors such as the urgency, value, and deadline of the tasks, and ensure that they can reasonably reflect the priority order of the tasks. This application does not make specific limitations on the allocation setting of priorities, and the priorities need to be set according to specific application scenarios.

[0030] Among them, the task resource is the demand for the computing core required during task execution.

[0031] S3. Real-time monitor the computing core utilization rate and execution progress of tasks, and adjust the computing core allocation in a timely manner according to the dynamic adjustment strategy in combination with historical task execution data.

[0032] Specifically, deploy and configure the AI acceleration card monitoring tool Prometheus or write a custom monitoring script to track the computing core utilization rate and execution progress of each task in real time, ensuring that the monitoring tool or script can accurately and timely collect task status data.

[0033] Set the monitoring frequency (per second, per minute, or every five minutes), store the computing core utilization rate and execution progress of the collected tasks in a database, and perform statistical analysis on this real-time data. Sort the tasks according to their importance, urgency, and execution time limit.

[0034] Among them, the dynamic adjustment strategy includes: S3a, for each task, compare its current computing core utilization rate with the set threshold, and judge whether the task needs to adjust the number of computing cores according to the comparison result.

[0035] S3b, if the computing core utilization rate of this task is lower than the threshold and the execution progress is lower than 1 / 3 of the task, trigger the dynamic adjustment mechanism and dynamically increase the number of computing cores of this task.

[0036] Among them, the increased number of computing cores needs to be comprehensively determined according to the task resources of the task, the current resource abundance of the AI acceleration card, and the priority of the remaining tasks.

[0037] Analyze the resource status of the total number of available computing cores, memory size, network bandwidth, etc. of the AI acceleration card. Set a reasonable computing core utilization rate threshold, such as 70% (the specific value needs to be determined according to the actual situation of the AI acceleration card), based on the resource abundance of the AI acceleration card. The threshold should not only make full use of the resources of the AI acceleration card but also avoid resource waste and AI acceleration card overload. For different types of tasks (such as compute-intensive, I / O-intensive, etc.), different thresholds can be set according to their task resources and characteristics. For example, for compute-intensive tasks, a slightly higher threshold can be set to ensure that they obtain sufficient computing resources.

[0038] The reference calculation formula for the increased number of computing cores is: increased number of computing cores = (target computing core utilization rate - current computing core utilization rate) / performance improvement coefficient of a single computing core × task resource coefficient × AI acceleration card resource abundance coefficient × task priority coefficient. Among them, each coefficient is estimated based on historical data or experience.

[0039] S3c, if the computing core utilization rate is higher than the threshold and the execution progress is higher than 2 / 3 of the task, evaluate the impact of reducing the number of computing cores on the execution speed of this task.

[0040] The evaluation method specifically includes: simulating the task execution speed after reducing the number of computing cores and comparing it with the current execution speed. According to the evaluation results, determine whether to reduce the number of computing cores and the reduced quantity. Assume that n computing cores are reduced, then simulate or estimate the task execution time T_new after reduction and compare it with the current execution time T_current. If T_new is within an acceptable range (such as not exceeding 10% of the original plan), then reducing n computing cores can be considered.

[0041] S3d, dynamically adjust the number of computing cores for this task according to the evaluation results, and release resources to other tasks that require computing resources.

[0042] S3e, update the status information such as the number of computing cores and the execution progress of the task in the database, and continue to monitor in real time.

[0043] S4, adopt a machine learning algorithm to analyze the change trend of task resources and adjust the computing core allocation strategy in advance.

[0044] Specifically, collect historical task execution data, including task type, execution time, required number of computing cores, execution duration, etc. After cleaning and normalizing the data, extract features such as task type, average task resource volume in the execution time period, and dependencies between tasks from the historical data, and construct a feature vector as the input of the machine learning model. Use the trained model to predict the task resources in a future time period, and output the prediction results, including the predicted task resources, change trends, etc. According to the prediction results, plan the allocation of computing cores in advance to ensure that there are sufficient computing cores available during high-task-resource periods of tasks.

[0045] The technical solutions in the embodiments of the present application described above have at least the following technical effects or advantages: By real-time monitoring and dynamically adjusting the computing core allocation, the utilization rate of AI acceleration card resources is effectively improved, ensuring the efficient execution of tasks. At the same time, using a machine learning algorithm to predict the task resource trend and plan the resource allocation in advance enhances the predictability and flexibility of the AI acceleration card. In addition, the dynamic adjustment of the task priority mechanism ensures that low-priority tasks can also be processed in a timely manner when resources are abundant, optimizing the task scheduling strategy and improving the overall performance and response speed of the AI acceleration card.

[0046] Example 2: In Example 1, the AI acceleration card has tried to dynamically adjust the computing core allocation by monitoring the computing core utilization rate and execution progress of tasks, in order to optimize resource utilization. However, this preliminary attempt mainly relies on the real-time status of tasks and the resource status of the AI acceleration card, without fully considering the characteristics and differences of the tasks themselves. In other words, Example 1 lacks a mechanism for detailed classification of tasks to more precisely match task requirements with AI acceleration card resources. Therefore, in order to further improve the rationality and efficiency of resource allocation, Example 2 has made important improvements on this basis. The idea of classifying tasks is proposed, and a "rated" number of computing cores is set for each type of task according to the type of task, task resources, and the resource status of the AI acceleration card. At the same time, a real-time waiting counter and a dynamic priority adjustment mechanism are introduced. By real-time monitoring the waiting time of tasks and combining historical task execution data, the actual scheduling priority of tasks is dynamically adjusted. This series of improvements aims to more finely manage the resource allocation during the task execution process, ensure that tasks can obtain computing core resources more fairly and efficiently, and thus significantly improve the overall performance and response speed of the AI acceleration card.

[0047] In some embodiments, the computing core utilization rate and execution progress of tasks are monitored in real time, and the computing core allocation is adjusted in a timely manner according to the dynamic adjustment strategy and in combination with historical task execution data. Step S3 includes: S31, perform a simple classification of several categories of tasks according to the type of task, task resources, and the resource status of the AI acceleration card.

[0048] S32, set a "rated" number of computing cores for each type of task, and allocate computing core resources according to the priority within the same type of tasks.

[0049] Among them, the rated is the number of computing cores that should be allocated to this type of task under normal circumstances.

[0050] S33, set a waiting counter for the tasks under each classification to record the waiting time of the tasks in the queue.

[0051] Specifically, set a waiting counter W for each task to record the waiting time of the task in the queue. Initially, W = 0. Whenever the AI acceleration card performs a resource allocation or task status update, for each task in the waiting queue, increment its waiting counter by 1.

[0052] S34, as the waiting time increases, increase the actual scheduling priority of the task.

[0053] Among them, the actual scheduling priority adjustment formula is: , is the initial priority of the task, which is determined by the task urgency, expected completion time, etc.; is the waiting time of the task in seconds, is the priority adjustment coefficient, which is used to control the influence degree of the waiting time on the priority; is to take the natural logarithm of the waiting time to avoid the priority increasing too fast linearly with the waiting time. Every certain period (such as every n seconds), the actual scheduling priorities of all tasks are updated and reordered to avoid tasks with low priority waiting for a long time without obtaining computing core resources.

[0054] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: Based on the type of tasks, task resources, and the status of AI acceleration card resources, a "rated" number of computing cores is customized for each type of task, ensuring the rationality of resource allocation. At the same time, by introducing a waiting counter and a dynamic priority adjustment mechanism, tasks can dynamically increase their priorities according to their waiting time during the waiting process, thus effectively avoiding the problem that low-priority tasks cannot obtain computing core resources for a long time. This design not only improves the execution efficiency of tasks but also enhances the response speed and fairness of the AI acceleration card. In addition, the adjustment formula for the actual scheduling priority fully considers the waiting time and initial priority of tasks, making resource allocation more accurate and efficient, and further improving the overall performance and stability of the AI acceleration card.

[0055] Embodiment 3: In Embodiment 2, the task dependency graph is used to manage and schedule the execution order of tasks in a multi-task AI acceleration card, and the resource allocation is optimized by clarifying the direct and indirect dependency relationships between tasks. However, although the task dependency graph can effectively represent the logical relationships between tasks, during the actual operation process, the priorities of tasks may change due to various factors, which directly affects the execution order and resource allocation of tasks. Especially in a dynamic environment, the adjustment of task priorities may cause a series of delays or advances in the execution times of dependent tasks, thereby affecting the efficiency of the entire task chain. For example, the priority increase of a critical task may require reclaiming computing core resources from other tasks to meet its immediate execution needs, while tasks with reduced priorities may release resources. Therefore, in order to accurately respond to the dynamic changes in task priorities and accordingly adjust the task execution order and resource allocation, it is particularly important to quickly identify and locally update the affected tasks. Therefore, in order to differentially handle the impact of task priority changes on task execution and resource allocation (to make up for the local deficiencies of the static task dependency graph), it is extremely crucial to accurately identify the affected tasks and timely adjust resources.

[0056] In some embodiments, during real-time monitoring of the computing core utilization rate and execution progress of tasks, according to the dynamic adjustment strategy, combined with historical task execution data, the computing core allocation is adjusted in a timely manner. Step S3 further includes: S35. Construct a task dependency graph for the classified tasks.

[0057] Specifically, traverse the task dependency graph using depth - first search (or breadth - first search, selected according to actual requirements) to identify and label the direct and indirect dependency relationships between tasks. Among them, nodes represent tasks, edges represent the dependency relationships between tasks, and the dependency relationships include prerequisite tasks, successor tasks, and parallel tasks.

[0058] S36. When the priority of a task changes, identify the affected tasks.

[0059] Specifically, when the priority of a certain task changes, immediately trigger the dependency analysis. Using the task dependency graph, starting from the task node with the priority change, trace all its prerequisite tasks (i.e., the tasks connected by the edges pointing to this node) and successor tasks (i.e., the tasks connected by the edges pointed out by this node). The tasks directly dependent on this task (successor tasks) will be affected because their start times may be delayed or advanced due to the priority change of the dependent task; the tasks indirectly dependent on this task may also be affected, and it is necessary to traverse the dependency graph through depth - first (or breadth - first) search to comprehensively identify all possible affected tasks.

[0060] S37. Perform local updates and re - allocate computing cores for the affected tasks.

[0061] Specifically, re - allocating computing cores includes: S37a. For the affected tasks, recalculate the actual scheduling priority of the currently affected tasks according to the formula in step S34.

[0062] S37b. According to the new actual scheduling priority, re - sort the affected tasks to determine the new task execution order and dependency relationships.

[0063] S37c. For tasks with increased priority, try to reclaim computing core resources from tasks with decreased priority; for tasks with decreased priority, promptly release the excess computing core resources they occupy for other tasks to use.

[0064] Among them, if the reclaimed resources are not sufficient to meet the requirements of tasks with increased priority, consider allocating from the reserved resources of the AI acceleration card or idle computing cores to ensure that the "rated" computing core quantity limit of tasks is not violated.

[0065] S37d. During the process of resource re - allocation, keep the total amount of resources allocated within the same type of tasks unchanged, that is, re - allocate resources within the same type of tasks.

[0066] The technical solutions in the embodiments of the present application above have at least the following technical effects or advantages: By constructing a task dependency graph and clarifying the direct and indirect dependencies between tasks, the task management AI accelerator card can more accurately grasp the logical relationships between tasks. When the task priorities change, the AI accelerator card can quickly trigger the dependency analysis, comprehensively identify the affected tasks, and perform local updates and re-allocate computing cores for these tasks. This solution not only improves the flexibility and response speed of task scheduling but also optimizes resource allocation, ensuring the efficiency and stability of task execution. At the same time, by reallocating resources within the same type of tasks, the balance of the total resource allocation is maintained, avoiding resource waste and conflicts. In addition, this solution also considers the coping strategies in case of resource shortages, ensuring that tasks can be successfully executed under the limit of the rated number of computing cores, further improving the efficiency of task completion.

[0067] Embodiment 4: In Embodiment 2, although the task scheduling AI accelerator card can sort and schedule according to the dependencies and priorities of tasks, its flexibility and accuracy may be limited when facing complex task dependency networks and dynamically changing task categories. Specifically, when the task dependencies change, especially when the successor tasks involve cross-category tasks, the original task classification and priority calculation methods may not accurately reflect the actual logical connections and dependency degrees between tasks. For example, due to different business logics or resource requirements, the execution order and priorities of different categories of tasks should be different. However, if only scheduled based on simple dependencies and priorities, these differences may be ignored, resulting in low task execution efficiency or unreasonable resource allocation. Therefore, in order to differentially handle the dependencies between cross-category tasks and improve the flexibility and accuracy of task scheduling, it is extremely crucial to classify the levels of task dependencies, identify and reclassify cross-category tasks, and dynamically adjust task priorities.

[0068] In some embodiments, the local update of the affected tasks includes: S371, divide the dependency levels of task dependencies and assign a level value to each dependency relationship.

[0069] Among them, if the predecessor task for executing task B is A, it means that task B depends on task A (A→B). Then task A is a first-level dependency task and task B is a second-level dependency task.

[0070] S372, when the dependency relationship changes and the successor tasks involve cross-category tasks, identify the cross-category tasks.

[0071] Among them, the cross-category task is defined as follows: Assume that the successor tasks of the first-category task C (with a third-level dependency B→C) are the first-category task D (with a fourth-level dependency) and the second-category task D (with a fourth-level dependency). If the first-category task D and the second-category task D have a parallel dependency on the first-category task C, then the successor tasks of task C are called cross-category tasks.

[0072] Specifically, the changes in task dependency relationships are monitored in real time, such as the addition, deletion, or modification of tasks. When the dependency relationship changes, it is checked whether the successor tasks belong to different task categories. If there are cross-category situations among the successor tasks, that is, the successor tasks and the current task do not belong to the same category, then these cross-category tasks are identified.

[0073] S373, according to the new dependency relationship and task priorities, reclassify the cross-category tasks and correct the simple classification of the tasks.

[0074] Specifically, according to the new dependency relationship, analyze the logical connection and dependency degree between the cross-category tasks and the current task. Based on the analysis results, reclassify the cross-category tasks and group them into the same task category.

[0075] S374, according to the tasks after reclassification, use the actual scheduling priority adjustment calculation formula to recalculate the priorities of the tasks.

[0076] S375, according to the recalculated priorities, sort and schedule the tasks and add them to the analysis of the change trend of task resources in the machine learning algorithm, and pre-adjust the allocation of computing cores.

[0077] The technical solutions in the embodiments of the present application described above have at least the following technical effects or advantages: By carefully dividing the task dependency levels and monitoring the changes in the dependency relationship in real time, the flexibility and accuracy of task scheduling are significantly improved. When the dependency relationship changes, especially when there are cross-category tasks, this solution can quickly identify and reclassify the affected tasks to ensure the rationality and consistency of task classification. At the same time, using the actual scheduling priority adjustment calculation formula to recalculate the task priorities makes the task sorting and scheduling more in line with the actual business requirements and resource conditions.

[0078] The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An AI acceleration card resource scheduling method based on an optimization model, characterized in that, Including: S1. Obtain and record the priority information and task resource data of all tasks in the AI acceleration card; S2. Construct a task resource allocation optimization model according to the task priority and task resource data, and initialize the optimization model parameters; S3. Real-time monitor the computing core utilization rate and execution progress of tasks, and according to the dynamic adjustment strategy, combine the historical task execution data to adjust the computing core allocation in a timely manner; S4. Analyze the change trend of task resources by using machine learning algorithms, and adjust the computing core allocation strategy in advance.

2. The AI acceleration card resource scheduling method based on an optimization model according to claim 1, wherein The dynamic adjustment strategy in S3 specifically includes: S3a. For each task, compare its current computing core utilization rate with the set threshold, and judge whether the task needs to adjust the number of computing cores according to the comparison result; S3b. If the computing core utilization rate of this task is lower than the threshold and the execution progress is lower than 1 / 3 of the task, trigger the dynamic adjustment mechanism and dynamically increase the number of computing cores of this task; S3c. If the computing core utilization rate is higher than the threshold and the execution progress is higher than 2 / 3 of the task, evaluate the impact of reducing the number of computing cores on the execution speed of this task S3d. Dynamically adjust the number of computing cores of this task according to the evaluation result, and release resources to other tasks that need computing resources; S3e. Update the status information such as the number of computing cores and execution progress of tasks in the database, and continue to monitor in real time.

3. The AI acceleration card resource scheduling method based on an optimization model according to claim 1, wherein The specific evaluation method includes: simulating the task execution speed after reducing the number of computing cores, and comparing it with the current execution speed. According to the evaluation result, determine whether to reduce the number of computing cores and the reduced number; assume that n computing cores are reduced, then simulate the task execution time T_new after reduction, and compare it with the current execution time T_current; if T_new is within an acceptable range, then it can be considered to reduce n computing cores.

4. The AI acceleration card resource scheduling method based on an optimization model according to claim 1, characterized in that S4 specifically includes: collecting historical task execution data, including task type, execution time, required number of computing cores, and execution duration; after cleaning and normalizing the data, extract the task type, average task resource volume during the execution period, and dependence relationship characteristics between tasks from the historical data, and construct a feature vector as the input of the machine learning model; use the trained model to predict the task resources in a certain future period, and output the prediction results, including the predicted task resources and change trends; according to the prediction results, plan the computing core allocation in advance.

5. The AI acceleration card resource scheduling method based on an optimization model according to claim 1, wherein S3 also includes: S31. Conduct a simple classification of several categories of tasks according to the task type, task resources, and the resource status of the AI acceleration card; S32. Set a "rated" number of computing cores for each category of tasks, and allocate computing core resources according to the priority within the same category of tasks; where the rated is the number of computing cores that should be allocated to this category of tasks under normal circumstances; S33. Set a waiting counter for each task under each classification to record the waiting time of the task in the queue; S34. As the waiting time increases, improve the actual scheduling priority of the task.

6. The AI acceleration card resource scheduling method based on the optimization model according to claim 5, wherein, The S33 includes: setting a waiting counter W for each task to record the waiting time of the task in the queue; initially, W = 0; whenever the AI acceleration card performs a resource allocation or task status update, for each task in the waiting queue, increment its waiting counter by 1.

7. The AI acceleration card resource scheduling method based on an optimization model according to claim 3, wherein In S34, the actual scheduling priority includes: The actual scheduling priority adjustment formula is: , is the initial priority of the task; is the waiting time of the task (in seconds), is the priority adjustment coefficient, which is used to control the influence degree of the waiting time on the priority; is to take the natural logarithm of the waiting time to avoid the priority growing too fast linearly with the waiting time, update the actual scheduling priorities of all tasks , and reorder them.

8. The AI acceleration card resource scheduling method based on an optimization model according to claim 1, characterized in that, The S3 further includes: S35, constructing a task dependency graph for the classified tasks; S36, identifying the affected tasks when the priority of a task changes; S37, performing local updates on the affected tasks and reallocating computing cores.

9. The AI accelerator card resource scheduling method based on the optimization model according to claim 8, wherein, In the S37, reallocating the computing cores further includes: S37a, for the affected tasks, recalculating the actual scheduling priority of the currently affected task according to the formula in step S34; S37b, reordering the affected tasks according to the new actual scheduling priority to determine the new task execution order and dependency relationship; S37c, for the tasks with increased priority, attempting to reclaim computing core resources from the tasks with decreased priority; for the tasks with decreased priority, promptly releasing the excess computing core resources they occupy for other tasks to use; S37d, during the resource reallocation process, keeping the total amount of resource allocation unchanged among tasks of the same type, that is, reallocating resources within the same type of tasks.

10. The AI acceleration card resource scheduling method based on the optimization model according to claim 8, wherein, The S37, performing local updates on the affected tasks, specifically includes: S371, dividing the dependency levels of the task dependencies and assigning a level value to each type of dependency relationship; S372, when the dependency relationship changes and there are cross-category tasks among the successor tasks, identifying the cross-category tasks; S373, reclassifying the cross-category tasks according to the new dependency relationship and task priority, and correcting the simple classification of the tasks; S374, according to the reclassified tasks, recalculating the priority of the tasks using the actual scheduling priority adjustment calculation formula; S375, according to the recalculated priority, sorting and scheduling the tasks and adding them to the analysis of the change trend of task resources in the machine learning algorithm to pre-adjust the allocation of computing cores.

Citation Information

Patent Citations

  • AI accelerator card resource scheduling method based on double optimization models

    CN118426971A