Multi-model collaborative reasoning AI agent dynamic task scheduling method and system
Patent Information
- Application Number
- CN202610710115.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-05-22
AI Technical Summary
[0005]为解决上述如何准确区分系统噪声与真实拥塞,并基于此做出稳定的调度决策的技术问题,本发明在如下的多个方面中提供方案
1、本发明通过采集历史数据的统计特征并计算风险调节因子与方向开关,对任务耗时均值进行有选择的惩罚调整,得到能够反映真实拥塞而非随机噪声的风险感知预测耗时;再通过构建有向无环图、计算有向路径惯性及动态向上秩,以抑制因子修正得到调度优先级,并依据动态最早完成时刻分配计算节点。整个技术方案能够在系统存在资源竞争和性能波动时准确区分拥塞与噪声,避免因局部抖动导致错误的惩罚和调度震荡,同时通过有向路径惯性压低不稳定支路任务的优先级,防止关键路径频繁翻转,从而实现稳定的调度决策、提高多模型协同推理任务的完成时间可靠性与系统整体吞吐量。
Smart Images

Figure CN122285230B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing. More specifically, this invention relates to a method and system for dynamic task scheduling of AI agents using multi-model collaborative reasoning. Background Technology
[0002] With the deep integration of edge computing and distributed artificial intelligence systems, modern industrial control and complex signal collaborative monitoring scenarios, such as multi-channel physical quantity balance state monitoring and large-scale dynamic target tracking, often deploy a large number of AI agents with independent computing capabilities. When performing tasks, these agents are no longer limited to single logical judgments, but need to invoke multiple inference models with different modalities and complexities, such as dynamic smoothing feature analysis models for high-frequency time-series signals and curvature feature extraction models for two-dimensional spatial images, to perform collaborative inference.
[0003] In actual operation, the task execution time of each computing node can fluctuate randomly due to factors such as resource contention, CPU frequency adjustments, and memory contention. Traditional heuristic scheduling strategies, such as shortest queue first and static HEFT algorithms, cannot distinguish between this random noise and real system congestion. They are prone to misinterpreting instantaneous fluctuations in local execution time as congestion trends, leading to frequent changes in scheduling decisions and causing short-term oscillations. These oscillations not only reduce system throughput but may also cause irregular fluctuations in the overall task completion time, affecting the real-time performance and stability of collaborative inference.
[0004] Therefore, how to accurately distinguish between system noise and real congestion, and make stable scheduling decisions based on this, is a technical problem that urgently needs to be solved in multi-model collaborative reasoning scenarios. Summary of the Invention
[0005] To address the aforementioned technical problem of accurately distinguishing between system noise and actual congestion, and making stable scheduling decisions accordingly, this invention provides solutions in several aspects.
[0006] In the first aspect, the dynamic task scheduling method for AI agents with multi-model collaborative reasoning includes: Historical data of each computing node executing the target task is collected and preprocessed to obtain the statistical characteristics of each computing node. The statistical characteristics include at least the mean and standard deviation of the time taken from the start to the end of the historical target task, the mean and standard deviation of the queue length of the computing node, and the covariance of the time taken from the start to the end of the historical target task and the queue length. The risk adjustment factor is calculated based on the covariance, and the direction switch is determined based on the changing trend of the ratio of the time taken from the start to the end of the current target task to the computational quantity measure for the previous target task. The average time taken from the start to the end of the historical target tasks on the computing node is penalized and adjusted based on the risk adjustment factor and the direction switch to obtain the risk perception prediction time of the computing node. A directed acyclic graph (DAG) of dependencies between target tasks is constructed. Based on the DAG, all directed paths from the initial task to each target task are determined. The stability metric of each directed path is calculated based on the stability of each computing node, and then accumulated to obtain the directed path inertia of the target task. Based on the successor relationship in the DAG and the risk perception prediction time of each computing node, the dynamic upward rank of each target task is calculated. Then, the dynamic upward rank is multiplied by the inhibition factor determined based on the directed path inertia to obtain the scheduling priority of the target task. Based on the scheduling priority from high to low, the computing node that minimizes the earliest dynamic completion time is selected and allocated for each target task.
[0007] Optionally, the calculation of the risk adjustment factor includes: The risk adjustment factor is the numerator calculated by multiplying the covariance of the time taken from the start to the end of the historical target task with the queue length, and the denominator by adding a very small constant to the product of the standard deviation of the time taken from the start to the end of the historical target task with the standard deviation of the queue length of the calculation node.
[0008] Optionally, the process of determining the direction switch includes: The unit computation time for executing the current target task is defined as the ratio of the time taken for the task to be executed from start to finish on the computing node to the computational quantity of the target task. If the unit computation time of the current target task is greater than the unit computation time of the previous target task executed on the same computing node, and the risk adjustment factor is greater than zero, then the direction switch is 1; otherwise, the direction switch is 0.
[0009] Optionally, the risk perception prediction time is the sum of the average time taken for the historical target task from start to finish and the congestion penalty increment; The congestion penalty increment is as follows: when the direction switch is 1 and the risk adjustment factor is greater than 0, the congestion penalty increment is equal to the product of the risk adjustment factor and the standard deviation of the time taken from the start to the end of the historical target task; otherwise, the congestion penalty increment is 0.
[0010] Optionally, the calculation of the directed path inertia includes: For each directed path in a directed acyclic graph from the initial task to the target task along directed edges, sum the stability of the computational nodes assigned to each target task along the path, and then sum the summation results for all directed paths to obtain the directed path inertia.
[0011] Optionally, the inhibition factor is equal to the sum of one minus one divided by one plus the square root of the directed path inertia.
[0012] Optionally, the calculation process of the dynamic upward rank includes: For a target task without a successor task, its dynamic upward rank is equal to the shortest risk perception prediction time of the target task across all computing nodes executing the target task. For other target tasks, their dynamic upward rank is equal to the shortest risk perception prediction time of the target task itself on all computing nodes executing the target task, plus the maximum value of the dynamic upward rank of all its direct successor tasks.
[0013] Optionally, the calculation of the earliest dynamic completion time includes: The moment when the scheduler begins to allocate computing nodes for new target tasks is recorded as the scheduling decision start time. For each computing node, the risk perception prediction time of each target task in the set of tasks to be executed that are queued but have not yet started is accumulated and then added to the scheduling decision start time to obtain the expected time when the computing node will clear the queue from the scheduling decision time. The larger of the expected time and the maximum actual completion time of all predecessor tasks of the new target task is taken as the actual start time. This is then added to the risk perception prediction time of the new target task on the computing node to obtain the earliest dynamic completion time of the new target task on the computing node.
[0014] Secondly, a dynamic task scheduling system for AI agents with multi-model collaborative reasoning includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the dynamic task scheduling method for AI agents with multi-model collaborative reasoning as described in any one of the claims is implemented.
[0015] The present invention has the following beneficial effects: 1. This invention collects statistical characteristics from historical data and calculates risk adjustment factors and direction switches to selectively penalize and adjust the average task time, obtaining a risk-perceived predicted time that reflects real congestion rather than random noise. Then, by constructing a directed acyclic graph, calculating directed path inertia and dynamic upward rank, and using suppression factors to correct scheduling priorities, computation nodes are allocated based on the earliest dynamic completion time. This entire technical solution can accurately distinguish between congestion and noise when the system experiences resource contention and performance fluctuations, avoiding erroneous penalties and scheduling oscillations caused by local jitter. Simultaneously, by using directed path inertia to lower the priority of unstable branch tasks, it prevents frequent critical path flips, thereby achieving stable scheduling decisions, improving the reliability of multi-model collaborative inference task completion time, and increasing the overall system throughput.
[0016] 2. By combining the risk adjustment factor with the direction switch, the prediction time is only increased when there is real congestion and the node performance continues to deteriorate. This avoids erroneous penalties during the node recovery phase or when there is only noise fluctuation, and effectively suppresses short-term scheduling oscillations caused by local jitter. Attached Figure Description
[0017] Figure 1 This is a flowchart of steps S1-S4 in the AI agent dynamic task scheduling method for multi-model collaborative reasoning according to an embodiment of the present invention.
[0018] Figure 2 This is a structural block diagram of the AI agent dynamic task scheduling system for multi-model collaborative reasoning according to an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0020] The scheduling method described in this invention is applied to a distributed system containing multiple heterogeneous computing nodes, each node deploying an agent capable of executing different AI inference models. These models have multiple modalities and varying computational complexities, such as dynamic smoothing feature analysis models for high-frequency time-series signals, curvature feature extraction models for two-dimensional spatial images, and multimodal feature fusion models. A complete collaborative inference task requires multiple models to cooperate and can be broken down into multiple atomic operations with data dependencies, each atomic operation corresponding to one inference operation of a specific model. For ease of description, each atomic operation is collectively referred to as the target task in this invention. The scheduler is responsible for dynamically allocating each target task to appropriate computing nodes for execution, minimizing the overall completion time and suppressing scheduling oscillations.
[0021] Reference Figure 1 The dynamic task scheduling method for AI agents with multi-model collaborative reasoning includes steps S1-S4, as follows: S1: Collect data from each computing node performing the target task and preprocess it to obtain the statistical characteristics of each computing node.
[0022] Each computing node in a multi-model collaborative inference system has independent computing resources, such as CPU and memory, and deploys one or more AI inference models. In this system, the task execution time of each computing node fluctuates randomly due to factors such as resource contention, CPU frequency adjustments, and memory contention. To distinguish this random noise from genuine system congestion, it is first necessary to collect historical task execution data from each computing node.
[0023] Specifically, a monitoring module is deployed on each computing node, which records the following three types of information whenever the computing node completes a target task: The time taken from the start to the end of the target task; the number of tasks already queued in the input queue when the target task arrives at the computing node, including tasks currently being executed, i.e., the queue length; the actual completion time of the target task, which will be used to update the statistical characteristics of the computing node and serve as the precursor completion time for subsequent task scheduling.
[0024] Furthermore, a sliding window is used to calculate the mean and standard deviation of the execution time of historical target tasks on the computing node from start to finish, as well as the mean and standard deviation of the queue length of the computing node. The length of the sliding window is a configurable hyperparameter; in this embodiment, the length is set to 10. The sliding window adopts a first-in, first-out (FIFO) approach, with the starting point being the 10th most recently completed task (the earliest retained task) and the ending point being the most recently completed task. Whenever a computing node completes a new task, the sliding window slides forward one step: discarding the earliest task data in the sliding window and adding the newly completed task data, always maintaining a record of the 10 most recent tasks within the sliding window. If the cumulative number of completed tasks on the computing node is less than 10, the starting point of the window is the first task, and the ending point is the most recently completed task. In this case, the calculation is performed using the actual number of collected data points. The sliding window only slides at a fixed length after the number of tasks reaches 10.
[0025] Furthermore, the covariance between the time taken for the historical target task from start to finish and the queue length is calculated. Calculating the covariance measures the direction of the linear correlation between the time taken and the queue length. When the covariance is greater than 0, it indicates that the longer the queue, the longer the time taken, which is a typical characteristic of real congestion. Subsequent scheduling penalties will be applied to this computing node based on this, so that the scheduler actively avoids congested computing nodes. When the covariance is less than or equal to 0, it indicates that the fluctuation mainly comes from random noise unrelated to the queue, such as CPU frequency adjustments and memory contention. Such fluctuations should not be used as the basis for scheduling penalties, otherwise it will lead to unnecessary scheduling oscillations.
[0026] Based on the data within the sliding window, the statistical characteristics of each computing node were obtained. Specifically, the statistical characteristics of each computing node include: the mean and standard deviation of the time taken from the start to the end of the historical target task, the mean and standard deviation of the queue length of the computing node, and the covariance between the time taken from the start to the end of the historical target task and the queue length.
[0027] S2: Calculate the risk perception prediction time for each computing node to execute the target task based on the statistical characteristics of each computing node.
[0028] In multi-model collaborative inference systems, the task execution time of individual computing nodes often fluctuates randomly due to factors such as resource contention and CPU frequency adjustments. Traditional scheduling methods easily misjudge this noise as genuine congestion, leading to short-term oscillations that frequently change scheduling decisions. Therefore, the goal of this step is to calculate a risk-aware prediction time, which reflects the expected time for a computing node to execute a task in its current state. Based on this prediction, penalties are imposed on truly congested computing nodes, even if their predicted values are higher than the historical average, thereby guiding the scheduler to proactively avoid these nodes.
[0029] First, a risk adjustment factor is calculated to distinguish between noise and genuine congestion. However, the risk adjustment factor alone is insufficient: the performance of a computing node may be recovering, meaning the time taken per unit of computation is decreasing. In this case, even if the risk adjustment factor indicates congestion, the predicted time taken by the computing node should not be increased, because increasing the predicted time will cause the scheduler to consider the computing node busier and avoid it, which will delay the recovery of the computing node. Therefore, it is also necessary to monitor the performance trend of computing nodes and introduce a directional switch for computing nodes to indicate whether the computing node is continuously deteriorating. Only when the computing node is truly congested and its performance is continuously deteriorating should the average time taken from the start to the end of the historical target task of the computing node be increased by the corresponding penalty amount, thus obtaining the risk-aware predicted time taken.
[0030] Specifically, the calculation process for the aforementioned risk adjustment factors is as follows: The covariance of the time taken for a historical target task to complete from start to finish versus the queue length measures the direction of the linear correlation between these two factors; a positive covariance indicates a congestion trend. However, the magnitude of the covariance is affected by both the dimensions and fluctuations of the time taken and the queue length, making direct comparisons between different computing nodes impossible. To eliminate the influence of dimensions and fluctuations, the covariance needs to be normalized to obtain a dimensionless risk adjustment factor. The risk adjustment factor is defined as follows: the numerator is the covariance of the time taken for a historical target task to complete from start to finish versus the queue length; the numerator is the product of the standard deviation of the time taken for a historical target task to complete from start to finish versus the standard deviation of the queue length. To avoid a zero denominator, a very small constant can be added to the denominator, such as... The range of values for the risk adjustment factor obtained in this way is: The closer to 1, the more obvious the actual congestion; if it is close to 0 or negative, it indicates that the congestion is mainly caused by noise.
[0031] Next, to determine whether the performance of a compute node is continuously deteriorating, it is necessary to compare the time taken per unit of computation for executing two consecutive target tasks. The time taken per unit of computation for executing the current target task is defined as the time taken for the task to complete on the compute node from start to finish, compared to the computational load metric of the task. This computational load metric can be directly obtained from the input data processed by the target task: for example, the sequence length when processing time-series signals, the total number of pixels when processing images, and the number of tokens when processing text.
[0032] If the time taken to execute the current target task per unit of computation is greater than the time taken to execute the previous target task on the same computing node, it indicates that the time taken to execute per unit of computation is increasing, and the node's performance is deteriorating. Based on this, a direction switch is defined: If the unit computation time for executing the current target task is greater than the unit computation time for the previous target task executed on the same computing node, and the risk adjustment factor is greater than zero, then the direction switch is defined as 1, allowing the application of penalties; otherwise, the direction switch is defined as 0, forcibly disabling penalties.
[0033] Finally, based on the risk adjustment factor, direction switch, and standard deviation of the time taken from the start to the end of the historical target task on the computing node, the congestion penalty increment is calculated.
[0034] Specifically, when the direction switch is 1 and the risk adjustment factor is greater than 0, the congestion penalty increment is the standard deviation between the risk adjustment factor and the time taken for the target task on the computing node from start to finish; when the direction switch is 0, the congestion penalty increment is directly 0.
[0035] Then, the congestion penalty increment calculated above is added to the average time taken from the start to the end of the historical target task on the computing node to obtain the risk perception prediction time for executing the current target task on the computing node.
[0036] Through the above design, the risk adjustment factor is used to quantify the correlation between time consumption and queue length, and penalties can only be imposed when there is real congestion; the direction switch further requires that the performance of the computing node continues to deteriorate before penalties are enabled. The combination of the two ensures that the computing node will not erroneously increase the predicted time consumption during the recovery phase or when there is only noise fluctuation, thereby effectively suppressing short-term scheduling oscillations caused by local jitter.
[0037] At this point, the risk perception prediction time for executing the target task on each computing node can be obtained.
[0038] S3: Calculate the scheduling priority of the target task based on the risk perception prediction time of executing the target task on each computing node.
[0039] In multi-model collaborative inference, there are often data dependencies between the target tasks corresponding to different models. For example, a multimodal fusion task may need to wait for the image feature extraction task and the time-series signal analysis task to be completed before it can begin; that is, the fusion task depends on the former two. This dependency can be described by a directed acyclic graph.
[0040] Furthermore, the stability of each computing node, i.e., the degree of fluctuation in execution time, affects the overall scheduling stability. Even though S2 has effectively suppressed the interference of local node jitter on single-node prediction, when a task on a non-critical path experiences a temporary increase in its expected remaining time due to short-term jitter, traditional algorithms may still incorrectly promote it to a critical path, leading to frequent switching of scheduling order and causing topology-level oscillations. To address this, this step first uses a directed acyclic graph to describe the dependencies between target tasks, and then introduces inertial features based on node stability to automatically lower the priority of unstable branches, thereby stabilizing the identification of critical paths.
[0041] The specific process of constructing the above directed acyclic graph is as follows: First, the dependencies are determined based on the data flow logic of multi-model inference: if the target task The execution requires the target task If the output is used as input, then there exists a... arrive directed edges ,in Called The precursor mission. Called The subsequent task is to ensure that there are no directed cycles in the directed acyclic graph; if circular dependencies exist, they should be transformed into an acyclic form by introducing intermediate storage or rearranging logic.
[0042] Furthermore, target tasks without predecessor tasks are marked as starting tasks, and target tasks without successor tasks are marked as ending tasks.
[0043] Furthermore, topological sorting of the directed acyclic graph can be performed using the Kahn algorithm or depth-first search to obtain the linear execution order of the target task.
[0044] The above operations yield a directed acyclic graph. , Represents the set of target tasks. It is a set of directed edges.
[0045] After explicitly modeling the collaborative reasoning task as a directed acyclic graph, in order to measure whether the directed path to a certain target task is stable, it is necessary to sum up the stability of the computation nodes corresponding to each target task on the directed path. Here, the directed path refers to a complete path from the initial task to a certain target task along the directed edge.
[0046] The stability of each computing node defines its stability; a higher stability indicates a more stable node. For each directed path from the initial task to a target task along directed edges, the stability of all computing nodes actually assigned to the target task along that path is summed to obtain the stability metric of that directed path. Then, the stability metrics of all directed paths from the initial task to the target task are summed to obtain the directed path inertia reaching the target task. A higher directed path inertia indicates that all upstream directed paths to the target task are more stable overall, and the critical directed path containing the target task is less likely to change due to local fluctuations. The stability of a computing node is equal to the mean of the execution time of historical target tasks on that node from start to finish, divided by the standard deviation.
[0047] Next, we calculate the dynamic upward rank of each target task, which is the optimal expected remaining time from the start of that target task to the end of the entire inference task. For a terminating task (i.e., a target task with no successors), its dynamic upward rank is equal to the shortest risk perception prediction time for executing that target task across all executable computation nodes. For other target tasks, their dynamic upward rank is equal to the shortest risk perception prediction time for that target task across all executable computation nodes, plus the maximum of the dynamic upward ranks of all its direct successors. This is because successor tasks can be executed in parallel, and the overall completion time depends on the branch with the longest execution time.
[0048] Finally, the dynamic upward rank of each target task calculated above is multiplied by a suppression factor related to the inertia of the directed path to the corresponding target task, so that the priority of the target task on the unstable directed path is reduced, thereby avoiding frequent reversals of the critical directed path.
[0049] The above-mentioned inhibitory factors satisfy the following relationship: In the formula, The above-mentioned inhibitory factors, The directed path inertia to reach the corresponding target task.
[0050] This suppression factor quantifies the impact of directed path stability on scheduling priority. When the directed path has high inertia (i.e., is stable), the suppression factor is close to 1, the dynamic upward rank is almost unaffected, and the priority remains unchanged. When the directed path has low inertia (i.e., is unstable), the suppression factor is significantly less than 1, the dynamic upward rank is greatly suppressed, thus reducing the likelihood of the target task being scheduled first. In this way, tasks on unstable directed paths will automatically be pushed back, avoiding frequent flips of the critical path due to local jitter.
[0051] Thus, the scheduling priorities of all target tasks can be obtained through the above calculations.
[0052] S4: Based on the scheduling priority of the target task, combined with the completion times of backlogged tasks and predecessor tasks, predict the earliest dynamic completion time and allocate execution computing nodes to the new target task.
[0053] After obtaining the scheduling priority of each target task, two practical problems still need to be solved: how to accurately estimate the real impact of the backlog of tasks on the computing node (queue length alone is not enough), and how to combine the completion time of the predecessor task and the idle time of the computing node to select the best computing node that minimizes the task completion time.
[0054] To predict when a new target task will begin execution on a specific computing node, it is necessary to first estimate the expected time required for all target tasks currently queued but not yet started on that computing node to complete. The moment when the scheduler begins allocating computing nodes for the new target task is recorded as the scheduling decision start time. .
[0055] For any given computing node, obtain the set of queued but not yet executed tasks. These tasks are arranged in the order they were assigned to that computing node by the scheduler. Each time the scheduler makes a decision, it selects the highest-priority task from the ready tasks and assigns it to a computing node. This task is then appended to the end of the local execution queue of that computing node. Therefore, the task sets in the set of tasks to be executed naturally maintain their scheduled order. The risk perception prediction time of each task in the set of tasks to be executed is summed and then compared with... Adding them together gives the expected time when the computing node will start clearing the queue from the scheduling decision time.
[0056] Furthermore, the actual completion time of a target task is constrained by two factors: first, the available time of the computing nodes allocated to the target task, i.e., the queue clearing time; and second, the actual completion time of all predecessor tasks of the target task. Because a new target task must wait for all predecessor tasks to complete before it can begin execution, its start time cannot be earlier than the maximum value of the completion times of all predecessor tasks. For predecessor tasks that have already been completed, their recorded actual completion times are used directly; for those that have not yet been completed, their corresponding risk perception prediction time is used for estimation.
[0057] Therefore, for a new target task, the larger of the expected time when the computing node starts clearing the queue from the scheduling decision time and the maximum value of the actual completion time of all the predecessor tasks of the new target task is taken as the actual start time. The risk perception prediction time of the new target task on the computing node is added to the actual start time. The sum is taken as the earliest dynamic completion time of the new target task on the computing node.
[0058] During scheduling decisions, target tasks are retrieved from the ready tasks in descending order of scheduling priority. For each target task, the computing node with the smallest dynamic earliest completion time is selected for execution, and the target task is appended to the end of the computing node's queue, recording its expected start and completion times. After a target task is actually completed, its actual data is fed back to the sliding window in S1 above, updating statistics and forming a closed-loop adaptive scheduling mechanism. The scheduler operates in an event-driven manner, triggering a decision whenever a target task is completed or a new target task arrives, repeating the above process until all target tasks are completed.
[0059] Through the above steps, the specific embodiment of the present invention realizes dynamic scheduling of each target task in the multi-model collaborative reasoning task, significantly suppresses short-term scheduling oscillations and critical path flipping, and improves the stability of the completion time of the collaborative reasoning task.
[0060] This invention also provides a dynamic task scheduling system for AI agents with multi-model collaborative reasoning. For example... Figure 2As shown, the system includes a package and a memory, the memory storing computer program instructions, which, when executed by the processor, implement the AI agent dynamic task scheduling method for multi-model collaborative reasoning according to the first aspect of the present invention.
[0061] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
[0062] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A dynamic task scheduling method for AI agents based on multi-model collaborative reasoning, characterized in that, include: Historical data of each computing node executing the target task is collected and preprocessed to obtain the statistical characteristics of each computing node. The statistical characteristics include at least the mean and standard deviation of the time taken from the start to the end of the historical target task, the mean and standard deviation of the queue length of the computing node, and the covariance of the time taken from the start to the end of the historical target task and the queue length. The risk adjustment factor is calculated based on the covariance, including: taking the covariance of the time taken from the start to the end of the historical target task to the queue length as the numerator, and adding a very small constant to the product of the standard deviation of the time taken from the start to the end of the historical target task and the standard deviation of the queue length of the calculation node as the denominator. The resulting quotient is the risk adjustment factor. The direction switch is determined based on the changing trend of the ratio of the time taken from the start to the end of the current sub-target task to the computational quantity metric between the previous target task and the current sub-target task. The computational quantity metric is the sequence length when processing time-series signals, the total number of pixels when processing images, or the number of words when processing text. Based on the risk adjustment factor and the direction switch, the average time taken from the start to the end of the historical target task on the computing node is penalized and adjusted to obtain the risk perception prediction time of the computing node. The risk perception prediction time is the sum of the average time taken from the start to the end of the historical target task and the congestion penalty increment. The congestion penalty increment is: when the direction switch is 1 and the risk adjustment factor is greater than 0, the congestion penalty increment is equal to the product of the risk adjustment factor and the standard deviation of the time taken from the start to the end of the historical target task; otherwise, the congestion penalty increment is 0. Construct a directed acyclic graph of dependencies between target tasks. Based on the directed acyclic graph, determine all directed paths from the initial task to each target task. Sum the stability of the computing nodes to which all target tasks are actually assigned on the directed path to obtain the stability metric of the directed path. Then, accumulate the stability metrics of all directed paths to obtain the directed path inertia of the target task. The stability of a computing node is equal to the mean of the time taken by historical target tasks from start to finish on that computing node divided by the standard deviation. The dynamic upward rank of each target task is calculated based on the successor relationships in the directed acyclic graph and the risk perception prediction time of each computing node. This includes: for target tasks without successor tasks, the dynamic upward rank is equal to the shortest risk perception prediction time of the target task on all computing nodes executing the target task; for other target tasks, the dynamic upward rank is equal to the shortest risk perception prediction time of the target task itself on all computing nodes executing the target task, plus the maximum value among the dynamic upward ranks of all its direct successor tasks. The scheduling priority of the target task is obtained by multiplying the dynamic upward rank by the inhibition factor determined according to the directed path inertia. The formula for calculating the inhibition factor is as follows: In the formula, As an inhibitor, Directed path inertia to reach the corresponding target task; Based on the scheduling priority from high to low, the computing node that minimizes the earliest dynamic completion time is selected and allocated for each target task.
2. The AI agent dynamic task scheduling method for multi-model collaborative reasoning according to claim 1, characterized in that, The process of determining the direction switch includes: The unit computation time for executing the current target task is defined as the ratio of the time taken for the target task to be executed from start to finish on the computing node to the computational quantity of the task. If the unit computation time of the current target task is greater than the unit computation time of the previous target task executed on the same computing node, and the risk adjustment factor is greater than zero, then the direction switch is 1; otherwise, the direction switch is 0.
3. The AI agent dynamic task scheduling method for multi-model collaborative reasoning according to claim 1, characterized in that, The calculation of the earliest dynamic completion time includes: The moment when the scheduler begins to allocate computing nodes for new target tasks is recorded as the scheduling decision start time. For each computing node, the risk perception prediction time of each target task in the set of tasks to be executed that are queued but have not yet started is accumulated and then added to the scheduling decision start time to obtain the expected time when the computing node will clear the queue from the scheduling decision time. The larger of the expected time and the maximum actual completion time of all predecessor tasks of the new target task is taken as the actual start time. This is then added to the risk perception prediction time of the new target task on the computing node to obtain the earliest dynamic completion time of the new target task on the computing node.
4. A dynamic task scheduling system for AI agents with multi-model collaborative reasoning, characterized in that, include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement the AI agent dynamic task scheduling method for multi-model collaborative reasoning according to any one of claims 1-3.
Citation Information
Patent Citations
Large model reasoning scheduling method and system in multi-node heterogeneous environment
CN121116653A
Industrial wireless network trusted scheduling method and device based on dynamic block chain
CN121568221A