A heterogeneous computing resource dynamic allocation method for vertical scene

CN122614588BActive Publication Date: 2026-09-29CHENGDU WEIAI DIGITAL MOMENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611055439.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-16
Publication Date
2026-09-29
Estimated Expiration
2046-07-16

AI Technical Summary

Technical Problem

[0006]为解决现有负载预测调度方案无法感知会话级请求簇发前兆,导致借用算力在神经网络处理单元负载骤升时无法及时回收的技术问题,本发明提出了一种面向垂类场景的异构计算资源动态分配方法

Benefits of technology

[0023]本发明通过在异构计算架构的资源调度架构中融合会话层请求特征与时序预测模块,使得调度系统具备了对亚秒级负载突变前兆的感知能力,弥补了分钟级负载预测在面对突发请求时的感知盲区。利用动态修正的预测负载与借用任务的最大剩余完成时间和预设的回收时间开销的实时对比,调度系统能够自适应地收紧安全借用额度上限,使神经网络处理单元算力在负载突升前能够以自然完成方式平稳回收,避免了硬中断对借用任务执行状态的破坏,提升了垂类场景下边缘侧一体化装备的运行稳定性与资源利用率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122614588B_ABST
    Figure CN122614588B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of heterogeneous computing, and particularly relates to a heterogeneous computing resource dynamic allocation method for vertical scene, which comprises the following steps: extracting request features containing the number of newly added sessions and the coefficient of variation of request arrival interval, and combining the historical session arrival rate average to construct a risk coefficient; inputting the load features into a time series prediction module to obtain an original prediction sequence, and modifying the original prediction sequence based on the coefficient to obtain a modified load with linear decay according to the prediction step; estimating the early warning time margin according to the modified load, combining the maximum remaining completion time of borrowed tasks, the recovery time overhead and the smooth buffer time to determine the upper limit of the safe borrowing quota; if the margin is insufficient, triggering soft recovery to migrate the intermediate state to the central processor for degraded execution to release computing power. The application can improve the perception ability of load mutation precursors and stabilize the recovery of computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heterogeneous computing technology. More specifically, this invention relates to a method for dynamically allocating heterogeneous computing resources for vertical computing scenarios. Background Technology

[0002] In vertical application scenarios such as intelligent tutoring in education, assisted medical diagnosis, and digital navigation in cultural tourism, edge-side integrated equipment typically employs a heterogeneous computing architecture consisting of a neural network processing unit (NN unit) and a central processing unit (CPU). The NN unit handles large-scale model inference tasks, while the CPU handles vertical application service tasks. Because the computational demands of these two types of tasks are staggered over time, vertical application domains can utilize some of the NN unit's computational power for computationally intensive operations such as scene rendering and vector retrieval when the NN unit's load is low, thereby improving overall resource utilization.

[0003] To mitigate resource waste caused by static allocation, existing solutions typically introduce load forecasting technology based on historical time-series data to adjust the allocation quota of the borrowing pool in advance. When the load is predicted to increase, the borrowing amount is reduced in advance to achieve preventative response to peak loads.

[0004] Currently, by introducing load forecasting and dynamic resource management technologies, task scheduling and migration control between devices are performed based on the resource pressure indicators of the devices. Chinese patent application CN122173201A discloses a unified orchestration and seamless migration method and system for heterogeneous computing pools. This application utilizes the collection of resource pressure indicators from source devices and triggers a pre-replication process before the source devices reach resource saturation, enabling the compression of service interruption time and achieving near-seamless migration of inference tasks when heterogeneous device resources approach saturation.

[0005] However, the aforementioned technical solutions can alleviate the task scheduling and migration problems when computing node resources are nearing saturation to some extent, enabling dynamic resource response. However, in vertical application scenarios, existing technologies suffer from several drawbacks. User inference requests are not uniformly distributed but rather exhibit strong clustering in sessions. For example, a user initiating an intelligent Q&A session often initiates multiple highly relevant follow-up questions within the next few tens of seconds. This clustering of requests within a session can cause the neural network processing unit load to surge from low to full load in less than a second. Existing monitoring schemes based on smoothing resource pressure indicators or time-series prediction modules trained on minute-level historical sequences have a much larger time perception range than the precursor signals of such sub-second-level abrupt changes, making it impossible to provide effective early warning output before such a change occurs. When this load abrupt change is detected, the computing power of the neural network processing unit already lent to the vertical application domain is executing rendering or retrieval tasks. A forced hard interruption at this time would disrupt the execution state of the borrowed tasks; waiting for the tasks to complete naturally would cause backlog and delays in the neural network processing unit's inference queue, severely degrading the quality of core inference services. This problem is further amplified during peak periods with high concurrent session counts. Summary of the Invention

[0006] To address the technical problem that existing load prediction and scheduling schemes cannot detect the precursors of session-level request clusters, resulting in the inability to promptly reclaim borrowed computing power when the load on neural network processing units surges, this invention proposes a dynamic allocation method for heterogeneous computing resources in vertical scenarios.

[0007] This invention provides a method for dynamically allocating heterogeneous computing resources for vertical scenarios. It is applied to heterogeneous computing architectures including central processing units (CPUs) and neural network processing units (NN units), and allows tasks carried by the CPU to borrow computing power from NN units for execution. The method includes: S1, extracting session layer request features including the number of new sessions and the coefficient of variation of request arrival intervals, and constructing a session burst risk coefficient by combining it with the historical average session arrival rate; S2, obtaining the original prediction sequence output by the time-series prediction module, and correcting the original prediction sequence based on the session burst risk coefficient by linearly decreasing the prediction step size, thus obtaining a corrected prediction load; S3, estimating the warning time margin based on the corrected prediction load, and dynamically determining the upper limit of the safe borrowing quota by combining the maximum remaining completion time of the borrowed task, the preset recycling time overhead, and the smoothing buffer time, and restricting new borrowing task applications based on the safe borrowing quota upper limit; S4, if the warning time margin is less than or equal to the sum of the maximum remaining completion time and the recycling time overhead, triggering a soft recycling process, migrating the intermediate state of the borrowed task to the CPU for degraded execution to release the computing power of the NN unit.

[0008] This invention introduces session layer behavioral features into the computing power allocation link, enabling the scheduling system to have real-time perception capabilities for sub-second load surge precursors. This compensates for the inherent blind spots in time granularity of existing minute-level time-series prediction modules, allowing the computing power of neural network processing units to be smoothly recovered in a natural manner before a load surge, avoiding the disruption of the borrowed task execution state by hard interrupts.

[0009] Preferably, the step of extracting session-layer request features including the number of new sessions and the coefficient of variation of request arrival interval, and constructing a session burst risk coefficient by combining the historical session arrival rate mean, includes: collecting the number of new sessions within the current active session window, and statistically analyzing the request arrival interval sequence of each session within the time window; calculating the ratio of the standard deviation to the mean of the request arrival interval sequence to obtain the coefficient of variation of request arrival interval; and multiplying the ratio of the number of new sessions to the historical session arrival rate mean by a multiplier, which is the sum of the coefficient of variation of request arrival interval and 1.

[0010] This invention can accurately measure the degree of clustering of requests within a session by collecting the number of new sessions and the interval sequence of request arrivals in real time, thereby constructing a session burst risk coefficient that reflects the precursors of sub-second mutations, effectively improving the sensitivity of short-term load prediction to session burst events.

[0011] Preferably, the method for determining the length of the time window includes: continuously recording a full sample of the interval between adjacent requests arriving in a historical session; updating the full sample on a daily basis and extracting percentile values; and multiplying the percentile values ​​by a preset multiple to obtain the length of the time window.

[0012] Preferably, the step of correcting the original prediction sequence based on the session burst risk coefficient by linearly decaying with increasing prediction step size to obtain the corrected prediction load includes: determining whether the session burst risk coefficient is greater than the historical mean benchmark; if it is greater than the historical mean benchmark, calculating the difference between the session burst risk coefficient and the historical mean benchmark, and multiplying the difference, the risk amplification coefficient, and the linear decay factor to obtain the correction term; using the sum of 1 and the correction term as a multiplication factor, multiplying the original prediction sequence by the multiplication factor to obtain the corrected prediction load.

[0013] This invention incorporates the session burst risk coefficient into the recent step correction of the original prediction sequence in a linear decay manner, so that the corrected prediction load can be synchronously increased after the emergence of session burst precursors. This effectively shortens the early warning response time lag while preserving the long-term prediction stability of the original time series module for macro load trends.

[0014] Preferably, the method for obtaining the risk amplification coefficient includes: obtaining a sequence of session sudden risk coefficients prior to the occurrence of a historical sudden event as an independent variable; taking the difference between the actual observed load peak and the original predicted sequence as a dependent variable; and obtaining the risk amplification coefficient by solving the least squares linear regression.

[0015] Preferably, estimating the warning time margin based on the modified predicted load includes: statistically analyzing the load level value sequence when request queuing backlog occurs in historical operating data to obtain the mean and standard deviation; subtracting the standard deviation from the mean to obtain the recovery trigger threshold; and calculating the time length from the current moment until the modified predicted load first exceeds the recovery trigger threshold to obtain the warning time margin.

[0016] Preferably, the step of dynamically determining the upper limit of the safe borrowing limit by combining the maximum remaining completion time of the borrowing task, the preset recovery time cost, and the smoothing buffer time includes: for each task currently running on the computing power of the borrowing neural network processing unit, tracking the current operator layer position of the borrowing task in the execution graph; querying the pre-stored operator layer time consumption file, accumulating the historical average execution time of each operator layer to obtain the remaining completion time, comparing the remaining completion time of all current borrowing tasks, and extracting the maximum value as the maximum remaining completion time; calculating the difference between the warning time margin and the sum of the maximum remaining completion time and the recovery time cost, and dividing the difference by the preset smoothing buffer time to obtain a ratio value; comparing the ratio value with 1 and taking the smaller value, and then comparing it with 0 and taking the larger value as the truncation coefficient, and multiplying the truncation coefficient by the upper limit of the total computing power of the borrowing pool to obtain the upper limit of the safe borrowing limit.

[0017] Preferably, the triggering of the soft recycling process, which migrates the intermediate state of the borrowed task to the central processing unit for degraded execution to release the computing power of the neural network processing unit, includes: pausing the task execution on the neural network processing unit side after the currently executing operator layer has finished running; writing the output tensor and execution progress pointer of the current layer into the shared memory mapping region to form a checkpoint snapshot; reading the checkpoint snapshot into the central processing unit memory, and having the central processing unit continue to complete the calculation of the remaining operator layers in degraded mode to release the computing power of the neural network processing unit.

[0018] This invention, through a checkpoint-driven soft recycling process, ensures that the intermediate execution state of the borrowing task is fully preserved and seamlessly transferred to the central processing unit for continued execution when the warning time margin is insufficient to support the natural completion of the borrowing task. This achieves lossless and rapid release of the computing power of the neural network processing unit.

[0019] Preferably, the step of writing the output tensor and execution progress pointer of the current layer into the shared memory mapping region to form a checkpoint snapshot includes: monitoring the load of the neural network processing unit during the checkpoint snapshot writing process; if the load of the neural network processing unit triggers an emergency threshold, then executing a hard interrupt fallback path; discarding the current borrowed task and sending a task failure notification to the application service side to resubmit it to the central processing unit queue for execution, ultimately forming the abnormal handling result of the checkpoint snapshot.

[0020] This invention avoids the risk of severely congesting the core inference resources of the neural network processing unit due to forcibly waiting for the snapshot to be written, by promptly executing the hard interrupt fallback path and resubmitting the task. Thus, the emergency fallback path covers extreme and sudden working conditions, ensuring the fault tolerance capability of the entire soft recycling process under abnormal circumstances.

[0021] Preferably, after releasing the computing power of the neural network processing unit, the method further includes: continuously updating the corrected prediction load; when the corrected prediction load within the future warning time margin is continuously higher than the recycling trigger threshold, the upper limit of the secure borrowing quota is maintained at 0, and no new borrowing applications are approved; when the corrected prediction load drops to a safe range below the recycling trigger threshold and the session burst risk coefficient falls below the average historical session arrival rate, the borrowing pool is gradually opened, and the borrowing is gradually opened according to the preset secure borrowing quota relationship, restoring the dynamic allocation of heterogeneous computing resources.

[0022] The beneficial effects of this invention are as follows:

[0023] This invention integrates session layer request features and timing prediction modules into the resource scheduling architecture of a heterogeneous computing architecture. This enables the scheduling system to perceive sub-second-level load surges, compensating for the blind spots in minute-level load prediction when facing sudden requests. By dynamically correcting the predicted load and comparing it in real time with the maximum remaining completion time of the borrowed task and the preset recovery time overhead, the scheduling system can adaptively tighten the upper limit of the safe borrowing quota. This allows the computing power of the neural network processing unit to be smoothly recovered naturally before the load surge, avoiding the disruption of the borrowing task execution state by hard interrupts. This improves the operational stability and resource utilization of edge-side integrated equipment in vertical scenarios.

[0024] Furthermore, the soft recycling process and hard interruption fallback path designed for extreme and sudden operating conditions ensure the continuity of the borrowed task's state and the fault tolerance of the scheduling system. The key parameters of each stage in the heterogeneous computing resource dynamic allocation method are based on objective historical data and real-time dynamic state evolution, enabling adaptive deployment in different vertical scenarios without manual intervention. This provides underlying computing power support for high-concurrency applications such as intelligent tutoring in education and assisted medical diagnosis. Attached Figure Description

[0025] Figure 1 This is a flowchart of a method for dynamically allocating heterogeneous computing resources for vertical scenarios in this invention; Figure 2 This is a line graph comparing the dynamic changes in inference response delay in this invention. Detailed Implementation

[0026] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0027] This invention discloses a method for dynamically allocating heterogeneous computing resources for vertical computing scenarios, referring to... Figure 1 This includes steps S1-S4: S1. Construct a session sudden risk coefficient.

[0028] In an optional embodiment, session layer request features are extracted and combined with the average historical session arrival rate to construct a session burst risk coefficient.

[0029] The system continuously collects six hardware load metrics at a preset sampling period: inference latency, memory usage, and inference batch size from the neural network processing unit's inference domain; and core utilization, memory usage, and I / O throughput from the central processing unit's application domain. For example, the sampling period can be set to 100ms. Since these six hardware load metrics reflect the current state of the device, they lack early warning capabilities for behavior patterns such as clustered requests within a session, which are prepared within hundreds of milliseconds. Therefore, within each sampling period, it is also necessary to synchronously extract session-level request features, i.e., collect the number of new sessions added per unit time within the currently active session window. Furthermore, for all currently active sessions, a request arrival interval sequence is calculated, consisting of the differences in arrival times of adjacent requests within the most recent time window for each session, and the coefficient of variation of this request arrival interval sequence is then calculated. It should be noted that the coefficient of variation of the request arrival interval... The dimensionless ratio is the ratio of the standard deviation of the requested arrival interval sequence to the mean of the requested arrival interval sequence.

[0030] To objectively reflect the typical duration of a complete request cluster cycle, the system continuously records a full sample of the arrival intervals of adjacent requests within historical sessions. It then uses a daily rolling update method to extract percentile values, multiplying these percentile values ​​by a preset factor to determine the length of the current time window. For example, the 95th percentile value can be used, and the preset multiple can be 3, that is, the time window length is 3 times the 95th percentile value. The length of the time window is sufficient to fully accommodate a complete clustering cycle from sparse requests to dense requests and then to the end of the dense request cycle.

[0031] Requested arrival interval coefficient of variation This characterizes the uniformity of the request distribution. Request arrival interval variation coefficient. A smaller value indicates that the sample values ​​of the interval sequence tend to be consistent, the distribution is more uniform, and the standard deviation is close to 0, thus achieving the desired interval coefficient of variation. Approaching 0; requesting the interval variation coefficient to reach A larger value indicates the coexistence of extremely short and extremely long intervals, with the variance of the interval distribution being much greater than the mean, suggesting the presence of clustering. Therefore, a value representing the interval variation coefficient is required. Significantly greater than 1.

[0032] Specifically, based on the analysis of high-concurrency request data from a large number of vertical scenarios, it is proven that the degree of request clustering within a session is positively correlated with the number of new sessions per unit time and the fluctuation of request arrival intervals. Therefore, based on empirical data, it is deduced that the number of new sessions... Average arrival rate of historical sessions The ratio multiplied by a multiplier, which is the coefficient of variation of the requested interval. The sum of 1 and 1 gives the session burst risk coefficient at the current moment. :

[0033] in, The risk factor for sudden changes in the session; The number of new sessions added per unit of time within the currently active session window; This is the average historical session arrival rate of the system during the current time period, which is obtained by the system through continuous rolling statistics of historical sampled data. The coefficient of variation of the requested arrival interval sequence.

[0034] When the session sudden risk factor A value greater than 1 indicates that the current session activity intensity has exceeded the historical average for the same period and that the request distribution is uneven, and that the session burst risk coefficient is high. The larger the value, the higher the risk of sudden events.

[0035] Due to the historical session arrival rate average Based on objective historical statistics, please provide the reach interval coefficient of variation. Length of the time window derived from real-time request sequence computation The objective quantile derivation from the historical interval distribution shows that the differences in vertical scenarios can naturally converge to their respective suitable parameter values ​​through historical data of different scenarios.

[0036] In this way, by introducing the session behavior dimension into the computing power allocation link, the scheduling system can obtain a real-time risk assessment value at the end of each sampling period, thereby providing sub-second-level aberration precursor perception input for subsequent prediction correction, breaking through the inherent limitation of minute-level historical sequences in terms of time granularity.

[0037] S2, Generate corrected predicted load.

[0038] In an optional embodiment, the original prediction sequence output by the time series prediction module is corrected by a recent step based on the session burst risk coefficient to obtain the corrected prediction load.

[0039] The temporal prediction module employs an attention-based temporal network with an encoder-decoder architecture. Using a 6-dimensional hardware load feature vector of a preset duration (e.g., 1 hour) as input, a bidirectional long short-term memory network traverses the historical sequence in both forward and backward directions to extract the hidden state at each time step. Subsequently, an additive attention layer uses the decoder's current hidden state as a query to calculate the contribution weights of each historical time step's hidden state to the current decoding step, and generates a context vector after weighted summation. The decoder uses the context vector and the output values ​​of the predicted time steps as input to autoregressively generate the original prediction sequence for the future preset duration. For example, the preset duration can be set to 5 minutes in the future. The value ranges from 1 to the total number of prediction steps. .

[0040] Before the time-series prediction module performs online forward inference, an offline training phase must be completed to obtain effective model parameters. Specifically, the 6-dimensional hardware load feature vector sequence collected from historical operating data over the past hour is used as the training set input, and the actual load level values ​​corresponding to the next 5 minutes on the time axis are extracted as supervision labels. During training iterations, mean squared error is used as the loss function to calculate the error between the autoregressive generated prediction sequence and the supervision labels. The optimizer is configured to calculate the gradient based on this error and backpropagate to update the weight matrices in the encoder and decoder networks until the value of the loss function decreases and tends to stabilize, reaching the preset convergence condition, thus solidifying the model parameters.

[0041] Since the training samples of the original prediction sequence come from minute-level historical data, the performance of session bursts in the historical sequence is often diluted by the stationary values ​​of adjacent sampling steps. This causes the output response of the time series prediction module to the precursors of burst events to lag significantly behind the actual moment of the abrupt change. Therefore, based on the session burst risk coefficient, the original prediction sequence output by the time series prediction module is corrected by linearly decreasing with the increase of the prediction step size, resulting in the corrected prediction load. The specific logic is as follows: determine the session burst risk coefficient. If the risk factor is greater than the historical average (baseline value is 1), calculate the difference between the session's sudden risk coefficient and the historical average, and then combine this difference with the risk amplification coefficient. and linear decay factor Multiplying yields the correction term; the sum of 1 and this correction term is used as a multiplication factor, and the original prediction sequence is multiplied by this multiplication factor to obtain the corrected prediction load. To eliminate the risk of over-correction for normal states through a linear decay mechanism, and combining the risk amplification coefficient obtained from historical event replays, the corrected prediction load relationship is constructed through theoretical derivation of existing prediction time series models:

[0042] in, This is the corrected predicted load; This is the original predicted sequence; This is a risk amplification factor. The risk factor for sudden changes in the session; To predict the total number of steps; This is the current prediction step number; It is a linear decay factor.

[0043] It should be noted that when When, linear decay factor When the value is close to 1, When, linear decay factor The decay to 0 means that the impact of session burst corrections is mainly concentrated in recent prediction steps, decreasing linearly to 0 with the prediction step size. This ensures that adjustments are triggered only when the risk factor for sudden events exceeds the historical average, avoiding overly conservative approaches under normal conditions. Risk amplification factor. Objective calibration is performed by replaying historical emergencies. Specifically, the process involves obtaining a sequence of session-based emergency risk coefficients prior to the event as the independent variable, and using the difference between the observed peak load and the original predicted sequence as the dependent variable. Least squares linear regression is then used to solve for the risk amplification coefficient. This minimizes the corrected short-term prediction error over the set of historical emergencies.

[0044] Compared to the original prediction output, which maintains a stable prediction curve when a sudden session outbreak occurs, the modified prediction output adds a linear decay correction term based on the real-time session outbreak risk coefficient to the multiplication factor. This causes the recent prediction load to be adjusted synchronously with the intensity of the session outbreak, avoiding the problem of prediction response lag. Moreover, the correction term naturally decays to 0 after the prediction step exceeds a certain range, without affecting the judgment of the macro trend.

[0045] Thus, the forecast load is corrected by incorporating the session burst risk coefficient into the recent step correction of the original forecast sequence in a linearly decaying manner. The frequency can be synchronously increased in the next sampling period after the emergence of the premonition of session clusters, which shortens the warning response time lag while preserving the long-term prediction stability of the original time series module for macro load trends.

[0046] S3. Determine the maximum safe borrowing limit.

[0047] In an optional embodiment, the warning time margin is estimated based on the modified predicted load, and the upper limit of the safe borrowing limit is dynamically determined in combination with the remaining completion time of the borrowing task.

[0048] Specifically, the load level values ​​during periods of request queuing and backlog in historical operational data are statistically analyzed to obtain the average value. with standard deviation To provide sufficient reaction time before the predicted load reaches the danger level, a recovery trigger threshold is constructed based on the lower boundary assessment principle in statistics. :

[0049] in, The threshold for triggering recycling; The mean; The standard deviation is denoted as .

[0050] Calculate the time from the current moment until the corrected predicted load first exceeds the recycling trigger threshold, and finally obtain the warning time margin. If the predicted load does not exceed the recycling trigger threshold throughout the entire prediction window. Then the warning time margin The upper limit of the total duration of the prediction window is taken. Since the recycling trigger threshold is taken as the lower boundary of the historical backlog load rather than the average value, when the predicted load approaches the recycling trigger threshold, the system has already entered the compression operation in advance, thus reserving time margin for the natural completion of the borrowing task.

[0051] For each task currently running on the borrowed neural network processing unit's computing power, track the current operator layer position of the borrowed task in the execution graph; query the pre-stored operator layer time consumption records, accumulate the historical average execution time of each operator layer, and obtain the remaining completion time. Compare the remaining completion times of all current borrowing tasks. Extract the maximum value as the maximum remaining completion time. When each borrowing task is initiated, the pre-stored operator layer time consumption file is queried according to the request type to calculate the initial remaining completion time. It is updated in real time as the task progresses. The operator layer time consumption file is updated by an exponentially weighted moving average of the historical execution times of similar tasks, so that the time consumption estimate can naturally track the slow drift of hardware status.

[0052] Then, calculate the warning time margin. With maximum remaining completion time and recycling time overhead The difference between the sums is calculated and then divided by the preset smoothing buffer time. The ratio value is obtained. To achieve a smooth contraction of the borrowing quota and avoid drastic fluctuations in computing power allocation, the ratio is based on the relationship between the early warning time margin and the total cost of task recovery, combined with a preset smoothing buffer time. and the maximum total computing power of the borrowing pool Establish a secure borrowing limit:

[0053] in, To ensure safe borrowing of the maximum amount; This is the maximum total computing power limit of the pool; This is a margin of time for early warning; The maximum remaining completion time; The fixed overhead time required for the reclamation operation itself is pre-stored by the system during the initialization phase by performing several benchmark measurements on the soft reclamation process and taking the average. The soft reclamation process includes three sub-operations: operator layer pause, tensor writing to the shared memory mapping region, and CPU memory reading. For example, the fixed overhead time... A typical value can be 50ms; The set smooth buffer time, typically set to 200ms to 500ms, is used to provide a gradually compressed buffer window before triggering hard reclamation.

[0054] When the warning time margin is much greater than the sum of the maximum remaining task completion time, the recovery time overhead, and the smoothing buffer time, Greater than 1, after and If the truncation coefficient is 1 after function truncation, then the maximum safe borrowing limit is reached. Take the maximum total computing power of the borrow pool The borrowing pool can be fully opened; when the warning time margin narrows and the drop enters the smoothing buffer time. When the cutoff coefficient starts to be less than 1 and gradually decreases within the specified range, the maximum safe borrowing limit is reached. Linear compression stops approving new borrowing applications, allowing existing borrowing tasks to converge naturally without interference, and the neural network processing unit returns to a clear state before the predicted load surge.

[0055] In the event of an extremely violent sudden incident, the warning time margin is... Further compressed to equal to or less than hour, Less than or equal to 0, after Cut off, at this point the maximum safe borrowing limit. It smoothly and precisely returns to 0. At the same time, this point in time accurately meets the conditions for mandatory triggering of subsequent soft reclamation.

[0056] In this way, by embedding the real-time comparison of the warning time margin with the maximum remaining completion time of the borrowing task, the preset recovery time overhead, and the smooth buffer time into the dynamic tightening logic of the borrowing quota, the scheduling system can complete the recovery of computing power in a lossless manner by naturally completing the borrowing task in most session burst scenarios, thus avoiding hard interruption paths.

[0057] S4. Trigger the soft recycling process and release computing power.

[0058] In an optional embodiment, a soft reclamation process is triggered to migrate the intermediate state of the borrowed task to the central processing unit for degraded execution in order to free up the computing power of the neural network processing unit.

[0059] If the warning time margin Less than or equal to the maximum remaining completion time Compared to the fixed overhead time required for the recycling operation itself The sum of When the system is established, there may be situations where the borrowing task cannot be completed naturally within the warning time margin. In this case, a soft recycling process is triggered, which migrates the intermediate state of the borrowing task to the central processing unit for downgraded execution to release the computing power of the neural network processing unit.

[0060] For tasks that utilize the computing power of the neural network processing unit (NNPU), they are organized into a directed execution graph at the operator level upon task submission. After the currently executing operator layer completes, task execution on the NNPU side is paused. The output tensor and execution progress pointer of the current layer are written to the shared memory mapping region to form a checkpoint snapshot. Cross-processor state persistence is achieved through address mapping from the NNPU's video memory in the PCIe-BAR space to the CPU's physical memory. The checkpoint snapshot is read into the CPU's memory, and the CPU continues to complete the computation of the remaining operator layers in degraded mode, ultimately releasing the NNPU's computing power. Because the soft reclamation trigger condition itself minimizes the remaining task load through secure borrowing quota compression logic, the number of remaining operator layers from the checkpoint position to task completion has been compressed, making the additional impact on the final response latency of vertical application services during CPU degrade execution within a controllable range. After the CPU degrade execution is completed, the status flag of the corresponding task is updated. Simultaneously, the NNPU's computing power is completely released and returned to the inference domain after the soft reclamation pause.

[0061] During the process of writing the output tensor of the current layer and the execution progress pointer to the shared memory mapping region to form a checkpoint snapshot, the load of the neural network processing unit during the checkpoint snapshot writing process is monitored; if the load of the neural network processing unit triggers the emergency threshold, a hard interrupt fallback path is executed; the current borrowed task is discarded, and a task failure notification is sent to the application service side to resubmit it to the central processing unit queue for execution, ultimately forming the abnormal handling result of the checkpoint snapshot.

[0062] After soft reclamation is completed and the computing power of the neural network processing unit is released, the predicted load is continuously updated and corrected. When predicting future warning time margin Internal Corrected Predicted Load Continuously above the recycling trigger threshold At that time, the maximum safe borrowing limit The value will remain at 0, and no new borrowing requests will be approved; when the projected load is revised... The risk level drops below the recycling trigger threshold and the session burst risk factor decreases. When the arrival rate falls below the historical average, the borrowing pool is opened and gradually opened according to the preset safe borrowing limit relationship, and finally the dynamic allocation of heterogeneous computing resources is restored, avoiding large-scale fluctuations in resource allocation and enabling the borrowing pool to be restored to an available state as soon as possible within a safe range.

[0063] Reference Figure 2This paper demonstrates the average response latency of inference tasks at five typical time points using the existing method with hard interrupts and the average response latency of the present invention's method with soft reclamation at five typical time points. Over time, the increase in the average response latency of the inference task using the existing method with hard interrupts is greater than that of the present invention's method with soft reclamation. The gap between the existing method with hard interrupts and the present invention's method with soft reclamation gradually widens. This proves the technical effect of the present invention's method with soft reclamation, which exhibits more stable latency control characteristics under sustained high load scenarios. Because the checkpoint-driven soft reclamation process completely preserves the intermediate execution state of the borrowed task and seamlessly migrates it to the central processing unit for continued execution when the warning time margin is insufficient to support the natural completion of the borrowed task, it achieves lossless and rapid release of the computing power of the neural network processing unit. This avoids the task state corruption caused by the hard interrupt method and the inference queue backlog caused by waiting for the task to complete naturally, thus verifying the rationality of the aforementioned latency control characteristics.

[0064] Thus, through the checkpoint-driven soft recycling process, when the warning time margin is insufficient to support the natural completion of the borrowing task, the intermediate execution state of the borrowing task is completely preserved and seamlessly transferred to the central processing unit for continued execution. This achieves lossless and rapid release of the computing power of the neural network processing unit, avoids the risk of task state damage caused by hard interruption, and at the same time, the emergency fallback path covers extreme working conditions, ensuring the fault tolerance capability of the soft recycling process under abnormal conditions.

[0065] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for dynamically allocating heterogeneous computing resources for vertical computing scenarios, applied to a heterogeneous computing architecture including a central processing unit (CPU) and a neural network processing unit (NN unit), and allowing tasks carried by the CPU to borrow computing power from the NN unit for execution, characterized in that, include: S1. Extract session-layer request features including the number of new sessions and the coefficient of variation of request arrival intervals, and construct a session burst risk coefficient by combining it with the historical session arrival rate average. This includes: collecting the number of new sessions within the current active session window and calculating the request arrival interval sequence for each session within the time window; calculating the ratio of the standard deviation to the mean of the request arrival interval sequence to obtain the coefficient of variation of request arrival intervals; multiplying the ratio of the number of new sessions to the historical session arrival rate average by a multiplier, which is the sum of the coefficient of variation of request arrival intervals and 1. S2. Obtain the original prediction sequence output by the time series prediction module, and perform a correction on the original prediction sequence based on the session burst risk coefficient, which linearly decays with the increase of the prediction step size, to obtain the corrected prediction load. This includes: determining whether the session burst risk coefficient is greater than the historical average benchmark; if it is greater than the historical average benchmark, the correction is performed on the original prediction sequence. S3. Based on the mean benchmark, calculate the difference between the session burst risk coefficient and the historical mean benchmark, and multiply the difference, risk amplification coefficient and linear decay factor to obtain the correction term; use 1 and the sum of the correction term as the multiplication factor, and multiply the original prediction sequence with the multiplication factor to obtain the corrected prediction load; S4. Estimate the warning time margin based on the corrected prediction load, and dynamically determine the safe borrowing limit based on the maximum remaining completion time of the borrowing task, the preset recovery time overhead and the smoothing buffer time, and restrict new borrowing task applications based on the safe borrowing limit; S5. If the warning time margin is less than or equal to the sum of the maximum remaining completion time and the recovery time overhead, trigger the soft recovery process, migrate the intermediate state of the borrowing task to the central processing unit for downgraded execution to release the computing power of the neural network processing unit.

2. The method for dynamic allocation of heterogeneous computing resources for vertical scenarios according to claim 1, characterized in that, The method for determining the length of the time window includes: continuously recording a full sample of the interval between adjacent requests arriving in a historical session; updating the full sample on a daily basis and extracting percentile values; and multiplying the percentile values ​​by a preset multiple to obtain the length of the time window.

3. The method for dynamic allocation of heterogeneous computing resources for vertical scenarios according to claim 1, characterized in that, The method for obtaining the risk amplification coefficient includes: obtaining a sequence of session sudden risk coefficients before the occurrence of historical sudden events as independent variables; taking the difference between the actual observed load peak and the original predicted sequence as dependent variables; and obtaining the risk amplification coefficient by solving the least squares linear regression.

4. The method for dynamic allocation of heterogeneous computing resources for vertical scenarios according to claim 1, characterized in that, The step of estimating the early warning time margin based on the modified predicted load includes: statistically analyzing the load level value sequence when request queuing backlog occurs in historical operating data to obtain the mean and standard deviation; subtracting the standard deviation from the mean to obtain the recovery trigger threshold; and calculating the time length from the current moment until the modified predicted load first exceeds the recovery trigger threshold to obtain the early warning time margin.

5. The method for dynamic allocation of heterogeneous computing resources for vertical scenarios according to claim 1, characterized in that, The method of dynamically determining the upper limit of the safe borrowing limit by combining the maximum remaining completion time of the borrowing task, the preset recovery time cost, and the smoothing buffer time includes: for each task currently running on the computing power of the borrowing neural network processing unit, tracking the current operator layer position of the borrowing task in the execution graph; querying the pre-stored operator layer time consumption file, accumulating the historical average execution time of each operator layer to obtain the remaining completion time, comparing the remaining completion time of all current borrowing tasks, and extracting the maximum value as the maximum remaining completion time; calculating the difference between the warning time margin and the sum of the maximum remaining completion time and the recovery time cost, and dividing the difference by the preset smoothing buffer time to obtain a ratio value; comparing the ratio value with 1 and taking the smaller value, and then comparing it with 0 and taking the larger value as the truncation coefficient, and multiplying the truncation coefficient by the upper limit of the total computing power of the borrowing pool to obtain the upper limit of the safe borrowing limit.

6. The method for dynamic allocation of heterogeneous computing resources for vertical scenarios according to claim 1, characterized in that, The triggering of the soft recycling process, which migrates the intermediate state of the borrowed task to the central processing unit for degraded execution to release the computing power of the neural network processing unit, includes: pausing task execution on the neural network processing unit side after the currently executing operator layer has finished running; writing the output tensor and execution progress pointer of the current layer into the shared memory mapping region to form a checkpoint snapshot; reading the checkpoint snapshot into the central processing unit memory, and having the central processing unit continue to complete the calculation of the remaining operator layers in degraded mode to release the computing power of the neural network processing unit.

7. The method for dynamic allocation of heterogeneous computing resources for vertical scenarios according to claim 6, characterized in that, The step of writing the output tensor and execution progress pointer of the current layer into the shared memory mapping region to form a checkpoint snapshot includes: monitoring the load of the neural network processing unit during the checkpoint snapshot writing process; if the load of the neural network processing unit triggers an emergency threshold, then executing a hard interrupt fallback path; discarding the current borrowed task and sending a task failure notification to the application service side to resubmit it to the central processing unit queue for execution, ultimately forming the abnormal handling result of the checkpoint snapshot.

8. The method for dynamic allocation of heterogeneous computing resources for vertical scenarios according to claim 4, characterized in that, After releasing the computing power of the neural network processing unit, the method further includes: continuously updating the corrected prediction load; when the corrected prediction load within the future warning time margin is continuously higher than the recycling trigger threshold, the upper limit of the secure borrowing quota is maintained at 0, and no new borrowing applications are approved; when the corrected prediction load drops to a safe range below the recycling trigger threshold and the session burst risk coefficient falls below the average historical session arrival rate, the borrowing pool is gradually opened, and the borrowing is gradually opened according to the preset secure borrowing quota relationship, restoring the dynamic allocation of heterogeneous computing resources.

Citation Information

Patent Citations

  • Heterogeneous computing power pool task adaptive migration and dynamic scheduling method and system

    CN121858288A

  • Container unified arrangement and non-inductive migration method and system for heterogeneous computing power pool

    CN122173201A