CPU (Central Processing Unit), GPU (Graphic Processing Unit) and NPU (Network Processing Unit) resource allocation method and system for training and calculating integrated machine

By integrating heterogeneous hardware parameters and real-time indicators, dynamically correcting the CPU, GPU, and NPU weight values, combined with task-led mode and hierarchical pressure response, the resource contention and efficiency bottlenecks in heterogeneous resource management are solved, and efficient resource utilization and stability guarantees are achieved in hybrid task scenarios.

CN120234155AInactive Publication Date: 2025-07-01DONGGUAN HUAMING TENG TECH CO LTD

Patent Information

Application Number
CN202510706424.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the existing technology deploys training and inference tasks in hybrid deployment, it is difficult to dynamically balance heterogeneous computing resources, resulting in resource competition and efficiency bottlenecks, which cannot meet the long-term stability of training tasks and the low-latency and high concurrency requirements of inference tasks.

Method used

By integrating the benchmark parameters of heterogeneous hardware and real-time dynamic indicators, the available weight values ​​of CPU, GPU, and NPU are dynamically corrected, combined with GPU stream processor utilization and task batch processing requirements, deep perception and efficient quantification of computing power resources are achieved, and dominant mode discrimination and hierarchical pressure response strategies are adopted to perform differentiated resource allocation.

Benefits of technology

Ensure the stability of computing power output in high temperature and high concurrency environments, realize accurate resource adaptation of training and reasoning tasks, and improve resource utilization and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234155A_ABST
    Figure CN120234155A_ABST
Patent Text Reader

Abstract

The invention relates to the field of heterogeneous computing resource management, in particular to a CPU (Central Processing Unit), GPU (Graphic Processing Unit) and NPU (Network Processing Unit) resource allocation method and system of a training and reckoning all-in-one machine. The invention discloses a CPU, GPU and NPU resource allocation system of a training and reckoning all-in-one machine. The system comprises a schedulable resource analysis module, a load analysis module and a resource allocation module. According to the method, the reference parameters and the real-time dynamic indexes of heterogeneous hardware are fused, so that deep perception and efficient quantification of computing power resources are realized; on the basis of an NPU temperature attenuation experiment, calibrating a computing power loss coefficient, and dynamically correcting available weight values of a CPU, a GPU and an NPU in combination with suitability analysis of a GPU stream processor utilization rate and a task batch processing demand and collaborative efficiency evaluation of a CPU dominant frequency and an IPC value; and the accuracy of resource availability prediction is improved, so that the system can still guarantee the stability of computing power output in complex environments such as high temperature and high concurrency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of heterogeneous computing resource management, and specifically to a method and system for allocating CPU, GPU, and NPU resources of a training and inference computing integrated machine. Background Art

[0002] With the rapid development of artificial intelligence technology, the heterogeneous computing architecture of training and inference integration has gradually become the mainstream form of artificial intelligence infrastructure; the collaborative scheduling of heterogeneous computing units such as CPUs, GPUs, and NPUs faces complex challenges: training tasks have high requirements for the long-term stability of computing resources, while inference tasks need to meet the real-time requirements of low latency and high concurrency; under the limited hardware resource configuration, how to dynamically balance the load of heterogeneous computing power, optimize resource utilization, and meet service quality has become a key issue restricting system efficiency.

[0003] Traditional resource allocation schemes are mostly based on static rules or single metrics, lacking comprehensive evaluation of multi-dimensional dynamic performance parameters; for example, existing methods usually ignore the impact of the temperature characteristics of NPUs on computing efficiency. In a high-temperature environment, the actual computing power of NPUs will be significantly attenuated due to chip downclocking, but this variable is not effectively quantified in resource scheduling; in addition, when training and inference tasks are deployed in a hybrid manner, most systems adopt a fixed priority strategy, which often causes resource competition: when training tasks occupy large resources for a long time, the latency of inference tasks may exceed the allowable range of the service level agreement; conversely, if excessive preference is given to inference tasks, the overall throughput of training jobs will be sacrificed; more seriously, conventional schemes lack sensitivity to the performance degradation of heterogeneous hardware. For example, when the CPU main frequency drops, the GPU stream processor utilization rate is insufficient, or the NPU cache hit rate decreases, the existing scheduling algorithms are difficult to perceive and trigger resource reallocation in a timely manner.

[0004] In view of the above problems, some dynamic scheduling methods have been proposed in the prior art, such as resource reservation strategies based on historical load prediction, or hybrid schemes that allocate computing power through weighted averaging; however, these methods have significant limitations: First, the performance characteristics differences between training tasks and inference tasks are not fully modeled. The sensitivities of the two to parameters such as batch size (BatchSize) and query rate (QPS) are different, and unified scheduling rules often lead to resource mismatches; Second, the priority decision-making mechanism is too simplistic, lacking a multi-dimensional scoring system (such as task type, submission source, business importance), making it difficult to accurately identify the core links that need to be prioritized in high-load scenarios; Third, environmental factors (such as NPU temperature fluctuations) and hardware real-time status (such as dynamic frequency, queue depth) are not incorporated into the resource allocation model, resulting in the inability to fully release the underlying computing power potential; Therefore, the present invention proposes a method and system for allocating CPU, GPU, and NPU resources of a training and inference computing integrated machine, which solves the resource contention and efficiency bottleneck in the hybrid training and inference scenario, realizes the global optimization of heterogeneous resource utilization rate, and at the same time ensures the low-latency and high-reliability requirements of critical services. Summary of the Invention

[0005] The present invention realizes in-depth perception and efficient quantification of computing power resources by integrating the benchmark parameters and real-time dynamic indicators of heterogeneous hardware; calibrates the computing power loss coefficient based on the NPU temperature attenuation experiment, combines the adaptability analysis of the GPU stream processor utilization rate and the task batch processing requirements, and the collaborative efficiency evaluation of the CPU main frequency and IPC value, and dynamically corrects the available weight values of the CPU, GPU, and NPU; improves the accuracy of resource availability prediction, enabling the system to ensure the stability of computing power output in complex environments such as high temperature and high concurrency.

[0006] A method for allocating CPU, GPU, and NPU resources of a training and inference computing integrated machine, comprising:

[0007] Obtain the CPU basic parameters, GPU basic parameters, and NPU basic parameters; the CPU basic parameters include the number of cores, the base main frequency, and the base IPC value; the GPU basic parameters include the number of stream processors and the FP32 theoretical computing power; the NPU basic parameters include the number of MAC units and the INT8 theoretical computing power;

[0008] Collect the CPU dynamic indicators, GPU dynamic indicators, and NPU dynamic indicators at the current performance analysis time point; the CPU dynamic indicators include the actual main frequency and the actual IPC value; the GPU dynamic indicators include the FP32 actual computing power and the video memory utilization rate; the NPU dynamic indicators include the temperature, cache hit rate, and INT8 actual computing power; then calculate the weights of the CPU, GPU, and NPU respectively based on the CPU dynamic indicators, GPU dynamic indicators, and NPU dynamic indicators, and further calculate the current global schedulable resources;

[0009] Based on the current task list, obtain the task classification identifier and task performance parameters for each task; the task classification identifier includes a training task or an inference task; the task performance parameters for the training task correspond to the batch size and the current iteration number; the task performance parameters for the inference task correspond to the queries per second and the time delay; at the same time, query the stream processor utilization rate of the current GPU and the NPU calculation queue depth, and calculate the load pressure factors of the CPU, GPU, and NPU respectively; based on the time delays of all current inference tasks, determine whether the current is an inference task-dominated mode or a training task-dominated mode, and calculate the global load pressure value in the corresponding mode;

[0010] Based on the current dominant mode and the global load pressure value, execute a differential resource allocation strategy.

[0011] Preferably, calculate the weights of the CPU, GPU, and NPU respectively based on the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics, and then calculate the current global schedulable resources. The specific operations are as follows:

[0012] The CPU basic parameters include the number of cores 、the base main frequency and the base IPC value ; the GPU basic parameters include the number of stream processors and the FP32 theoretical computing power ; the NPU basic parameters include the number of MAC units and the INT8 theoretical computing power ;

[0013] The CPU dynamic metrics include the actual main frequency and the actual IPC value ; the GPU dynamic metrics include the FP32 actual computing power and the video memory utilization rate ; the NPU dynamic metrics include the temperature 、the cache hit rate and the INT8 actual computing power ;

[0014] Based on the currently obtained CPU dynamic metrics, use the formula to calculate and obtain the CPU weight ; where is the preset main frequency weight, is the preset IPC weight;

[0015] Based on the currently obtained GPU dynamic metrics, simultaneously obtain the current task batch size and the maximum effective processing volume of the GPU , and use the formula to calculate and obtain the GPU weight ;

[0016] Based on the currently obtained NPU dynamic metrics, use the formula to calculate and obtain the NPU weight ; where is the NPU reference temperature, i.e., the optimal operating temperature; is the cache efficiency threshold, and when the NPU cache hit rate is lower than this value, a performance penalty is triggered; is the preset temperature decay coefficient;

[0017] Finally, use the formula to calculate and obtain the current global schedulable resources .

[0018] Preferably, calculate the load pressure factors of the CPU, GPU, and NPU respectively. The specific operations are as follows:

[0019] Obtain the current utilization rate of the GPU's stream processors and the depth of the NPU calculation queue ;

[0020] Use the formula to calculate and obtain the CPU load pressure factor;

[0021] Use the formula to calculate and obtain the GPU load pressure factor;

[0022] Use the formula to calculate and obtain the NPU load pressure factor; is the queue load threshold, indicating that when the queue depth exceeds this critical point, the processing delay increases significantly or the chip efficiency drops sharply; this formula indicates that when the queue depth exceeds or the cache misses, the NPU pressure increases.

[0023] Preferably, based on the time delays of all current inference tasks, determine whether the current is an inference task-dominated mode or a training task-dominated mode. The specific operations are as follows:

[0024] Set the maximum allowable time delay . If there is any time delay for the current inference tasks, then it is determined that the current is an inference task-dominated mode; otherwise, it is a training task-dominated mode.

[0025] Preferably, calculate the global load pressure value in the corresponding mode. The specific operations are as follows:

[0026] Use the formula to calculate and obtain the global load pressure value .

[0027] Preferably, based on the current dominant mode and the global load pressure value, a differential resource allocation strategy is executed, and the specific operations are as follows:

[0028] If the current global load pressure value is less than 0.5 and the current dominant mode is the training task mode, then use the formula to calculate and obtain the resource release amount in the training mode , and recycle SCU units from the training task and release them to the global resource pool; where is the total number of SCUs currently allocated to the training task;

[0029] If the current global load pressure value is less than 0.5 and the current dominant mode is the inference task mode, then use the formula to calculate and obtain the resource allocation amount in the inference mode , and reserve SCU units;

[0030] If and the current dominant mode is the training task mode, then use the formula to calculate and obtain , represents the adjusted Batch size;

[0031] If and the current dominant mode is the inference task mode, then use the formula to calculate and obtain the adjusted maximum request processing rate ; is the theoretical maximum QPS; is the maximum queue capacity;

[0032] If , then for any task currently in progress, use the formula to calculate and obtain the termination determination value of the current task, where is the priority of the current task, is the maximum allowed time delay. If the termination determination value of the current task, then terminate the current task and recycle the resources occupied by the task to the global resource pool.

[0033] A CPU, GPU, and NPU resource allocation system for a training and inference computing integrated machine, including:

[0034] The schedulable resource analysis module includes a basic parameter acquisition unit, a dynamic index collection unit, and a resource calculation unit. The basic parameter acquisition unit is used to acquire CPU basic parameters, GPU basic parameters, and NPU basic parameters. The dynamic index collection unit is used to collect CPU dynamic indexes, GPU dynamic indexes, and NPU dynamic indexes at the current performance analysis time point. The resource calculation unit is used to calculate the weights of the CPU, GPU, and NPU respectively using the CPU dynamic indexes, GPU dynamic indexes, and NPU dynamic indexes, and then calculate the current global schedulable resources.

[0035] The load analysis module includes a task query unit, a stress factor calculation unit, a mode discrimination unit, and a global load calculation unit. The task query unit is used to obtain the task classification identifier and task performance parameters of each task based on the current task list. The stress factor calculation unit is used to query the stream processor utilization rate of the current GPU and the NPU calculation queue depth, and calculate the load stress factors of the CPU, GPU, and NPU respectively. The mode discrimination unit is used to determine whether the current is an inference task-dominated mode or a training task-dominated mode based on the time delay of all current inference tasks. The global load calculation unit is used to calculate the global load stress value in the corresponding mode.

[0036] The resource allocation module is used to execute a differential resource allocation strategy through the current dominant mode and the global load stress value.

[0037] The present invention has the following advantages:

[0038] 1. By integrating the benchmark parameters and real-time dynamic indexes of heterogeneous hardware, the present invention realizes in-depth perception and efficient quantification of computing power resources. Based on the NPU temperature attenuation experiment to calibrate the computing power loss coefficient, combined with the adaptability analysis of the GPU stream processor utilization rate and the task batch processing requirements, and the collaborative efficiency evaluation of the CPU main frequency and IPC value, the available weight values of the CPU, GPU, and NPU are dynamically corrected. The accuracy of resource availability prediction is improved, so that the system can still ensure the stability of computing power output in complex environments such as high temperature and high concurrency.

[0039] 2. The present invention innovatively proposes a dynamic discrimination and hierarchical pressure response strategy for the "dominant mode", achieving precise resource adaptation for training and inference tasks. In high-load scenarios, the global pressure value is calculated in real time through the load pressure factor, and combined with the task priority score, differential scheduling is performed: for the inference task dominant mode, the latency requirements of high-priority requests are preferentially guaranteed, and by dynamically restricting the QPS upper limit of batch inference, low-timeliness tasks are prevented from squeezing core business resources; in the training task dominant mode, the batch processing scale is flexibly adjusted to ensure the balance between model iteration efficiency and resource occupancy. More importantly, when the load pressure exceeds the threshold, the system automatically triggers the task termination scoring algorithm to forcibly recycle inefficient task resources to preferentially meet the needs of high-value services, solving the problem that it is difficult to achieve both service quality and resource utilization rate in mixed-load scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 FIG. is a schematic structural diagram of the CPU, GPU, and NPU resource allocation system of the training and inference computing integrated machine adopted in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0042] Embodiment 1. A method for allocating CPU, GPU, and NPU resources of a training and inference computing integrated machine includes:

[0043] Obtain the CPU basic parameters, GPU basic parameters, and NPU basic parameters; the CPU basic parameters include the number of cores, the base main frequency, and the base IPC value; the GPU basic parameters include the number of stream processors and the FP32 theoretical computing power; the NPU basic parameters include the number of MAC units and the INT8 theoretical computing power; the number of cores is the physical basis for the parallel processing ability of the CPU and determines the number of threads that can be executed simultaneously; the main frequency reflects the single-core computing speed and, together with IPC (instructions per clock cycle), determines the single-thread peak performance; the number of stream processors is the smallest unit for the GPU to execute calculations, and the number directly affects the parallel computing ability; FP32 (32-bit single-precision floating-point number) is a high-precision numerical format, and FP32 performance = number of SMs × number of floating-point units per single SM × main frequency; this indicator directly affects the resource allocation weight of tasks that require high-precision calculations (such as parameter updates during large model training); MAC determines the parallel ability of integer / low-precision matrix multiplication and addition; INT8 (8-bit signed integer) is a low-precision fixed-point number format, and INT8 performance = number of MAC units × frequency × 2 (INT8 can execute two operations per cycle); this indicator directly affects the throughput of quantized inference tasks.

[0044] Collect CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics at the current performance analysis time point; CPU dynamic metrics include the actual main frequency and the actual IPC value; GPU dynamic metrics include the actual FP32 computing power and the video memory utilization rate; NPU dynamic metrics include temperature, cache hit rate, and actual INT8 computing power; subsequently, calculate the weights of the CPU, GPU, and NPU based on the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics respectively, and then calculate the current global schedulable resources; the cache hit rate represents the success rate of accessing the internal cache of the NPU, and a low hit rate means frequent access to the external DDR, increasing latency and power consumption;

[0045] Based on the current task list, obtain the task classification identifier and task performance parameters of each task; the task classification identifier includes training tasks or inference tasks. Training tasks focus on long-term occupation of computing resources (such as GPUs) and high video memory requirements, allowing a certain delay; Inference tasks: have high requirements for real-time performance (low latency), are usually deployed as online services, and resource allocation needs to meet the SLA; the task performance parameters of training tasks correspond to the batch size and the current iteration number; the task performance parameters of inference tasks correspond to the queries per second and the time delay; at the same time, query the utilization rate of the stream processors of the current GPU and the depth of the NPU calculation queue, and calculate the load pressure factors of the CPU, GPU, and NPU respectively; based on the time delays of all current inference tasks, determine whether the current is the inference task dominant mode or the training task dominant mode, and calculate the global load pressure value in the corresponding mode;

[0046] Based on the current dominant mode and the global load pressure value, execute a differential resource allocation strategy; in the inference or training dominant mode, the system dynamically allocates resources according to the global load pressure: in the inference mode, tasks are preferentially scheduled to high-computing-power hardware (such as NPU / GPU) to meet the low-latency requirements, and excess load is shunted based on the queue depth and temperature; in the training mode, the GPU resources are locked to ensure the computing efficiency of large batches, and at the same time, the CPU is used to assist in preprocessing or the NPU cache is used to optimize the data stream; the global load pressure value drives elastic scaling, such as restricting resource preemption under high load and allowing oversubscription under low load, thereby improving resource utilization, ensuring the performance of high-priority tasks, and preventing hardware overload or overheating.

[0047] Calculate the weights of the CPU, GPU, and NPU based on the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics respectively, and then calculate the current global schedulable resources. The specific operations are as follows:

[0048] The CPU basic parameters include the number of cores and the base main frequency and the base IPC value ; the GPU basic parameters include the number of stream processors and the theoretical FP32 computing power ; The basic parameters of the NPU include the number of MAC units and the theoretical INT8 computing power ;

[0049] The dynamic indicators of the CPU include the actual main frequency and the actual IPC value ; The dynamic indicators of the GPU include the actual FP32 computing power and the video memory utilization rate ; The dynamic indicators of the NPU include temperature , cache hit rate and the actual INT8 computing power ;

[0050] Based on the currently obtained CPU dynamic indicators, use the formula to calculate and obtain the CPU weight ; Among them, is the preset main frequency weight, is the preset IPC weight, and The default values of are 0.6 and 0.4 respectively, which can be adjusted by professionals according to experience;

[0051] Based on the currently obtained GPU dynamic indicators, simultaneously obtain the current task batch size and the maximum effective processing capacity of the GPU , use the formula to calculate and obtain the GPU weight ;

[0052] Based on the currently obtained NPU dynamic indicators, use the formula to calculate and obtain the NPU weight ; Among them, is the NPU reference temperature, that is, the optimal operating temperature recommended by the manufacturer; is the cache efficiency threshold, and performance penalty is triggered when the NPU cache hit rate is lower than this value; is the temperature attenuation coefficient;

[0053] Temperature attenuation coefficient The calibration steps are as follows:

[0054] Step 1: Run the NPU in the temperature control box with a fixed load; adjust the ambient temperature, for example, gradually increase from 40°C to 80°C, and record the change of the actual INT8 computing power with temperature;

[0055] Step 2: Assume that the actual INT8 computing power satisfies the exponential decay model:

[0056]

[0057] Perform linear regression on the logarithmically transformed data to obtain the slope, which is ;

[0058] Finally, use the formula to calculate and obtain the current global schedulable resources ; In this invention, this part is based on the dynamic performance indicators of heterogeneous hardware (CPU / GPU / NPU), quantifies the real-time efficiency weights of each hardware through mathematical modeling, and finally calculates the global schedulable resources; the CPU weight reflects the performance attenuation of its actual main frequency and IPC (instruction efficiency) relative to the benchmark. A decrease in the main frequency or a decrease in IPC will reduce the value; the GPU weight evaluates the availability by combining the actual computing power and batch processing efficiency (to avoid reducing the utilization rate due to small batches); the NPU weight introduces a temperature penalty term (exponential decay when the temperature is too high) and cache hit rate (penalty when the hit rate is low) to comprehensively calculate the efficiency; finally, the global resource quantity is obtained by multiplying the weights of each hardware by its core configuration (such as the number of CPU cores and the number of GPU stream processors), dynamically measuring the real-time computing power capacity of the heterogeneous cluster, and providing a quantitative basis for task allocation; its principle is to fuse and model through multi-dimensional dynamic parameters (performance, temperature, task efficiency), normalize the capabilities of heterogeneous hardware into a unified schedulable resource unit, and achieve precise resource matching and load balancing.

[0059] Calculate the load pressure factors of the CPU, GPU, and NPU respectively. The specific operations are as follows:

[0060] Obtain the stream processor utilization rate of the current GPU and the NPU calculation queue depth ;

[0061] Use the formula to calculate and obtain the CPU load pressure factor; a decrease in the main frequency or a decrease in IPC efficiency will increase the pressure factor, reflecting the deterioration of CPU performance;

[0062] Use the formula to calculate and obtain the GPU load pressure factor; an overload of either the video memory or the computing unit will cause the load pressure to rise;

[0063] Use the formula to calculate and obtain the NPU load pressure factor; is the queue load threshold, indicating that when the NPU is in the recommended working mode, when the queue depth reaches this value, it is considered to enter the performance-sensitive area (that is, after the queue depth reaches this critical point, the processing delay increases significantly or the chip efficiency drops suddenly). This parameter should be determined based on the characteristics of the NPU hardware architecture, usually provided by the manufacturer in the technical documentation or calibrated through expert experience; this formula indicates that when the queue depth exceeds or the cache is not hit, the NPU pressure increases;

[0064] Construct a load pressure factor through the bottleneck indicators of heterogeneous hardware cores to specifically evaluate the performance bottlenecks of each device: The CPU load factor is based on the main frequency drop and the decay of instruction efficiency (the degree of deviation between the actual value and the benchmark value), reflecting the decline of core computing efficiency; the GPU pressure takes the square of the maximum value of its video memory and stream processor utilization rate, highlighting the non-linear impact of any resource overload on the overall performance; the NPU pressure is triggered by the queue depth and cache misses. The more critical the queue or the lower the cache hit rate, the higher the pressure index, reflecting the combined impact of task accumulation and the decline in data transfer efficiency; its principle is to perform index fusion modeling around the core bottlenecks with different hardware characteristics, quantify the resource competition pressure within a single device, provide a fine-grained load perception basis for the core scheduler, and guide the hierarchical flow control or elastic migration strategy.

[0065] Based on the time delay of all current inference tasks, determine whether the current is the inference task dominant mode or the training task dominant mode. The specific operations are as follows:

[0066] Set the maximum allowable time delay , if there is any time delay of an inference task currently , then it is determined that the current is the inference task dominant mode, otherwise it is the training task dominant mode; Dynamically determine the system dominant mode by monitoring the time delay of inference tasks: If the current delay of any inference task exceeds 80% of the preset maximum delay, then trigger the inference dominant mode, otherwise it is the training dominant mode; Its core principle is to give priority to ensuring the quality of service of inference tasks that are sensitive to real-time. When the delay approaches the risk threshold, the system switches to the inference mode to optimize resource allocation and avoid exceeding the delay standard; On the contrary, within the safe range of delay, maintain the training mode and allow training tasks to fully occupy resources; Dynamically trigger mode switching through delay to achieve an elastic balance between low-latency requirements (inference) and high-throughput demands (training).

[0067] Calculate the global load pressure value in the corresponding mode. The specific operations are as follows:

[0068] Use the formula to calculate and obtain the global load pressure value .

[0069] Based on the current dominant mode and the global load pressure value, execute a differentiated resource allocation strategy. The specific operations are as follows:

[0070] If the current global load pressure value is less than 0.5 and the current is the training task dominant mode, then use the formula to calculate and obtain the resource release amount in the training mode , and recycle SCU units from the training tasks and release them to the global resource pool; Among them, is the total number of SCUs currently allocated to the training task;

[0071] If the current global load pressure value is less than 0.5 and the current is the inference task dominant mode, then use the formula to calculate and obtain the resource allocation amount in the inference mode , and reserve SCU units, and only allow tasks with a priority greater than 7 points to occupy;

[0072] If and the current is the training task dominant mode, then use the formula to calculate and obtain , represents the adjusted Batch size;

[0073] If and the current is the inference task dominant mode, then use the formula to calculate and obtain the adjusted maximum request processing rate ; is the theoretical maximum QPS; is the maximum queue capacity;

[0074] If , then for any task currently in progress, use the formula to calculate and obtain the termination determination value of the current task , where is the priority of the current task, is the maximum allowed time delay. If the termination determination value of the current task , then terminate the current task and recycle the resources occupied by the task to the global resource pool; if the current is the training task dominant mode, then use the recycled resources for the training task; if the current is the inference task dominant mode, then allocate the recycled resources to high-priority inference tasks;

[0075] The determination method of task priority is as follows:

[0076] Priority score = task baseline score + submission level score + business addition score; the task baseline scores of online inference tasks, batch inference tasks, and training tasks are 5 points, 3 points, and 1 point in sequence; the submission level score is determined based on the submitter of the task, and the submission level scores corresponding to the user terminal, business background, and administrator are 1 point, 2 points, and 3 points in sequence; the business addition scores of ordinary business, regulatory business, and core link are 1 point, 2 points, and 3 points in sequence; if the calculation result of the priority score is greater than 10 points, then set it to 10;

[0077] Example 2, a CPU, GPU, and NPU resource allocation system for a training and inference computing all-in-one machine, as Figure 1As shown in the figure, it includes:

[0078] A schedulable resource analysis module, including a basic parameter acquisition unit, a dynamic index collection unit, and a resource calculation unit; the basic parameter acquisition unit is used to acquire CPU basic parameters, GPU basic parameters, and NPU basic parameters; the CPU basic parameters include the number of cores, the base main frequency, and the base IPC value; the GPU basic parameters include the number of stream processors and the FP32 theoretical computing power; the NPU basic parameters include the number of MAC units and the INT8 theoretical computing power; the dynamic index collection unit is used to collect CPU dynamic indexes, GPU dynamic indexes, and NPU dynamic indexes at the current performance analysis time point; the CPU dynamic indexes include the actual main frequency and the actual IPC value; the GPU dynamic indexes include the FP32 actual computing power and the video memory utilization rate; the NPU dynamic indexes include temperature T, cache hit rate, and INT8 actual computing power; the resource calculation unit is used to calculate the weights of the CPU, GPU, and NPU respectively using the CPU dynamic indexes, GPU dynamic indexes, and NPU dynamic indexes, and then calculate the current global schedulable resources;

[0079] A load analysis module, including a task query unit, a stress factor calculation unit, a mode discrimination unit, and a global load calculation unit; the task query unit is used to obtain the task classification identifier and task performance parameters of each task based on the current task list; the task classification identifier includes a training task or an inference task; the task performance parameters of the training task correspond to the batch size and the current iteration number; the task performance parameters of the inference task correspond to the queries per second and the time delay; the stress factor calculation unit is used to query the stream processor utilization rate of the current GPU and the NPU calculation queue depth, and calculate the load stress factors of the CPU, GPU, and NPU respectively; the mode discrimination unit is used to determine whether the current is an inference task dominant mode or a training task dominant mode based on the time delay of all current inference tasks; the global load calculation unit is used to calculate the global load stress value in the corresponding mode;

[0080] A resource allocation module, which is used to execute a differential resource allocation strategy through the current dominant mode and the global load stress value.

[0081] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention. The parts not described in detail in this specification belong to the prior art well known to those skilled in the art.

Claims

1. A method for allocating CPU, GPU, and NPU resources of a training computing power integrated machine, characterized in that, Including: Obtain the basic parameters of the CPU, GPU, and NPU; The basic parameters of the CPU include the number of cores, the base main frequency, and the base IPC value; The basic parameters of the GPU include the number of stream processors and the FP32 theoretical computing power; The basic parameters of the NPU include the number of MAC units and the INT8 theoretical computing power; Collect the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics at the current performance analysis time point; the CPU dynamic metrics include the actual main frequency and the actual IPC value; the GPU dynamic metrics include the FP32 actual computing power and the video memory utilization rate; the NPU dynamic metrics include the temperature, the cache hit rate, and the INT8 actual computing power; then calculate the weights of the CPU, GPU, and NPU respectively based on the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics, and further calculate the current global schedulable resources; Based on the current task list, obtain the task classification identifier and task performance parameters of each task; the task classification identifier includes a training task or an inference task; The task performance parameters of the training task correspond to the batch size and the current iteration number; the task performance parameters of the inference task correspond to the queries per second and the time delay; at the same time, query the utilization rate of the stream processors of the current GPU and the depth of the NPU calculation queue, and calculate the load pressure factors of the CPU, GPU, and NPU respectively; based on the time delays of all current inference tasks, determine whether the current is an inference task dominant mode or a training task dominant mode, and calculate the global load pressure value in the corresponding mode; Execute a differential resource allocation strategy based on the current dominant mode and the global load pressure value.

2. The CPU, GPU, and NPU resource allocation method for a training and computing integrated machine according to claim 1, wherein Calculate the weights of the CPU, GPU, and NPU respectively based on the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics, and further calculate the current global schedulable resources. The specific operations are as follows: The basic parameters of the CPU include the number of cores , the base main frequency and the base IPC value ; The basic parameters of the GPU include the number of stream processors and the theoretical FP32 computing power ; The basic parameters of the NPU include the number of MAC units and the theoretical computing power of INT8 ; The dynamic indicators of the CPU include the actual main frequency and the actual IPC value ; The dynamic indicators of the GPU include the actual FP32 computing power and the video memory utilization rate ; The dynamic indicators of the NPU include the temperature , the cache hit rate and the actual INT8 computing power ; Based on the currently obtained CPU dynamic metrics, use the formula to calculate and obtain the CPU weight ; where is the preset main frequency weight, is the preset IPC weight; Based on the currently obtained GPU dynamic metrics, simultaneously obtain the current task batch size and the maximum effective processing capacity of the GPU , and use the formula to calculate and obtain the GPU weight ; Based on the currently obtained NPU dynamic metrics, use the formula to calculate and obtain the NPU weights ; where is the NPU reference temperature, i.e., the optimal operating temperature; is the cache efficiency threshold, and when the NPU cache hit rate is lower than this value, a performance penalty is triggered; is the preset temperature attenuation coefficient; Finally, use the formula to calculate and obtain the current globally schedulable resources .

3. A method for allocating CPU, GPU, and NPU resources of a training and computing integrated machine according to claim 2, wherein, Calculate the load pressure factors of the CPU, GPU, and NPU respectively. The specific operations are as follows: Obtain the stream processor utilization rate of the current GPU and the NPU calculation queue depth ; Use the formula to calculate and obtain the CPU load pressure factor; Use the formula to calculate and obtain the GPU load pressure factor; Use the formula to calculate and obtain the NPU load pressure factor; is the queue load threshold, indicating that when the queue depth exceeds this critical point, the processing delay will increase significantly or the chip efficiency will drop sharply; this formula means that when the queue depth exceeds or there is a cache miss, the NPU pressure increases.

4. A method for allocating CPU, GPU, and NPU resources of a training and computing integrated machine according to claim 3, characterized in that, Based on the time delays of all current inference tasks, determine whether the current is an inference task dominant mode or a training task dominant mode. The specific operations are as follows: Set the maximum allowable time delay If there is a time delay for any current inference task then it is determined that the current is the inference task dominant mode; otherwise, it is the training task dominant mode.

5. A method for allocating CPU, GPU, and NPU resources of a training and computing integrated machine according to claim 4, characterized in that, Calculate the global load pressure value in the corresponding mode. The specific operations are as follows: Use the formula to calculate and obtain the global load pressure value .

6. A method for allocating CPU, GPU, and NPU resources of a training and computing integrated machine according to claim 5, characterized in that, Execute a differential resource allocation strategy based on the current dominant mode and the global load pressure value. The specific operations are as follows: If the current global load pressure value is less than 0.5 and the current mode is dominated by training tasks, then use the formula to calculate the resource release amount in the training mode , and recycle SCU units from the training tasks and release them to the global resource pool; where is the total number of SCUs currently allocated to the training tasks; If the current global load pressure value is less than 0.5 and the current is the inference task dominant mode, then use the formula to calculate and obtain the resource allocation amount in the inference mode , and reserve SCU units; If and it is currently in the training task dominant mode, then use the formula to calculate and obtain , indicating the adjusted Batch size; If and the current is the inference task dominant mode, then use the formula to calculate and obtain the adjusted maximum request processing rate ; is the theoretical maximum QPS; is the maximum capacity of the queue; If , for any task currently in progress, use the formula to calculate and obtain the termination determination value of the current task, where is the priority of the current task, is the maximum allowable time delay. If the termination determination value of the current task, then terminate the current task and recycle the resources occupied by the task to the global resource pool.

7. A CPU, GPU, and NPU resource allocation system for a training and computing integrated machine, characterized in that, The system is applied to the CPU, GPU, and NPU resource allocation method of a training and inference computing integrated machine according to any one of claims 1-6 above, including: A schedulable resource analysis module, including a basic parameter acquisition unit, a dynamic metric acquisition unit, and a resource calculation unit; the basic parameter acquisition unit is used to obtain the basic parameters of the CPU, GPU, and NPU; the dynamic metric acquisition unit is used to collect the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics at the current performance analysis time point; the resource calculation unit is used to calculate the weights of the CPU, GPU, and NPU respectively using the CPU dynamic metrics, GPU dynamic metrics, and NPU dynamic metrics, and further calculate the current global schedulable resources; The load analysis module includes a task query unit, a stress factor calculation unit, a mode discrimination unit, and a global load calculation unit; the task query unit is used to obtain the task classification identifier and task performance parameters of each task based on the current task list; the stress factor calculation unit is used to query the stream processor utilization rate of the current GPU and the NPU calculation queue depth, and calculate the load stress factors of the CPU, GPU, and NPU respectively; the mode discrimination unit is used to determine whether the current is an inference task-dominated mode or a training task-dominated mode based on the time delay of all current inference tasks; the global load calculation unit is used to calculate the global load pressure value in the corresponding mode; The resource allocation module is used to execute a differential resource allocation strategy through the current dominant mode and the global load pressure value.

Citation Information

Patent Citations

  • Large model scheduling method and device based on NPU computing power

    CN119336457A

  • Artificial neural network module for performing artificial neural network operation on plurality of subgraphs and operating method thereof

    US20230105810A1

  • KR20240067788A

Cited By

  • Artificial intelligence reasoning service method and system of high-density interconnection structure

    CN121387559A

  • Method for computing power heterogeneous scheduling under large model training reasoning framework

    CN121411905A

  • A method for heterogeneous scheduling of computing power under a large model training inference framework

    CN121411905B

  • Automatic safety monitoring system for computing power server

    CN121434026A

  • Method and device for detecting cooperation efficiency between processors

    CN121597287A