Deep neural network inference method based on idle CPU resource for maximizing resource utilization rate in training cluster

By strategically allocating CPU cores for training and inference tasks, the method addresses CPU underutilization and interference, enhancing resource efficiency and throughput in GPU servers during deep neural network training.

WO2025143765A1PCT designated stage expired Publication Date: 2025-07-03RES & BUSINESS FOUND SUNGKYUNKWAN UNIV

Patent Information

Application Number
PCT/KR2024/021059
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing methods for utilizing CPU resources in GPU servers during deep neural network training fail to efficiently allocate CPU cores, leading to significant underutilization and interference with training tasks, especially when CPU-intensive inference workloads are involved.

Method used

A method that classifies CPU cores into groups for training tasks, designates unallocated cores as a 'U' group for online inference, and executes batch inference tasks on idle CPU cores, prioritizing training tasks when needed, while deploying multiple instances per thread to maximize throughput without degrading training performance.

Benefits of technology

Effectively utilizes idle CPU cycles for inference workloads, providing deterministic service without interfering with training, thereby maximizing resource utilization and throughput, especially for CPU-intensive tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024021059_03072025_PF_FP_ABST
    Figure KR2024021059_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A deep neural network inference method based on an idle CPU resource according to the present specification may comprise the steps of: when a training task is reserved and all GPUs are used, classifying a CPU core into a training task-specific group, and classifying an unallocated idle CPU core into an unallocated group (U group); executing a training task by using a CPU core of the training task-specific group and, when there is a request for an online inference task, executing the online inference task by using the idle CPU core of the U group; and when there is a request for a batch inference task, additionally executing the batch inference task by using at least one of the idle CPU core of the U group and an idle CPU core of the training task-specific group.
Need to check novelty before this filing date? Find Prior Art

Description

A deep neural network inference method based on idle CPU resources to maximize resource utilization in a training cluster.

[0001] This specification relates to a method for deep neural network inference based on idle CPU resources to maximize resource utilization in a training cluster.

[0002] A deep neural network (DNN) is an artificial neural network (ANN) consisting of multiple hidden layers between the input layer and the output layer.

[0003] DNNs, like general ANNs, can model complex non-linear relationships.

[0004] For example, in a DNN structure for an object identification model, each object can be represented as a hierarchical composition of image primitives.

[0005] At this time, additional layers can gradually combine the features of the collected lower layers.

[0006] This characteristic of DNN allows it to model complex data with a smaller number of units (nodes) compared to similarly performed ANNs.

[0007] Training DNN models is generally known to be complex and time-consuming, and more and more companies are building GPU clusters to address this issue.

[0008] In a GPU cluster, GPU nodes provide both GPU and CPU resources. Therefore, tasks can be categorized as either GPU or CPU tasks. GPU tasks require both GPU and CPU resources, while CPU tasks require only CPU resources.

[0009] In the field of training using DNNs, several approaches have been proposed to increase CPU utilization on GPU servers.

[0010] For example, there's a method for optimal CPU core allocation. This method distributes CPU cores based on actual training demand, reducing the occurrence of unused CPU cores.

[0011] However, a significant drawback of this approach is that it fails to account for the constraints imposed by the high CPU-to-GPU ratio on GPU servers. This imbalance often leaves CPU cores unallocated even when GPUs reach their maximum capacity. Consequently, even when the exact number of CPU cores required for the training task is allocated, an excess of unused CPU cores persists.

[0012] Another example is GPU sharing. This approach allows for easy distribution of a single GPU across multiple training jobs, increasing the number of training jobs and thus increasing the number of CPU cores utilized and overall efficiency.

[0013] However, this approach is limited by the specific framework in use and only partially addresses the issue of unused CPU cores on GPU servers due to significant CPU-to-GPU ratios.

[0014] Another example is nondeterministic behavior. Recent innovations have enabled CPU-intensive tasks to be executed on GPU servers. However, this approach lacks consistent results that can translate into beneficial results for inference workloads.

[0015] The inventors analyzed publicly available cloud logs and a series of experiments, and as shown in Fig. 1, it was found that during the training process of the GPU server, many CPU cores are often kept idle, wasting computational resources.

[0016] Figure 1 shows the utilization of CPU cores when training on four GPUs.

[0017] However, we can see that neither the image classification models such as ResNet101 and VGG19 nor the language processing models such as BERT-Large and BERT-base fully utilize all available CPU cores.

[0018] The difference in utilization between different models contributes to the fact that, while GPU time increases, the preprocessing ratio decreases when preprocessing time remains the same. Furthermore, text preprocessing is generally simpler and easier than image preprocessing, reducing CPU usage for CPU-side computations of language models.

[0019] However, despite these variations, all models utilize only a portion of the allocated CPU cores.

[0020] The technical problem that this specification seeks to solve is to solve at least one of the above problems.

[0021] Another technical challenge that this specification aims to address is to provide a method for deep neural network inference based on idle CPU resources that runs CPU-intensive inference workloads utilizing batch inference workload and online inference workload strategies on GPU servers.

[0022] Another technical challenge that this specification aims to address is to provide a method for deep neural network inference based on idle CPU resources, in which the CPU-intensive deep neural network inference workload does not interfere with the main training task.

[0023] Another technical challenge that this specification aims to address is to provide a deep neural network inference method based on idle CPU resources that can effectively utilize wasted CPU cycles to obtain free inference workload throughput by providing deterministic inference services on GPU servers without degrading the performance of training jobs.

[0024] Another technical challenge that this specification seeks to address is to provide a deep neural network inference method based on idle CPU resources, which allocates idle CPU cores to inference services during the training process of a GPU server.

[0025] Another technical challenge that this specification aims to address is to provide a deep neural network inference method based on idle CPU resources that can maximize resource utilization in a training cluster by efficiently utilizing wasted CPU cores without degrading the performance of existing GPU tasks.

[0026] The technical problems to be achieved in this specification are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which this specification pertains from the description below.

[0027] A deep neural network inference method based on idle CPU resources according to the present specification may include the steps of: when a training task is scheduled and all GPUs are used, classifying CPU cores into groups for each training task, and classifying unallocated idle CPU cores into a U group (Unallocated group); executing a training task using the CPU cores of the group for each training task, and executing the online inference task using the idle CPU cores of the U group when there is a request for an online inference task; and additionally executing a batch inference task using at least one of the idle CPU cores of the U group and the idle CPU cores of the group for each training task when there is a request for a batch inference task.

[0028] And when the idle CPU cores of the group for each training task execute the batch inference task, the idle CPU cores of the group for each training task can execute the training task by setting it to a higher priority than the batch inference task.

[0029] With this configuration, whenever a training task requires CPU cores, it can be smoothly prioritized over batch inference tasks.

[0030] The time for which the idle CPU cores of the above U group perform an online inference task can be set to be the same as the training time of the group with the minimum training time among the groups for each training task.

[0031] With this configuration, throughput can be maximized while maintaining the required latency.

[0032] When running the above batch inference job, instead of running a single instance with multiple threads, you can deploy multiple instances, each with a single thread.

[0033] This configuration addresses the common bottleneck where all but one thread completes execution, while the remaining threads are interrupted by the ongoing training process.

[0034] A method for deep neural network inference based on idle CPU resources according to the present specification may include a first step of retrieving a training task from a queue; a second step of assigning a GPU to the training task; a third step of determining whether there is a GPU remaining that has not been assigned to the training task; a fourth step of returning to the first step if the result of the third step is "yes"; a fifth step of registering an event handler if the result of the third step is "no"; and a sixth step of executing the inference task.

[0035] The fifth step may include a step 5-1 of reorganizing the groups by classifying CPU cores into groups for each training task, classifying unallocated CPU cores into a U group (Unallocated group); a step 5-2 of updating the lease time for each training task; and a step 5-3 of executing the training task using the CPU cores allocated to the training task.

[0036] The above 6th step may include a 6-1 step of determining whether there is a request for an online inference task; a 6-2 step of executing an online inference task using idle CPU cores of the U group when the result of the 6-1 step is "yes"; and a 6-3 step of executing a batch inference task using idle cores of the U group when the result of the 6-1 step is "no".

[0037] The idle CPU cores of the above U group can proceed to step 6-3 to execute batch inference tasks after completing the online inference task in step 6-2.

[0038] The idle CPU cores in each group of training tasks can additionally run batch inference tasks.

[0039] When the idle CPU cores of the group for each training task execute the batch inference task, the idle CPU cores of the group for each training task can execute the training task by setting it to a higher priority than the batch inference task.

[0040] The time for which the idle CPU cores of the above U group perform an online inference task can be set to be the same as the training time of the group with the minimum training time among the groups for each training task.

[0041] When running the above batch inference task, multiple instances, each using a single thread, can be deployed.

[0042] According to this specification, CPU-intensive inference workloads can be executed on GPU servers utilizing batch inference workload and online inference workload strategies.

[0043] And CPU-intensive deep neural network inference workloads can avoid interfering with the primary training task.

[0044] And because it can provide deterministic inference services on GPU servers without compromising the performance of training jobs, it can effectively utilize wasted CPU cycles to obtain free inference workload throughput.

[0045] And during the training process on the GPU server, idle CPU cores can be allocated to the inference service.

[0046] The effects that can be obtained from this specification are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the technical field to which this specification belongs from the description below.

[0047] The accompanying drawings, which are incorporated in and are intended to aid in the understanding of the present specification and are a part of the detailed description, provide embodiments of the present specification and, together with the detailed description, explain the technical features of the present specification.

[0048] Figure 1 is a graph showing the utilization of CPU cores for each model in a training task.

[0049] Figure 2 is a flow chart of a deep neural network inference method based on idle CPU resources according to the present specification.

[0050] Figure 3 is a flow chart showing the detailed configuration of the step of registering the event handler illustrated in Figure 2.

[0051] Figure 4 is a diagram showing the process of classifying CPU cores in Figure 3.

[0052] Figure 5 is a flow chart showing the detailed configuration of the steps for executing the inference task illustrated in Figure 2.

[0053] Figure 6 is a diagram showing the process of executing the inference task in Figure 5.

[0054] Figure 7 is a graph showing the normalized repetition time for each model.

[0055] Figure 8 is a graph showing the processing capacity per second of batch inference tasks for each model.

[0056] Figure 9 is a graph showing the P99th percentile delay time for each model.

[0057] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components are given the same reference numbers and redundant descriptions thereof will be omitted.

[0058] The suffixes "assembly" and "part" used for components in the following description are given or used interchangeably only for the convenience of writing specifications, and do not have distinct meanings or roles in themselves.

[0059] In addition, when describing the embodiments disclosed in this specification, if it is determined that a detailed description of a related known technology may obscure the gist of the embodiments disclosed in this specification, the detailed description is omitted.

[0060] In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of this specification.

[0061] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.

[0062] When it is said that a component is "coupled to" or "in contact with" another component, it should be understood that it may be directly coupled to or in direct contact with that other component, but there may also be other components present in between.

[0063] On the other hand, when it is said that a component is "directly coupled" to or "in direct contact with" another component, it should be understood that there are no other components in between.

[0064] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0065] In this application, terms such as “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in the specification, but should be understood not to exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0066] Hereinafter, the present specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers and redundant descriptions thereof will be omitted.

[0067] Hereinafter, with reference to FIGS. 2 to 6, a method for deep neural network inference based on idle CPU resources according to the present specification will be described.

[0068] FIG. 2 is a flow chart of a deep neural network inference method based on idle CPU resources according to the present specification, FIG. 3 is a flow chart showing a detailed configuration of a step of registering an event handler shown in FIG. 2, and FIG. 4 is a diagram showing a process of classifying CPU cores in FIG. 3.

[0069] And Fig. 5 is a flow chart showing the detailed configuration of the steps for executing the inference task illustrated in Fig. 2, and Fig. 6 is a drawing showing the process for executing the inference task in Fig. 5.

[0070] A method for performing deep neural network inference based on idle CPU resources according to the present specification comprises the steps of: when a training task is scheduled and all GPUs are used, classifying CPU cores into groups for each training task, and classifying unallocated idle CPU cores into a U group (Unallocated group); executing a training task using the CPU cores of the group for each training task, and executing the online inference task using the idle CPU cores of the U group when there is a request for an online inference task; and additionally executing a batch inference task using at least one of the idle CPU cores of the U group and the idle CPU cores of the group for each training task when there is a request for a batch inference task.

[0071] And when the idle CPU cores of the group for each training task execute the batch inference task, the idle CPU cores of the group for each training task can execute the training task by setting it to a higher priority than the batch inference task.

[0072] With this configuration, whenever a training task requires CPU cores, it can be smoothly prioritized over batch inference tasks.

[0073] The time for which the idle CPU cores of the above U group perform an online inference task can be set to be the same as the training time of the group with the minimum training time among the groups for each training task.

[0074] With this configuration, throughput can be maximized while maintaining the required latency.

[0075] When running the above batch inference job, instead of running a single instance with multiple threads, you can deploy multiple instances, each with a single thread.

[0076] This configuration addresses the common bottleneck where all but one thread completes execution, while the remaining threads are interrupted by the ongoing training process.

[0077] More specifically, the idle CPU resource-based deep neural network inference method according to the present specification is described. In the idle CPU resource-based deep neural network inference method according to the present specification, a first step of retrieving a training task from a queue is executed.

[0078] Afterwards, a second step of allocating a GPU to a training task is executed, a third step of determining whether there are GPUs remaining that are not allocated to the training task is executed, and if the result of the third step is "yes", a fourth step of returning to the first step is executed, and if the result of the third step is "no", a fifth step of registering an event handler and a sixth step of executing an inference task are executed.

[0079] In the fifth step of registering an event handler, as illustrated in Fig. 3, the CPU cores are classified into groups for each training task, and the groups are reorganized. Step 5-1 of classifying unallocated CPU cores into a U group (Unallocated group) is executed, and step 5-2 of updating the lease time for each training task is executed, and then step 5-3 of executing the training task using the CPU cores allocated to the training task is executed.

[0080] Figure 4 shows the strategic allocation and grouping of CPU cores for various training tasks and their durations on a single server node.

[0081] In Figure 4, the CPU cores marked in pink, yellow, and green are groups allocated for each training task, and the CPU cores marked in white are unallocated idle CPU cores and are classified as the U group (Unallocated group).

[0082] Group C, shown in pink, has a lease time of 4 hours, Group B, shown in yellow, has a lease time of 3 hours, and Group A, shown in green, has a lease time of 2 hours. Therefore, Groups A through C execute training tasks during their lease times.

[0083] The key aspect here is to strategically utilize the unallocated idle CPU cores of the U group.

[0084] These idle CPU cores are designated to run online inference tasks with the goal of achieving the desired throughput within the Service Level Objective (SLO) constraints.

[0085] However, this execution is conditional. That is, it starts only when there are no GPUs available on the server, and the time for which idle CPU cores perform online inference tasks can be set equal to the training time of the group with the minimum training time among the groups for each training task.

[0086] That is, in the case of Fig. 4, the time for which the idle CPU core performs the online inference task can be set to 2 hours, which is the lease time of group A indicated in green.

[0087] This is to secure the relevant CPU cores and GPUs after a certain period of time has passed.

[0088] The availability of these GPUs could potentially attract new training tasks that may require three or more CPU cores.

[0089] Therefore, the idle CPU cores of the U group are allocated for such emergency situations.

[0090] Additionally, the idle CPU cores in group U can be focused on online inference tasks, while groups related to training (group A, group B, group C) can be additionally designated to accommodate batch inference tasks.

[0091] In this example scenario, the online inference job runs in the U group for a specified 2 hours when GPUs are unavailable.

[0092] At the same time, batch inference tasks are planned to run together with training tasks assigned to groups A, B, and C to ensure an efficient and dynamic resource allocation mechanism.

[0093] Batch inference tasks can be executed using CPU cores that are intermittently idle among the CPU cores allocated to Group A, Group B, and Group C.

[0094] And when the idle CPU cores of the group for each training task execute the batch inference task, the idle CPU cores of the group for each training task can execute the training task by setting it to a higher priority than the batch inference task.

[0095] Referring to FIGS. 5 and 6, the sixth step of executing a batch inference task may include a step 6-1 of determining whether there is a request for an online inference task, a step 6-2 of executing an online inference task using an idle CPU core of the U group when the result value of the step 6-1 is “yes”, and a step 6-3 of executing a batch inference task using an idle core of the U group when the result value of the step 6-1 is “no”.

[0096] And, the idle CPU core of the U group can proceed to step 6-3 and execute a batch inference task after completing the online inference task in step 6-2.

[0097] The idle CPU cores in each group of training tasks can additionally run batch inference tasks.

[0098] When the idle CPU cores of the group for each training task execute the batch inference task, the idle CPU cores of the group for each training task can execute the training task by setting it to a higher priority than the batch inference task.

[0099] The time for which the idle CPU cores of the above U group perform an online inference task can be set to be the same as the training time of the group with the minimum training time among the groups for each training task.

[0100] When running the above batch inference task, multiple instances, each using a single thread, can be deployed.

[0101] Hereinafter, the effectiveness of a deep neural network inference method based on idle CPU resources according to the present specification is described.

[0102] Fig. 7 is a graph showing the normalized repetition time for each model, Fig. 8 is a graph showing the throughput per second of the batch inference task for each model, and Fig. 9 is a graph showing the P99th percentile delay time for each model.

[0103] Given that the applicant's approach runs batch inference jobs alongside training jobs on the same set of CPU cores, it is important to show that these batch instances do not interfere with the training job.

[0104] Figure 7 shows the normalized iteration times for training various DNN models when running on 6 CPU cores.

[0105] In this experiment, the GPU server consists of two Intel Xeon Gold 5218 processors, each with 16 cores, for a total of 32 CPU cores. The GPU server also includes four Nvidia Titan RTX GPUs, each with 24 GB of memory. Six different models were selected for evaluation.

[0106] The baseline normalizes all iteration times, i.e., when the training task is the only task utilizing CPU cores.

[0107] If batch inference instances are run alongside training jobs on these 6 cores without prioritization, the overall performance of the training jobs will degrade.

[0108] The degree of performance degradation of a training task depends on the CPU core utilization of the training task and the size of the training model.

[0109] For Bert-Large, Bert-Base, ResNet101, and VGG19, the overall performance does not deteriorate significantly because the CPU-side computational burden is small.

[0110] However, for DeepSpeech and MobileNetV2, which have relatively small GPU-side computations and high preprocessing requirements, sharing batch inference instances and CPU cores has a significant impact on training performance.

[0111] On the other hand, when running a training job using real-time classes, there are no signs of interference from batch instances.

[0112] The applicant's approach allows batch inference tasks to utilize idle CPU core time without impacting the performance of the training workload.

[0113] Figure 8 shows the throughput of a batch inference job on six CPU cores when running with a training workload.

[0114] For training, the applicant utilized PyTorch with Nvidia CUDA support to take advantage of the computational capabilities of GPUs.

[0115] During inference, the applicant utilized a PyTorch build that supports AVX-512 vector operations on CPUs.

[0116] Due to interference from the training workload, the overall performance of the batch inference task was dependent on the slowest thread, so there was no significant progress.

[0117] However, assigning multiple instances to batch inference tasks using a single thread per instance increased overall throughput. In particular, throughput increased by 16x and 26x for Bert-Large and Bert-Base, respectively.

[0118] For DeepSpeech, the throughput increased by 30x, and finally, for ResNet101, MobileNetV2, and VGG19, the throughput increased by 99x, 411x, and 8x, respectively.

[0119] Thus, it was found that the applicant's approach is optimized to maximize throughput in the case of batch inference.

[0120] Referring to Figure 9, for the online inference task, when calculating the possible configurations and latency for a given request rate, the execution time was considered to be 10% more than the profiled execution time to account for interference that may occur when multiple inference instances are executed on the same machine.

[0121] Thanks to the 10%, the applicant was able to safely perform inference without worrying about interference. To systematically understand interference, the applicant used regression analysis to model it by combining different tasks.

[0122] We evaluated approaches to create appropriate configurations for different request rates.

[0123] This figure shows the P99th percentile latency of the three models at different request rates.

[0124] The wait time constraint was set to 200ms and the idle core group was set to 16 CPU cores.

[0125] The applicant has found that the applicant's approach provides an appropriate configuration that satisfies delay constraints and request requirements at all different request rates.

[0126] Specifically, for request rates of 10 and 11, the applicant's approach created a single instance with 16 CPU cores to meet latency requirements. For higher request rates, the applicant's configuration began by creating two instances, each with 8 cores.

[0127] The 99th percentile latency for all these request rates is less than 200ms.

[0128] Additionally, an event occurred where a single instance took slightly longer than the actual execution time, which increased the overall P99th percentile.

[0129] It will be apparent to those skilled in the art that this specification may be embodied in other specific forms without departing from the essential characteristics thereof. Therefore, the foregoing detailed description should not be construed in any way as limiting but rather as illustrative. The scope of this specification should be determined by a reasonable interpretation of the appended claims, and all changes within the scope of equivalents herein are intended to be included within the scope of this specification.

Claims

1. When a training task is scheduled and all GPUs are used, the CPU cores are grouped by training task, and the unallocated idle CPU cores are grouped into the U group (Unallocated group); A step of executing a training task using the CPU cores of the group for each training task, and executing the online inference task using the idle CPU cores of the U group when there is a request for an online inference task; and When there is a request for a batch inference task, a step of additionally executing a batch inference task using at least one of the idle CPU cores of the U group and the idle CPU cores of the training task-specific group. A deep neural network inference method based on idle CPU resources to maximize resource utilization in a training cluster, including:

2. In paragraph 1, A method for deep neural network inference based on idle CPU resources for maximizing resource utilization in a learning cluster, characterized in that when the idle CPU cores of the group for each training task execute the batch inference task, the idle CPU cores of the group for each training task execute the training task by setting it to a higher priority than the batch inference task.

3. In paragraph 1 or 2, A deep neural network inference method based on idle CPU resources for maximizing resource utilization in a learning cluster, characterized in that the time for which the idle CPU cores of the above U group perform an online inference task is the same as the training time of the group with the minimum training time among the groups for each training task.

4. In paragraph 3, A deep neural network inference method based on idle CPU resources for maximizing resource utilization in a training cluster, characterized in that when executing the above batch inference task, multiple instances, each using a single thread, are distributed.

5. The first step is to get a training job from the queue; Step 2: Assigning GPUs to training tasks; A third step is to determine whether there are any remaining GPUs that are not allocated to the training task; If the result value of the above step 3 is “Yes”, the 4th step of returning to the 1st step; If the result value of the above step 3 is "No", the fifth step of registering an event handler; and Step 6: Executing the inference task A deep neural network inference method based on idle CPU resources to maximize resource utilization in a training cluster, including:

6. In paragraph 5, The fifth step above is, Step 5-1: Reorganize the groups by classifying the CPU cores into groups according to the training task, and classify the unallocated CPU cores into the U group (Unallocated group); Step 5-2: Updating the rental time for each training task; and Step 5-3: Executing the training task using the CPU cores assigned to the above training task. A deep neural network inference method based on idle CPU resources to maximize resource utilization in a training cluster, including:

7. In paragraph 6, The above 6th step is, Step 6-1: Determining whether there is a request for an online inference task; If the result value of the above step 6-1 is "Yes", step 6-2 executing an online inference task using the idle CPU cores of the U group; and If the result of the above step 6-1 is “No”, step 6-3 executes the batch inference task using the idle cores of the U group. A deep neural network inference method based on idle CPU resources to maximize resource utilization in a training cluster, including:

8. In paragraph 7, The idle CPU cores of the above U group are: A deep neural network inference method based on idle CPU resources for maximizing resource utilization in a learning cluster, characterized in that after completing the online inference task in the above step 6-2, the method proceeds to the above step 6-3 to execute a batch inference task.

9. In paragraph 8, A deep neural network inference method based on idle CPU resources for maximizing resource utilization in a training cluster, characterized in that idle CPU cores of each group for each training task additionally execute batch inference tasks.

10. In Article 9, A method for deep neural network inference based on idle CPU resources for maximizing resource utilization in a learning cluster, characterized in that when the idle CPU cores of the group for each training task execute the batch inference task, the idle CPU cores of the group for each training task execute the training task by setting it to a higher priority than the batch inference task.

11. In any one of paragraphs 7 to 10, A deep neural network inference method based on idle CPU resources for maximizing resource utilization in a learning cluster, characterized in that the time for which the idle CPU cores of the above U group perform an online inference task is the same as the training time of the group with the minimum training time among the groups for each training task.

12. In Article 11, A deep neural network inference method based on idle CPU resources for maximizing resource utilization in a training cluster, characterized in that when executing the above batch inference task, multiple instances, each using a single thread, are distributed.

Citation Information

Patent Citations

  • Virtual machine scheduling method and apparatus

    JP2021521518A

  • Construction Masonry Brick Walls Coupling Apparatus

    KR102221247B1

  • Terras Folding System for Secondary Battery Cell

    KR102240008B1

  • KR20210034558A

Cited By

  • Network parameter configuration method, equipment, system and storage medium

    CN121907685A