Method, device, computer device and readable storage medium for processing a job

CN113157403BActive Publication Date: 2026-09-25CAMBRICON TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010012302.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-07
Publication Date
2026-09-25
Estimated Expiration
2040-01-07

AI Technical Summary

Technical Problem

[0004]然而,上述分配过程中由于亲和性绑定原则的因素,往往会导致作业需要长时间的等待,进而严重影响作业的执行效率

Benefits of technology

[0028]本申请实施例提供了一种作业处理的方法、装置、计算机设备及可读存储介质。当满足预设处理条件时,CPU根据目标任务包含的目标作业的作业属性,在各节点中确定与目标作业相匹配的第一节点。其中,作业属性包括执行目标作业所需的运算单元的目标数目。然后,CPU通过第一节点和执行目标任务的运算单元所在的第二节点,执行目标任务包含的目标作业。这样,当目标作业需要等待较长时间才可以被该目标作业所等待执行的运算单元执行时,CPU可以通过第一节点和第二节点共同执行该目标作业,从而减少该目标作业的等待时长,提高该目标作业的执行效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113157403B_ABST
    Figure CN113157403B_ABST
Patent Text Reader

Abstract

The application relates to a job processing method and device, computer equipment and a readable storage medium. The method comprises the following steps: when it is detected that a target job meets a preset processing condition, a target node matched with the target job is determined in each node according to a job attribute of the target job, the job attribute comprising a target number of operation units required for executing the target job; and the target job is processed to the target node, so that the target job is executed by the target node. The application can reduce the waiting time of the target job and improve the execution efficiency of the target job.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, and readable storage medium for job processing. Background Technology

[0002] Currently, NUMA (Non-Uniform Memory Access Architecture) is widely used in chip design for artificial intelligence applications. Chips based on NUMA typically include a processor with multiple processing units and multiple memory units. These processing units are usually divided into multiple processing unit groups, and each processing unit group is equipped with at least one memory unit. A processing unit group and its corresponding memory unit constitute a node. In this way, the reading and writing of data required by the processing units within a node can be achieved through the memory units within that node.

[0003] During chip operation, tasks need to be assigned to specific nodes for execution. The specific allocation process is as follows: First, determine the memory size required to execute the task. Then, based on the memory units corresponding to each node, determine the target node with sufficient remaining memory space. For example, the node with the largest remaining memory space can be selected as the target node, or a node with more remaining memory space than the required memory size can be randomly selected as the target node. Then, based on the affinity binding principle, the task is assigned to the target node for execution.

[0004] However, due to the affinity binding principle, the above allocation process often leads to long waiting times for tasks, which in turn seriously affects the efficiency of task execution. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, apparatus, computer equipment, and readable storage medium for processing jobs to address the aforementioned technical problems.

[0006] Firstly, a method for processing jobs is provided, the method comprising:

[0007] When the preset processing conditions are met, based on the job attributes of the target job contained in the target task, a first node matching the target job is determined in each node, wherein the job attributes include the target number of computing units required to execute the target job;

[0008] The target job contained in the target task is executed through the first node and the second node where the computing unit that executes the target task is located.

[0009] As an optional implementation, determining the first node matching the target job among the nodes based on the job attributes of the target job included in the target task includes:

[0010] For each of the nodes, if the number of idle computing units in that node is greater than or equal to the target number, then that node is determined as the first node.

[0011] As an optional implementation, the method further includes:

[0012] Obtain the idle time of the idle computing units in each node;

[0013] If there are tasks containing multiple jobs waiting to be executed in the list of splittable tasks, and the idle time of each idle computing unit is greater than or equal to the preset time threshold, then the preset processing conditions are met.

[0014] As an optional implementation, before determining the first node matching the target job among the nodes based on the job attributes of the target job included in the target task when the preset processing conditions are met, the method further includes:

[0015] Obtain the target task to be executed, and determine the information of each dimension of the target task and the target number of computing units required to execute the target task;

[0016] If the ratio of the product of the information in each dimension to the number of targets is greater than 1, then the target task is added to the list of scalable tasks.

[0017] The affinity mask of the target task is modified according to the preset affinity mask modification rules.

[0018] As an optional implementation, before executing the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located, the method further includes:

[0019] In the mask used for the target task, the position corresponding to the first node is set to 1.

[0020] As an optional implementation, before executing the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located, the method further includes:

[0021] If the bits corresponding to the first node and the second node in the affinity mask and usage mask of the target job are both 1, then the step of executing the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located is executed.

[0022] As an optional implementation, the affinity mask of the target job is the same as the affinity mask of the target task, and the usage mask of the target job is the same as the usage mask of the target task.

[0023] Secondly, a job processing apparatus is provided, the apparatus comprising:

[0024] The first determining module is used to determine, when the preset processing conditions are met, a first node matching the target job in each node according to the job attributes of the target job contained in the target task, wherein the job attributes include the target number of computing units required to execute the target job;

[0025] The execution module is used to execute the target job contained in the target task through the first node and the second node where the computing unit for executing the target task is located.

[0026] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor executes the computer program to implement the steps of any of the methods in the first aspect.

[0027] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the first aspects.

[0028] This application provides a method, apparatus, computer device, and readable storage medium for job processing. When preset processing conditions are met, the CPU determines a first node matching the target job among various nodes based on the job attributes of the target job contained in the target task. The job attributes include the target number of computing units required to execute the target job. Then, the CPU executes the target job contained in the target task through the first node and the second node where the computing units executing the target task are located. In this way, when the target job needs to wait a long time before it can be executed by the computing units it is waiting to execute, the CPU can execute the target job through both the first and second nodes, thereby reducing the waiting time of the target job and improving its execution efficiency. Attached Figure Description

[0029] Figure 1 A schematic diagram of an intelligent processor provided in an embodiment of this application;

[0030] Figure 2 A flowchart illustrating a method for job splitting and affinity mask modification provided in an embodiment of this application;

[0031] Figure 3 A flowchart illustrating a job processing method provided in an embodiment of this application;

[0032] Figure 4 A flowchart illustrating a method for determining processing conditions provided in an embodiment of this application;

[0033] Figure 5 This is a schematic diagram of the structure of a job processing apparatus provided in an embodiment of this application;

[0034] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0035] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0036] It should be understood that the terms "first," "second," "third," and "fourth," etc., in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.

[0037] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0038] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0039] This application provides a job processing method that can be applied to a chip. The chip may include a NUMA-based intelligent processor and a general-purpose processor. The general-purpose processor may be a CPU (central processing unit), etc. The NUMA-based intelligent processor may be an accelerator processor, an IPU (Intelligent Processing Unit), a GPU (Graphics Processing Unit), or other types of processors; this application does not limit the specific type. Specifically, this method can be applied to the aforementioned chip, where the general-purpose processor (CPU) can execute the job processing method to distribute multiple jobs to at least one processing unit in the intelligent processor for execution. The specific execution process of the job processing method of this application is described below.

[0040] Optionally, the intelligent processor of this NUMA architecture includes a processor with multiple computing units and multiple storage units. These computing units are typically divided into multiple computing unit groups, each computing unit group being equipped with at least one storage unit. A computing unit group and its corresponding storage unit constitute a node. The reading and writing of data required by the computing units within a node can be achieved through the storage units within that node, and different nodes can read and write data through a communication interface. Figure 1 This is a schematic diagram of a NUMA-based intelligent processor provided in an embodiment of this application. Figure 1 As shown, the intelligent processor contains 16 processing units and 4 storage units. The intelligent processor is divided into 4 nodes, each containing 4 processing units and 1 storage unit.

[0041] Figure 1This diagram illustrates an intelligent processor only. In other possible implementations, each node may contain more than four processing units and one storage unit, which may include multiple sub-storage units. For example, each node may include four sub-nodes, meaning each node may include 16 processing units. Each sub-node contains four processing units and one sub-storage unit, and the four sub-nodes can be arranged in the same manner as four nodes. Furthermore, the above job processing method can be executed among the various sub-nodes of a single node; the specific execution process is detailed in the following description of the job processing method.

[0042] When a task is scheduled to the software queue, the processor, based on the number of computing units required to execute the task, allocates the desired computing units from the node containing the task's data storage unit, and increments the wait reference count (clu_wait_ref) of each desired computing unit by 1. For example, ... Figure 1 As shown, the number of computing units required for this task is 2, and the storage unit for storing the task data is storage unit 1. Then the processor can determine computing unit 1 and computing unit 2 as the computing units expected by the task in node 1, and increment the waiting reference count of computing unit 1 and computing unit 2 by 1.

[0043] Once the processor determines the computational unit to execute the task, the task is scheduled to the hardware queue, and the real reference count (i.e., clu_real_ref) of each computational unit executing the task is incremented by 1. For example, as Figure 1 As shown, after the processor determines that the arithmetic units executing the task are arithmetic unit 1 and arithmetic unit 2, the processor can increment the actual reference count of arithmetic unit 1 and arithmetic unit 2 by 1.

[0044] When the task is completed, the waiting reference counts of each arithmetic unit expected by the task are decremented by 1, and the actual reference counts of each arithmetic unit executing the task are decremented by 1. For example, after arithmetic unit 1 and arithmetic unit 2 complete the task, the processor can decrement the waiting reference counts and actual reference counts of arithmetic unit 1 and arithmetic unit 2 by 1. If the arithmetic units expected by the task are migrated, the waiting reference counts of each source arithmetic unit expected by the task are decremented by 1, and the waiting reference counts of each destination arithmetic unit expected by the task are incremented by 1. For example, as... Figure 1 As shown, if the computational units expected by the task are moved from computational units 1 and 2 to computational units 3 and 4, the processor will decrement the wait reference counts of computational units 1 and 2 by 1 and increment the wait reference counts of computational units 3 and 4 by 1.

[0045] This application first describes the division of tasks and the modification of affinity masks, such as... Figure 2 As shown, the specific processing procedure is as follows:

[0046] Step 201: Obtain the target task to be executed, and determine the target task's dimensional information and the target number of computing units required to execute the target task.

[0047] In implementation, once a task (i.e., the target task) is scheduled to the software queue, the processor can determine the target task's dimensions (i.e., dimX, dimY, and dimZ) and the target number of computing units (i.e., kernel_class) required to execute the target task. Then, the processor can calculate the ratio of the product of the dimensions (i.e., dimX * dimY * dimZ) to the target number and determine if this ratio is greater than 1. If the ratio is greater than 1, it indicates that the target task can be split into multiple jobs, and the processor executes step 202. If the ratio is less than or equal to 1, it indicates that the target task cannot be split into multiple jobs.

[0048] Step 202: If the ratio of the product of the information in each dimension to the number of targets is greater than 1, then add the target task to the list of splittable tasks.

[0049] In implementation, if the ratio is greater than 1, it indicates that the target task can be split into multiple jobs. Accordingly, the processor can add the target task to the list of splittable tasks. This list of splittable tasks stores tasks that can be split into multiple jobs; it can be a linked list or other types of lists, which are not limited in this embodiment. Furthermore, once all jobs contained in a task in the list of splittable tasks have been executed, the processor can remove that task from the list.

[0050] Optionally, when it is determined that the target task can be split into multiple jobs, the target task can be sent to the scheduler. The scheduler can then split the target task into multiple jobs based on its various dimensions and task attributes such as the target number of computing units required. Further optionally, the scheduler can be a hardware scheduler located on a chip, which may include multiple circuit modules such as a task splitting unit. Of course, the scheduler can also be a software scheduler; no specific limitation is made here.

[0051] Step 203: Modify the affinity mask of the target task according to the preset affinity mask modification rules.

[0052] In implementation, after the processor splits the target task into multiple target jobs, it can modify the affinity mask of the target task according to a preset affinity mask modification rule. The affinity mask of the target task indicates which nodes can execute the target task. The affinity mask includes bits representing the total number of nodes included in the intelligent processor. Each bit uniquely corresponds to one node. If a bit is 1, it means the node corresponding to that bit can execute the target task; if a bit is 0, it means the node corresponding to that bit cannot execute the target task. The affinity mask modification rule can be set by technicians based on the range of nodes processed by the job. For example, if the affinity mask modification rule is that the task is migrated to all nodes, and the original affinity mask of the target task is 0001, then the processor can modify the affinity mask of the target task to 1111 according to the affinity mask modification rule. For example, if the affinity mask modification rule allows a task to migrate to nodes 3 and 4, and the original affinity mask of the target task is 0001, then the processor can modify the affinity mask of the target task to 1101 according to the affinity mask modification rule. Steps 202 and 203 above are not sequential.

[0053] It should be noted that since the target job is obtained by splitting the target task, the affinity mask of the target job is the same as the affinity mask of the target task.

[0054] The following will describe a job processing method provided in this application with reference to specific embodiments, such as... Figure 3 As shown, the specific processing procedure is as follows.

[0055] Step 301: When the preset processing conditions are met, determine the first node that matches the target job among all nodes based on the job attributes of the target job contained in the target task. The job attributes include the target number of computing units required to execute the target job.

[0056] In implementation, after the processor schedules a task (i.e., the target task) to the software queue, it can determine whether the target task can be split into multiple target jobs. If the target task can be split into multiple target jobs, it can be determined that the target task is a task whose affinity can be relaxed. The processor can then further determine whether preset processing conditions are met. The process of determining whether the preset processing conditions are met will be described in detail later. When the preset processing conditions are met, the processor can determine the first node matching the target job among all nodes based on the job attributes of the target job. The job attributes of the target job include the target number of computing units required to execute the target job.

[0057] Optionally, the processor determines the first node that matches the target job in each node based on the job attributes of the target job as follows: for each node, if the number of idle computing units in the node is greater than or equal to the target number, then the node is determined as the first node.

[0058] In implementation, when preset processing conditions are met, for each node, the processor can obtain the number of idle computing units (i.e., those with a true reference count of 0) in that node. Then, the processor can determine whether the number of idle computing units in that node is greater than or equal to the target number. If the number of idle computing units in that node is greater than or equal to the target number, it means that the node can execute the target job contained in the target task. Accordingly, the processor can identify that node as the first node. If the number of idle computing units in that node is less than the target number, it means that the node cannot execute the target job contained in the target task. Accordingly, that node is not the first node.

[0059] As an optional implementation method, such as Figure 4 As shown, the processor determines whether the preset processing conditions are met through the following process:

[0060] Step 401: Obtain the idle time of the idle computing units in each node.

[0061] In practice, for each node, the processor can obtain the idle time of the idle computing units (i.e., the actual reference count is equal to 0) in that node.

[0062] Step 402: If there are tasks containing multiple jobs waiting to be executed in the list of splittable tasks, and the idle time of each idle computing unit is greater than or equal to the preset time threshold, then the preset processing conditions are met.

[0063] In implementation, after obtaining the idle time of each idle computing unit, the processor can further determine whether there are any tasks containing multiple jobs waiting to be executed in the splittable task list, and whether any idle time in each idle computing unit is greater than or equal to a preset time threshold. This preset time threshold can be set by technicians based on experience. If there are tasks containing multiple jobs waiting to be executed in the splittable task list, and the idle time of each idle computing unit is greater than or equal to the preset time threshold, then the processor can allocate the tasks to be executed in the splittable task list to the idle computing units in each node for execution. Accordingly, the processor can determine that the preset processing conditions are met. If there are no tasks waiting to be executed in the splittable task list, or the idle time of each idle computing unit is not greater than or equal to the preset time threshold, then the processor cannot allocate the tasks to be executed in the splittable task list to the idle computing units in each node for execution. Accordingly, the processor can determine that the preset processing conditions are not met.

[0064] Step 302: Execute the target job contained in the target task through the first node and the second node where the computing unit for executing the target task is located.

[0065] In implementation, after the processor determines the first node, it can execute the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located. In this way, when the target job needs to wait for a long time before it can be executed by the computing unit that the target job is waiting to be executed, the processor can execute the target job through the first node and the second node, thereby reducing the waiting time of the target job and improving its execution efficiency.

[0066] As an optional implementation, before the processor executes the target job contained in the target task through the first node and the second node, the processor can also modify the usage mask of the target job according to the determined first node. The specific process is as follows: in the usage mask of the target task, the position corresponding to the first node is set to 1.

[0067] In implementation, the usage mask of the target task is used to indicate which node among the nodes is determined to execute the target task. The usage mask includes bits representing the total number of nodes in the intelligent processor. Each bit uniquely corresponds to one node. If a bit is 1, it means that the node corresponding to that bit is determined to execute the target task; if a bit is 0, it means that the node corresponding to that bit is not executing the target task. After the processor determines the first node of the target job, before scheduling the target job to the hardware queue, the bit corresponding to the first node in the usage mask of the target task can be set to 1. For example, the original usage mask of the target task is 0001. Assuming that the first node is node 2 and node 4, the modified usage mask of the target task is 1011.

[0068] It should be noted that since the target job is obtained by splitting the target task, the usage mask of the target job is the same as that of the target task.

[0069] As an optional implementation, before the processor executes the target job contained in the target task through the first node and the second node, the processor can also determine whether the first node and the second node can execute the target job based on the affinity mask and the usage mask of the target job. The specific processing procedure is as follows: if the bits corresponding to the first node and the second node in the affinity mask and the usage mask of the target job are both 1, then the step of executing the target job contained in the target task through the first node and the second node where the arithmetic unit of the target task is located is executed.

[0070] In implementation, after obtaining the affinity mask and usage mask of the target job, the processor, for each first node, can determine whether all the corresponding bits of the first node are 1 in the affinity mask and usage mask. If all the corresponding bits of the first node are 1, it means that the first node can execute the target job. Similarly, for the second node where the arithmetic unit executing the target task is located, the processor can determine whether all the corresponding bits of the second node are also 1 in the affinity mask and usage mask. If all the corresponding bits of the second node are 1, it means that the second node can execute the target job. Accordingly, the processor can execute the target job contained in the target task through the first node and the second node. If there is a 0 in the corresponding bits of the first node, it means that the first node cannot execute the target job, and the processor will only execute the target job through the second node. For example, if the affinity mask of the target job is 1101, the mask used is 1001, the first node is node 2 and node 4, and the second node is node 1, then the bits corresponding to node 1 are all 1, the bits corresponding to node 4 are all 1, and the bits of node 2 in the affinity mask are 0. The processor can execute the target job through nodes 1 and 4.

[0071] This application provides a method for job processing. When preset processing conditions are met, the processor determines a first node matching the target job among all nodes based on the job attributes of the target job contained in the target task. The job attributes include the target number of computing units required to execute the target job. Then, the processor executes the target job contained in the target task through the first node and the second node where the computing units executing the target task are located. In this way, when the target job needs to wait a long time before it can be executed by the computing units it is waiting for, the processor can execute the target job through both the first and second nodes, thereby reducing the waiting time and improving the execution efficiency of the target job.

[0072] This application also provides a job processing apparatus, such as... Figure 5 As shown, the device includes:

[0073] The first determining module 510 is used to determine, when the preset processing conditions are met, a first node matching the target job in each node according to the job attributes of the target job contained in the target task. The job attributes include the target number of computing units required to execute the target job.

[0074] The execution module 520 is used to execute the target job contained in the target task through the first node and the second node where the computing unit for executing the target task is located.

[0075] As an optional implementation, the first determining module 510 is specifically used for:

[0076] For each node, if the number of idle computing units in that node is greater than or equal to the target number, then that node is designated as the first node.

[0077] As an optional implementation, the device further includes:

[0078] The acquisition module is used to obtain the idle time of idle computing units in each node;

[0079] The second determining module is used to determine if the preset processing conditions are met if there is a task containing multiple jobs waiting to be executed in the list of divisible tasks, and the idle time of each idle computing unit is greater than or equal to a preset time threshold.

[0080] As an optional implementation, the device further includes:

[0081] The third determination module is used to obtain the target task to be executed, and to determine the target task's various dimensions of information and the target number of computing units required to execute the target task.

[0082] Add a module to add the target task to the list of splittable tasks if the ratio of the product of the information in each dimension to the number of targets is greater than 1.

[0083] The modification module is used to modify the affinity mask of the target task according to the preset affinity mask modification rules.

[0084] As an optional implementation, the device further includes:

[0085] The setting module is used to set the position corresponding to the first node to 1 in the usage mask of the target task.

[0086] As an optional implementation, the device further includes:

[0087] The fourth determining module is used to trigger the execution module 520 to execute the steps of the target job contained in the target job through the first node and the second node where the computing unit of the target task is located, if the corresponding bits of the first node and the usage mask of the target job are both 1.

[0088] As an optional implementation, the affinity mask of the target job is the same as the affinity mask of the target task, and the usage mask of the target job is the same as the usage mask of the target task.

[0089] This application provides a job processing apparatus. When preset processing conditions are met, the CPU determines a first node matching the target job among all nodes based on the job attributes of the target job contained in the target task. The job attributes include the target number of computing units required to execute the target job. Then, the CPU executes the target job contained in the target task through the first node and the second node where the computing units executing the target task are located. In this way, when the target job needs to wait a long time before it can be executed by the computing units it is waiting for, the CPU can execute the target job through both the first and second nodes, thereby reducing the waiting time and improving the execution efficiency of the target job.

[0090] In one embodiment, a computer device is provided, such as Figure 6 As shown, it includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the above-described method steps for processing the job.

[0091] In one embodiment, a computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the above-described job processing method.

[0092] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this disclosure. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.

[0093] It should be further explained that, although Figure 2-4 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2-4 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0094] It should be understood that the above-described device embodiments are merely illustrative, and the device disclosed herein can be implemented in other ways. For example, the division of units / modules described in the above embodiments is only a logical functional division, and other division methods may be used in actual implementation. For example, multiple units, modules, or components may be combined, integrated into another system, or some features may be ignored or not executed.

[0095] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments disclosed herein can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0096] If the integrated unit / module is implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0097] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution disclosed herein, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0098] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0099] The foregoing can be better understood in accordance with the following terms:

[0100] Clause A1 corresponds to right 1; Clause A2 corresponds to right 2; Clause A3 corresponds to right 3; Clause A4 corresponds to right 4; Clause A5 corresponds to right 5; Clause A6 corresponds to right 6; Clause A7 corresponds to right 7; Clause A8 corresponds to right 8; Clause A9 corresponds to right 9.

[0101] For example, Clause A1, a method for processing a job, said method comprising:

[0102] When the preset processing conditions are met, based on the job attributes of the target job contained in the target task, a first node matching the target job is determined in each node, wherein the job attributes include the target number of computing units required to execute the target job;

[0103] The target job contained in the target task is executed through the first node and the second node where the computing unit that executes the target task is located.

[0104] Clause A2, the method described in Clause A1, wherein determining the first node matching the target job among the nodes based on the job attributes of the target job contained in the target task includes:

[0105] For each of the nodes, if the number of idle computing units in that node is greater than or equal to the target number, then that node is determined as the first node.

[0106] Clause A3. The method described in Clause A1 further includes:

[0107] Obtain the maximum idle time of the idle computing units in each node;

[0108] If there are tasks containing multiple jobs waiting to be executed in the list of splittable tasks, and the idle time of each idle computing unit is greater than or equal to the preset time threshold, then the preset processing conditions are met.

[0109] Clause A4. According to the method described in Clause A1, before determining the first node matching the target job among the nodes based on the job attributes of the target job included in the target task when the preset processing conditions are met, the method further includes:

[0110] Obtain the target task to be executed, and determine the information of each dimension of the target task and the target number of computing units required to execute the target task;

[0111] If the ratio of the product of the information in each dimension to the number of targets is greater than 1, then the target task is added to the list of scalable tasks.

[0112] The affinity mask of the target task is modified according to the preset affinity mask modification rules.

[0113] Clause A5. According to the method described in Clause A1, before executing the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located, the method further includes:

[0114] In the mask used for the target task, the position corresponding to the first node is set to 1.

[0115] Clause A6. According to the method described in Clause A1, before executing the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located, the method further includes:

[0116] If the bits corresponding to the first node and the second node in the affinity mask and usage mask of the target job are both 1, then the step of executing the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located is executed.

[0117] Clause A7. As described in Clause A6, the affinity mask of the target job is the same as the affinity mask of the target task, and the usage mask of the target job is the same as the usage mask of the target task.

[0118] Clause A8. A job processing apparatus, said apparatus comprising:

[0119] The first determining module is used to determine, when the preset processing conditions are met, a first node matching the target job in each node according to the job attributes of the target job contained in the target task, wherein the job attributes include the target number of computing units required to execute the target job;

[0120] The execution module is used to execute the target job contained in the target task through the first node and the second node where the computing unit for executing the target task is located.

[0121] Clause A9. A computer device including a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the computer program to implement the steps of any one of Clauses A1 to A7.

[0122] Clause A10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of Clauses A1 to A7.

[0123] The embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this disclosure. Furthermore, any changes or modifications made by those skilled in the art based on the ideas of this disclosure, and on the specific implementation methods and application scope of this disclosure, are all within the scope of protection of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.

Claims

1. A method for processing jobs, characterized in that, For chips applied to NUMA architecture, the method includes: When a target task can be broken down into multiple target jobs, the target task is determined to be a task whose affinity can be relaxed, and the affinity mask of the target task is modified. The affinity mask is used to represent the nodes in each node that can execute the target task. When the target task is a task that can relax affinity and meets the preset processing conditions, according to the job attributes of the target job included in the target task, a first node matching the target job is determined in each node. The job attributes include the target number of computing units required to execute the target job. The node consists of a computing unit group and a storage unit corresponding to the computing unit group. The computing unit group contains multiple computing units. Based on the determined first node, modify the usage mask of the target task, wherein the usage mask is used to represent the node among the nodes that is determined to execute the target task; If, based on the affinity mask and usage mask of the target task, it is determined that the first node and the second node where the computing unit executing the target task is located can execute the target job, then the target job contained in the target task is executed through the first node and the second node where the computing unit executing the target task is located.

2. The method according to claim 1, characterized in that, The step of determining the first node matching the target task among each node based on the task attributes of the target task includes: For each of the nodes, if the number of idle computing units in that node is greater than or equal to the target number, then that node is determined as the first node.

3. The method according to claim 1, characterized in that, The method further includes: Obtain the idle time of the idle computing units in each node; If there are tasks containing multiple jobs waiting to be executed in the list of splittable tasks, and the idle time of each idle computing unit is greater than or equal to the preset time threshold, then the preset processing conditions are met.

4. The method according to claim 1, characterized in that, Modifying the affinity mask of the target task includes: Obtain the target task to be executed, and determine the information of each dimension of the target task and the target number of computing units required to execute the target task; If the ratio of the product of the information in each dimension to the number of targets is greater than 1, then the target task is added to the list of scalable tasks. The affinity mask of the target task is modified according to the preset affinity mask modification rules.

5. The method according to claim 1, characterized in that, The step of modifying the usage mask of the target task based on the determined first node includes: In the mask used for the target task, the position corresponding to the first node is set to 1.

6. The method according to claim 1, characterized in that, Before executing the target job contained in the target task through the first node and the second node where the computing unit executing the target task is located, the method further includes: If the bits corresponding to the first node and the second node in the affinity mask and usage mask of the target job are both 1, it is determined that the first node and the second node where the computing unit executing the target task is located can execute the target job.

7. The method according to claim 6, characterized in that, The affinity mask of the target job is the same as the affinity mask of the target task, and the usage mask of the target job is the same as the usage mask of the target task.

8. A work processing apparatus, characterized in that, A chip used in a NUMA architecture, the device comprising: The first determining module is configured to, when a target task can be split into multiple target jobs, determine that the target task is a task whose affinity can be relaxed, and modify the affinity mask of the target task, wherein the affinity mask is used to represent the nodes in each node that can execute the target task; when the target task is a task that can be relaxed and meets preset processing conditions, determine a first node matching the target job in each node according to the job attributes of the target jobs included in the target task, wherein the job attributes include the target number of computing units required to execute the target job, and the node consists of a computing unit group and a storage unit corresponding to the computing unit group, wherein the computing unit group contains multiple computing units; An execution module is configured to modify the usage mask of the target task based on the determined first node, wherein the usage mask is used to indicate the node among the nodes that is determined to execute the target task; and if, based on the affinity mask and the usage mask of the target task, it is determined that the first node and the second node where the computing unit executing the target task is located can execute the target job, the target job contained in the target task is executed through the first node and the second node where the computing unit executing the target task is located.

9. A computer device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Task-dynamic dispatching method under distributed computation mode in cloud computing environment

    CN102073546A

  • Data processing method, data processing device, terminal, and readable storage medium

    CN108491263A

  • Resource scheduling method, apparatus, and computer-readable storage medium

    CN109144710A

  • Method, apparatus, and computer program product for providing a self-tunable parameter used for dynamically yielding an idle processor

    US20060048160A1