Reduction Scheduling Method and Device
By setting multiple thread bundles inside the execution unit and using vector registers and scalar registers to perform reduction operations, the problem of low reduction efficiency in the prior art is solved, and more efficient computing power utilization is achieved.
Patent Information
- Application Number
- CN202410801566.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-06-20
AI Technical Summary
The prior art is less efficient than frequent access to shared memory and multiple iterations when reducing operations using hardware chips.
By setting multiple thread bundles inside the execution unit, reducing operations using vector registers and scalar registers, access to shared memory is reduced, and reducing within the thread bundle and within the execution unit is realized.
It improves the efficiency of reduction operations, reduces frequent access and iteration of shared memory, and improves computing power.
Smart Images

Figure CN118796389B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of chips, and particularly to a reduction scheduling method and apparatus. Background Art
[0002] Reduction refers to performing a certain operation on multiple input data to reduce the multiple input data to a single result. It usually reduces a vector including multiple vector elements to a single value. The operation can be minimum, maximum, sum, sum of squares, logical AND, logical OR, vector dot product, etc. Computers usually perform reduction operations using software. Performing reduction operations using software is not conducive to improving operation efficiency, cannot improve computing power, and does not meet the development needs of the large language model era. Therefore, a method of using a hardware chip for reduction has been proposed.
[0003] Typically, in the method of using a hardware chip for reduction, each reduction vector in the target task is loaded into the shared memory shared by multiple execution units (EUs) in the chip. Then, a tree reduction or interleaved reduction is performed in the shared memory to obtain the reduction result of the target task. Therefore, the entire task reduction process requires frequent access to the shared memory and multiple iterations in the shared memory, resulting in low reduction efficiency. Summary of the Invention
[0004] Embodiments of the present disclosure provide a reduction scheduling method and apparatus, which can improve the efficiency of reduction operations.
[0005] According to one aspect of the present disclosure, a reduction scheduling method is provided for a scheduling unit that schedules execution units. The execution units include reduction units, vector registers, and scalar registers. Among them, the vector registers include vector register regions corresponding to each warp, and the vector register regions include vector register rows corresponding to each thread in the warp; the scalar registers include scalar register regions corresponding to each warp, and the scalar register regions include scalar register bits corresponding to each thread in the warp. The reduction scheduling method includes:
[0006] For a target reduction task, set a plurality of warps corresponding to the execution units;
[0007] Store each reduction vector of the target reduction task in each vector register row;
[0008] Through each warp, execute a first reduction instruction to, for each thread in the warp, use the reduction unit to reduce each vector element of the reduction vector in the vector register row corresponding to the thread, obtain a first reduction result, and store it in the scalar register bit corresponding to the thread;
[0009] Execute a second reduction instruction through a target warp among the multiple warps corresponding to the execution unit, so as to utilize the reduction unit to reduce each scalar register bit in the scalar register within the execution unit to obtain a second reduction result.
[0010] According to an aspect of the present disclosure, a reduction scheduling device is provided, including a scheduling unit and at least one execution unit. The execution unit includes a reduction unit, a vector register, and a scalar register. Among them, the vector register includes a vector register area corresponding to each warp, and the vector register area includes vector register rows corresponding to each thread in the warp; the scalar register includes a scalar register area corresponding to each warp, and the scalar register area includes scalar register bits corresponding to each thread in the warp; the scheduling unit is configured to:
[0011] For a target reduction task, set multiple warps corresponding to the execution unit;
[0012] Store each vector to be reduced of the target reduction task into each vector register row;
[0013] Through each warp, execute a first reduction instruction to, for each thread in the warp, utilize the reduction unit to reduce each vector element of the vector to be reduced in the vector register row corresponding to the thread to obtain a first reduction result, which is stored in the scalar register bit corresponding to the thread;
[0014] Execute a second reduction instruction through a target warp among the multiple warps corresponding to the execution unit, so as to utilize the reduction unit to reduce each scalar register bit in the scalar register within the execution unit to obtain a second reduction result.
[0015] Optionally, the reduction unit is specifically configured to:
[0016] Reduce the scalar register bits corresponding to each thread in a single warp in the scalar register to obtain a first intermediate reduction result:
[0017] Reduce the first intermediate reduction results of each warp in the scalar register to obtain the second reduction result.
[0018] Optionally, the reduction unit is specifically further configured to:
[0019] Allocate thread indices to each thread in a single warp in the scalar register according to a predetermined index allocation rule;
[0020] Reduce the scalar register bits corresponding to the threads with the same thread index in each of the thread bundles to obtain a second intermediate reduction result:
[0021] Reduce the second intermediate reduction results corresponding to the respective thread indices in the scalar register to obtain the second reduction result.
[0022] Optionally, the scheduling unit is further specifically configured to:
[0023] Enable the target thread bundle to have access rights to each of the scalar register areas in the scalar register, while other thread bundles among the multiple thread bundles only have access rights to the scalar register areas corresponding to the other thread bundles.
[0024] Optionally, the execution unit is multiple execution units in the computing unit;
[0025] The scheduling unit is further specifically configured to:
[0026] Execute a third reduction instruction through the target thread bundle of the target execution unit among the multiple execution units to reduce the second reduction results of the respective execution units to obtain a third reduction result.
[0027] Optionally, the scheduling unit is further specifically configured to:
[0028] Obtain the number of idle thread bundles in each execution unit;
[0029] Obtain the vector register capacity in the execution unit;
[0030] Based on the number of idle thread bundles and the vector register capacity, determine the target execution unit among the multiple execution units.
[0031] Optionally, the computing unit further has a shared memory, and the second reduction results of the respective execution units are stored in the shared memory. The scheduling unit is further specifically configured to:
[0032] Read the second reduction results of the respective execution units in the shared memory into the reduction unit of the target execution unit;
[0033] Reduce the second reduction results of the respective execution units in the reduction unit of the target execution unit to obtain the third reduction result.
[0034] Optionally, the computing unit further has a shared memory, and the second reduction results of the respective execution units are stored in the shared memory. The scheduling unit is further specifically configured to:
[0035] Using shared memory atomic operations, reduce the second reduction results of each of the execution units in the shared memory to obtain the third reduction result.
[0036] Optionally, the scheduling unit is further specifically configured to:
[0037] Through the reduction unit in each execution unit, obtain the first storage location of the second reduction result in the scalar register of the execution unit;
[0038] Read the second reduction result from the scalar register of each execution unit according to the first storage location, and write the read second reduction result into the shared memory.
[0039] Optionally, the shared memory has a first area and a second area, and the scheduling unit is further specifically configured to:
[0040] Through the reduction unit in each execution unit, write the first storage location of the second reduction result in the scalar register of the execution unit into the first area;
[0041] Read the first storage location from the first area, read the second reduction result from the scalar register of each execution unit according to the first storage location, and write the read second reduction result into the second area.
[0042] Optionally, the computing unit is multiple computing units in a computing component;
[0043] The scheduling unit is further specifically configured to:
[0044] Execute a fourth reduction instruction through the target warp of the target execution unit of the target computing unit among the multiple computing units to reduce the third reduction results of each computing unit to obtain a fourth reduction result.
[0045] Optionally, the scheduling unit is further specifically configured to:
[0046] Notify each computing unit to store the third reduction result of the computing unit in the shared memory of the computing unit, and obtain the shared memory identifier of the computing unit;
[0047] Search for the shared memory of the computing unit according to the shared memory identifier, obtain the third reduction results of each computing unit from the found shared memories of each computing unit, and reduce the third reduction results of each computing unit to obtain a fourth reduction result.
[0048] Optionally, the scheduling unit is further specifically configured to:
[0049] Read the third reduction results of the respective computing units into the reduction unit register area of the target execution unit of the target computing unit;
[0050] Reduce the third reduction results of the respective computing units in the reduction unit register area of the target execution unit of the target computing unit to obtain the fourth reduction result.
[0051] Optionally, the scheduling unit is further specifically configured to:
[0052] If each thread in a single warp has obtained the first reduction result, broadcast a synchronization message to other warps;
[0053] Executing the second reduction instruction through a target warp among the multiple warps corresponding to the execution unit includes: if the target warp among the multiple warps corresponding to the execution unit has received the synchronization messages of all other warps, execute the second reduction instruction.
[0054] Optionally, the scheduling unit is further specifically configured to:
[0055] Obtain the number of vectors to be reduced of the target reduction task, the maximum number of threads in a warp, and the maximum number of warps in the execution unit;
[0056] Calculate a first product of the maximum number of threads and the maximum number of warps;
[0057] If the number of vectors to be reduced is not greater than the first product, determine a first number based on the ratio of the number of vectors to be reduced to the maximum number of threads;
[0058] Set the first number of warps in the execution unit.
[0059] Optionally, the execution unit is multiple execution units in the computing unit;
[0060] The scheduling unit is further specifically configured to:
[0061] If the number of vectors to be reduced is greater than the first product, determine a second number based on the ratio of the number of vectors to be reduced to the first product;
[0062] If the second number is not greater than the number of execution units included in the computing unit, for the second number of execution units, set the multiple warps corresponding to each execution unit.
[0063] Optionally, the computing unit is multiple ones of the computing units in the computing component;
[0064] Specifically, the scheduling unit is further configured to:
[0065] If the second number is greater than the number of execution units included in the computing unit, determine a third number based on the ratio of the second number to the number of execution units;
[0066] If the third number is not greater than the number of computing units included in the computing component, for the third number of the computing units, set multiple warp corresponding to the execution units of each computing unit.
[0067] Optionally, the scheduling unit is further configured to:
[0068] Determine multiple vector register regions corresponding to the multiple warp set;
[0069] Successively store each vector to be reduced of the target reduction task into the vector register rows in each of the vector register regions.
[0070] Optionally, the scheduling unit is further configured to:
[0071] Obtain the remaining processing capacity of each warp;
[0072] Obtain the capacity of the vector register region corresponding to each warp;
[0073] Based on the remaining processing capacity and the capacity of the vector register region, determine the target warp among the multiple warp.
[0074] Optionally, the scheduling unit is further configured to:
[0075] Based on the remaining processing capacity, determine a first score of the warp;
[0076] Based on the capacity of the vector register region, determine a second score of the warp;
[0077] Based on the first score and the second score, determine the total score of the warp;
[0078] Based on the total scores of the respective warp, determine the target warp among the multiple warp.
[0079] In the embodiment of the present disclosure, a reduction unit, vector registers, and scalar registers are provided inside the execution unit. The vector registers include vector register regions corresponding to each warp, and each vector register region includes vector register rows corresponding to each thread in the warp. The scalar registers include scalar register regions corresponding to each warp, and each scalar register region includes scalar register bits corresponding to each thread in the warp. In the embodiment of the present disclosure, each vector to be reduced in the target reduction task is stored in each vector register row. When performing the reduction of each thread in the warp, a first reduction instruction is executed. For each thread in the warp, each vector element of the vector to be reduced in the vector register row corresponding to the thread is reduced to obtain a first reduction result, which is stored in the scalar register bit corresponding to the thread, thus realizing the reduction of threads within the warp. When performing the overall reduction inside the execution unit, a second reduction instruction is executed to reduce each scalar register bit in the scalar registers within the execution unit to obtain a second reduction result. The entire process does not require a shared memory outside the execution unit. By means of the vector registers and scalar registers provided inside the execution unit, the reduction of each thread within the warp and the overall reduction inside the execution unit are realized. Therefore, the entire task reduction process does not require frequent access to the shared memory, and the execution of the first reduction and the second reduction instruction does not require repeated iteration during reduction, improving the reduction efficiency.
[0080] Other features and advantages of the present disclosure will be described in the following specification, and in part, will be obvious from the specification, or can be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. They are used together with the embodiments of the present disclosure to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0082] Figure 1 is an architecture diagram of a system to which the reduction scheduling method according to an embodiment of the present disclosure is applied;
[0083] Figure 2 is an overall flowchart of a reduction scheduling method according to an embodiment of the present disclosure;
[0084] Figure 3 is a schematic diagram of a reduction scheduling method according to an embodiment of the present disclosure;
[0085] Figure 4A and Figure 4B is a schematic diagram of scalar registers and vector registers provided according to an embodiment of the present disclosure;
[0086] Figure 5 Yes Figure 2 A flowchart of step 210 in setting multiple warps corresponding to an execution unit;
[0087] Figure 6 Yes Figure 5 A schematic diagram of setting multiple warps corresponding to an execution unit;
[0088] Figure 7 A flowchart of setting multiple warps corresponding to an execution unit in the case where the execution unit is multiple execution units in a computing unit according to an embodiment of the present disclosure;
[0089] Figure 8 A flowchart of setting multiple warps corresponding to an execution unit in the case where the computing unit is multiple computing units in a computing component according to an embodiment of the present disclosure;
[0090] Figure 9 A schematic diagram of a computing component according to an embodiment of the present disclosure;
[0091] Figure 10 Yes Figure 2 A flowchart of step 220 in storing a vector to be reduced into each vector register row;
[0092] Figure 11 Yes Figure 2 A flowchart of step 240 in executing a second reduction instruction;
[0093] Figure 12 Yes Figure 2 Another flowchart of step 240 in executing a second reduction instruction;
[0094] Figure 13 A flowchart of selecting a target warp among multiple warps according to an embodiment of the present disclosure;
[0095] Figure 14 Yes Figure 13 A schematic diagram of selecting a target warp among multiple warps;
[0096] Figure 15 Yes Figure 13 A flowchart of step 1330 in determining a target warp among multiple warps;
[0097] Figure 16 A flowchart of setting access permissions according to an embodiment of the present disclosure;
[0098] Figure 17 A flowchart of executing a third reduction instruction according to an embodiment of the present disclosure;
[0099] Figure 18 Yes Figure 17 It is a flowchart of the third reduction instruction executed in step 1710 in
[0100] Figure 19 Yes Figure 17 It is another flowchart of the third reduction instruction executed in step 1710 in
[0101] Figure 20 It is a flowchart of storing the second reduction result into the shared memory according to an embodiment of the present disclosure.
[0102] Figure 21 Yes Figure 20 It is a schematic diagram of storing the second reduction result into the shared memory in
[0103] Figure 22 It is another flowchart of storing the second reduction result into the shared memory according to an embodiment of the present disclosure.
[0104] Figure 23 Yes Figure 22 It is a schematic diagram of storing the second reduction result into the shared memory in
[0105] Figure 24 It is a flowchart of selecting a target execution unit among multiple execution units according to an embodiment of the present disclosure.
[0106] Figure 25 Yes Figure 24 It is a schematic diagram of selecting a target execution unit among multiple execution units in
[0107] Figure 26 It is a flowchart of executing the fourth reduction order according to an embodiment of the present disclosure.
[0108] Figure 27 Yes Figure 26 It is a flowchart of the fourth reduction order executed in step 2610 in
[0109] Figure 28 Yes Figure 27 It is a schematic diagram of executing the fourth reduction order in
[0110] Figure 29 Yes Figure 27 It is a flowchart of obtaining the fourth reduction result in step 2720 in
[0111] Figure 30 It is a flowchart of synchronization of multiple threads in a warp according to an embodiment of the present disclosure.
[0112] Figure 31 It is a block diagram of a reduction scheduling device according to an embodiment of the present disclosure. Detailed implementation manners
[0113] To make the objectives, technical solutions and advantages of the present disclosure more comprehensible, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present disclosure and are not intended to limit the present disclosure.
[0114] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:
[0115] Warp: A warp is the basic execution unit in a Streaming Multiprocessor (SM). When a grid is launched (launching a grid is equivalent to launching a kernel, and each kernel corresponds to its own grid), the grid contains thread blocks. After the thread blocks are assigned to a certain SM, they are divided into multiple warps. In a warp, all threads execute in a single-instruction multiple-thread manner, executing the same instruction at each step, but the data processed is private data, that is, the data corresponds to the thread.
[0116] Warp scheduling: It refers to the way of parallel execution of multiple threads or instructions, combining them into a thread bundle for scheduling and execution. In this scheduling method, the central processing unit will start several threads or instructions simultaneously and improve the execution efficiency through pipeline technology.
[0117] Execution Unit (EU): It is the execution unit in a microprocessor, which is responsible for the execution of instructions, performing arithmetic operations, logical operations, shift operations and other computing tasks. In fact, it has both the functions of a controller and an arithmetic unit.
[0118] Shared Memory: It refers to the memory of a certain capacity size in a multi-processor computer system that can be accessed by different processors. Since multiple processors need to access the memory quickly, the memory needs to be cached. After any cached data is updated, since other processors may also need to access it, the shared memory needs to be updated immediately, otherwise different processors may use different data.
[0119] Reduction refers to performing a certain operation on multiple input data, reducing multiple input data to a single result. It usually reduces a vector containing multiple vector elements to a single value. This operation can be taking the minimum, maximum, summation, sum of squares, logical AND, logical OR, vector dot product, etc. Computers usually perform reduction operations using software. Performing reduction operations using software is not conducive to improving operation efficiency, cannot improve computing power, and does not meet the development needs of the large language model era. Therefore, a method of using a hardware chip for reduction is proposed.
[0120] The usual method of using a hardware chip for reduction loads each vector to be reduced in the target task into the shared memory shared by multiple execution units (EUs) in the chip. Then, a tree reduction or interleaved reduction is used in the shared memory to obtain the reduction result of the target task. Therefore, the entire task reduction process requires frequent access to the shared memory and multiple iterations in the shared memory, resulting in low reduction efficiency.
[0121] Based on this, the embodiments of the present disclosure provide a reduction scheduling method and apparatus, which can improve the efficiency of reduction operations.
[0122] The system architecture applied in the embodiments of the present disclosure
[0123] Figure 1 is the system architecture diagram applied to the reduction scheduling method according to the embodiments of the present disclosure. It is mainly a computing component, and the computing component includes a scheduling unit for scheduling thread blocks and multiple computing units.
[0124] The computing component is a computer processing device with a certain computing ability, capable of performing reduction scheduling for a given task to obtain a reduction result. The computing component can be a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a neural processing unit (NPU), a tensor processing unit (TPU), etc. When the given task is a graphics processing related task, such as image recognition, animation rendering, etc., the computing component is usually set to a GPU. And when the given task is other extensive computing tasks, such as running an operating system, the computing component is usually set to a CPU.
[0125] The computing component includes multiple computing units, and the scheduling unit refers to a device that schedules thread blocks corresponding to the computing component to multiple computing units for reduction processing through the thread blocks at the computing units. A computing unit is a unit for measuring computing resources in a computer system, usually including a computing core and memory. Among them, the computing core is determined based on the computing component. When the computing component is a CPU, the computing core is a CPU core; when the computing component is a GPU, the computing core is a GPU core. The computing unit can support the requirements of various computing tasks.
[0126] In addition, the computing unit includes a scheduling unit for scheduling warps and multiple execution units. The scheduling unit can schedule multiple warps corresponding to the computing unit to multiple execution units for reduction processing. The execution unit is an execution unit in a microprocessor, which is responsible for the execution of instructions and performs computing tasks such as arithmetic operations, logical operations, and shift operations. A shared memory is also provided in the computing unit, and the shared memory can be accessed by different execution units and has a large-capacity memory.
[0127] The reduction scheduling method provided by the embodiments of the present disclosure is applied to Figure 1 the system architecture shown, which can improve the efficiency of the reduction operation.
[0128] General description of the embodiments of the present disclosure
[0129] The reduction scheduling method refers to the process of performing reduction scheduling on a target reduction task on computing components such as CPUs and GPUs to obtain a reduction result. This reduction scheduling method can improve the reduction efficiency.
[0130] The reduction scheduling method provided by the embodiments of the present disclosure is used for the scheduling unit, and the scheduling unit is used to schedule the execution unit. The execution unit includes a reduction unit, vector registers, and scalar registers. Among them, the vector registers include vector register areas corresponding to each warp, and the vector register areas include vector register rows corresponding to each thread in the warp; the scalar registers include scalar register areas corresponding to each warp, and the scalar register areas include scalar register bits corresponding to each thread in the warp. Referring to Figure 2 , the reduction scheduling method provided by the embodiments of the present disclosure includes:
[0131] Step 210: Set multiple warps corresponding to the execution unit for the target reduction task;
[0132] Step 220: Store each vector to be reduced of the target reduction task into each vector register row;
[0133] Step 230: Through each warp, execute a first reduction instruction to, for each thread in the warp, use a reduction unit to reduce each vector element of the vector to be reduced in the vector register row corresponding to the thread, obtaining a first reduction result and storing it in the scalar register bit corresponding to the thread;
[0134] Step 240: Through the target warp among the multiple warps corresponding to the execution unit, execute a second reduction instruction to, using the reduction unit, reduce each scalar register bit in the scalar register within the execution unit, obtaining a second reduction result.
[0135] It should be noted that referring to Figure 3 , the execution unit includes a reduction unit, vector registers, and scalar registers. Among them, the reduction unit refers to a unit that receives a reduction instruction and executes an operation corresponding to the reduction instruction. The reduction unit can perform reduction calculations on the data in the vector registers and scalar registers. The vector register is a component for temporarily storing vectors, and the scalar register is a component for temporarily storing scalars. The vector register and scalar register are components for temporarily storing data during the execution of the target reduction task.
[0136] The vector register includes multiple vector register regions, and each vector register region corresponds to a warp one by one. Referring to Figure 4A , the vector register includes a total of M vector register regions, and the M vector register regions correspond to warp 1, warp 2,..., warp M respectively. In the vector register region, multiple vector register rows corresponding to each thread in the warp are set. Each vector register row is used to store a vector, and each bit in the vector register row corresponds to an element of the vector corresponding to the vector register row. For example, Figure 4B the first row of the vector register corresponding to warp 1 stores vector A1, and the first position in the first row stores the first element A11 of vector A1. The vector register row corresponds to each thread in the warp. For example, Figure 4B warp 1 in
[0137] includes N threads, and each thread corresponds to a vector register row. Figure 4A The scalar register shown includes a total of M scalar register regions. Warp 1 is correspondingly set with scalar register region 1, and warp M is correspondingly set with scalar register region M. The scalar register region includes multiple scalar register bits, and the scalar register bits correspond to the threads. For example, Figure 4B S11 in
[0138] The following provides a detailed description of steps 210 to 240.
[0139] In step 210, for a target reduction task, multiple warps corresponding to the execution unit are set.
[0140] The target reduction task refers to a task that requires reduction calculation. The target reduction task can be set as needed. Additionally, for a computing component, different computing tasks correspond to different target reduction tasks, and the same computing task may include multiple reduction tasks.
[0141] Suppose the computing component needs to calculate the normalized exponential function, i.e., the softmax function. The purpose of Softmax is to present the results of multi-classification in the form of probabilities. It maps the outputs of multiple neurons to the interval (0, 1) for multi-classification. Assume we have an array X, and x i represents the i-th element in X, then the softmax value of this element is:
[0142]
[0143] where n is the number of elements in the array X. During the calculation of the softmax function, the value of n is usually large, so the calculation amount of the softmax function is relatively large. Therefore, that is, the sum of the exponential function values of multiple elements in the array is used as the target reduction task.
[0144] Based on the target reduction task, multiple warps corresponding to the execution unit are set. The number of warps is determined based on the calculation amount of the target reduction task, or rather, the size of the data. The more data that needs to be processed in the target reduction task, the larger the number of warps set.
[0145] In step 220, each vector to be reduced in the target reduction task is stored in each vector register row.
[0146] The vector to be reduced refers to the vector that needs to be reduced in the target reduction task. The vector to be reduced corresponds to the vector register row. For Figure 4B the vector register shown, the vector to be reduced A1 is stored in the first vector register row in vector register 1, and each bit of the vector register row stores an element of the vector to be reduced.
[0147] For the target reduction task can be to divided into multiple vectors to be reduced and stored in each vector register row. Among them, etc. are the vector elements in their corresponding vectors to be reduced.
[0148] In step 230, through each warp, a first reduction instruction is executed to reduce each vector element of the vector to be reduced in the vector register row corresponding to the thread in the warp by using a reduction unit, so as to obtain a first reduction result, which is stored in the scalar register bit corresponding to the thread.
[0149] The first reduction instruction refers to an instruction for performing reduction calculation on multiple vectors to be reduced in a vector register. The scheduling unit can execute the first reduction instruction through a warp. Specifically, the scheduling unit sends the first reduction instruction to the reduction unit through the warp. After receiving the first reduction instruction, the reduction unit determines the warp corresponding to the first reduction instruction, and reduces each vector element of the vector to be reduced in the vector register row corresponding to each thread in the warp to obtain a first reduction result.
[0150] It should be noted that the reduction can be set as needed, and it can be taking the minimum, maximum, summation, sum of squares, logical AND, logical OR, vector dot product, etc.
[0151] In the embodiment of the present disclosure, a reduction unit is used to reduce each vector element in the vector to be reduced, and the reduction unit can obtain the first reduction result of the vector to be reduced without iterating multiple times.
[0152] Referring to Figure 4A , since the vectors to be reduced are stored in vector register banks 1 to M, the first reduction instruction needs to be executed through warps 1 to M. Among them, the number of threads in warp 2 is N, and the multiple threads in warp 2 respectively correspond to vectors to be reduced B1, vectors to be reduced B2,..., vectors to be reduced BN. For warp 2, the reduction unit is used to reduce each vector element of the vectors to be reduced B1, vectors to be reduced B2,..., vectors to be reduced BN respectively to obtain a first reduction result. Among them, the first reduction result corresponding to the vector to be reduced B1 is S21, the first reduction result corresponding to the vector to be reduced B2 is S22, and the first reduction result corresponding to the vector to be reduced BN is S2N.
[0153] After obtaining the first reduction result, the first reduction result is stored in the scalar register bit corresponding to the thread. Since Figure 4B the vector to be reduced A1 in the vector register bank 1 in corresponds to the scalar register bit in the first row and the first column in the scalar register bank 1 for the same thread, the first reduction result S11 corresponding to the vector to be reduced A1 is stored in this scalar register bit.
[0154] In step 240, the target warp among multiple warps corresponding to the execution unit executes a second reduction instruction to utilize the reduction unit to reduce each scalar register bit in the scalar register within the execution unit, obtaining a second reduction result.
[0155] The target warp is one of the multiple warps corresponding to the execution unit. Referring to Figure 4A , the target warp can be any one of warp 1 to warp M.
[0156] The second reduction instruction is an instruction for reducing multiple first reduction results in the scalar register. The second reduction instruction is executed through the target warp. After receiving the second reduction instruction, the reduction unit reduces the first reduction results corresponding to each scalar register bit in the scalar register within the execution unit, obtaining a second reduction result.
[0157] Referring to Figure 4B , in the scalar register, each scalar register bit in scalar register area 1 stores first reduction results S11 to S1N, each scalar register bit in scalar register area 2 stores first reduction results S21 to S2N, and each scalar register bit in scalar register area M stores first reduction results SM1 to SMN. The first reduction results corresponding to each scalar register bit in the scalar register are reduced to obtain a second reduction result Final.wp.
[0158] It should be noted that after obtaining the second reduction result, the second reduction result can be stored in the scalar register area corresponding to the target warp.
[0159] Referring to Figure 3 , the scheduling unit can execute the first reduction instruction through each warp. During the execution of the first reduction instruction, the reduction unit reduces each vector element of the vector to be reduced in the vector register row corresponding to each thread in the warp, obtaining a first reduction result, and stores the obtained first reduction result in the scalar register bit corresponding to the thread. Then, the scheduling unit determines the target warp among multiple warps and executes the second reduction instruction through the target warp to utilize the reduction unit to reduce the first reduction results corresponding to each scalar register bit in the scalar register, obtaining a second reduction result.
[0160] In the embodiments of the above steps 210 to 240, a reduction unit, vector registers, and scalar registers are provided inside the execution unit. The vector registers include vector register regions corresponding to each warp, and each vector register region includes vector register rows corresponding to each thread in the warp. The scalar registers include scalar register regions corresponding to each warp, and each scalar register region includes scalar register bits corresponding to each thread in the warp. In the embodiments of the present disclosure, each vector to be reduced in the target reduction task is stored in each vector register row. When performing the reduction of each thread in the warp, a first reduction instruction is executed. For each thread in the warp, each vector element of the vector to be reduced in the vector register row corresponding to the thread is reduced to obtain a first reduction result, which is stored in the scalar register bit corresponding to the thread, thereby implementing the reduction of the threads within the warp. When performing the overall reduction within the execution unit, a second reduction instruction is executed to reduce each scalar register bit in the scalar registers within the execution unit to obtain a second reduction result. The entire process does not require a shared memory outside the execution unit. By means of the vector registers and scalar registers provided within the execution unit, the reduction of each thread within the warp and the overall reduction within the execution unit are achieved. Therefore, the entire task reduction process does not require frequent access to the shared memory, and the execution of the first reduction and the second reduction instruction does not require repeated iteration during reduction, thereby improving the reduction efficiency.
[0161] The above is the overall description of steps 210 to 240. Since step 230 has been described in sufficient detail above, only the specific implementation processes of steps 210, 220, and 240 will be described in detail below.
[0162] Detailed description of step 210
[0163] In step 210, for the target reduction task, a plurality of warps corresponding to the execution unit are set.
[0164] In one embodiment, referring to Figure 5 , step 210 includes:
[0165] Step 510, obtaining the number of vectors to be reduced in the target reduction task, the maximum number of threads in a warp, and the maximum number of warps in the execution unit;
[0166] Step 520, calculating a first product of the maximum number of threads and the maximum number of warps;
[0167] Step 530, if the number of vectors to be reduced is not greater than the first product, determining a first number based on the ratio of the number of vectors to be reduced to the maximum number of threads;
[0168] Step 540, setting a first number of warps in the execution unit.
[0169] The following will describe steps 510 to 540 in detail.
[0170] In step 510, obtain the number of vectors to be reduced for the target reduction task, the maximum number of threads in a warp, and the maximum number of warps in an execution unit.
[0171] The number of vectors to be reduced refers to the number of vectors to be reduced corresponding to the target reduction task.
[0172] The maximum number of threads refers to the maximum value of the number of threads that a warp can accommodate. A warp usually includes 32 consecutive threads, so the maximum number of threads in a warp is 32.
[0173] The maximum number of warps refers to the maximum number of warps corresponding to the execution unit. For Figure 4A the shown execution unit, the corresponding maximum number of warps is M.
[0174] In step 520, calculate the first product of the maximum number of threads and the maximum number of warps.
[0175] The first product is the product of the maximum number of threads and the maximum number of warps, which can also be regarded as the maximum value of the number of threads that the execution unit can accommodate. For example, for Figure 4A the shown execution unit, the corresponding first product is MN.
[0176] In step 530, if the number of vectors to be reduced is not greater than the first product, determine the first number based on the ratio of the number of vectors to be reduced to the maximum number of threads.
[0177] The first number refers to the number of warps in the execution unit. If the number of vectors to be reduced is not greater than the first product, then the vector registers in a single execution unit are sufficient to place multiple vectors to be reduced corresponding to the target reduction task. At this time, calculate the ratio of the number of vectors to be reduced to the maximum number of threads, and round up based on this ratio to obtain the first number. When the number of vectors to be reduced is not greater than the first product, if the number of vectors to be reduced is MN - 1 and the maximum number of threads is N, calculate the ratio of the number of vectors to be reduced to the maximum number of threads (MN - 1) / N, and round up to obtain the first number as M.
[0178] In step 540, set the first number of warps in the execution unit.
[0179] After determining the value of the first number, set the first number of warps in the execution unit, and correspond the warps to the vector register area and the scalar register area one by one.
[0180] Refer to Figure 6, after determining the target reduction task, obtain the number of vectors to be reduced, the maximum number of threads, and the maximum number of warps. Take the product of the maximum number of threads and the maximum number of warps as the first product. When the number of vectors to be reduced is not greater than the first product, determine the first number based on the ratio of the number of vectors to be reduced to the maximum number of threads, so as to set the first number of warps in the execution unit.
[0181] In the embodiments of steps 510 to 540 above, when the number of vectors to be reduced is not greater than the first product, the first number of warps is set in the execution unit to execute the target reduction task using a single execution unit, maximizing the computing power of the execution unit, reducing resource waste, and improving the reduction efficiency.
[0182] In one embodiment, the execution unit is multiple execution units in the computing unit. Refer to Figure 7 , after step 520, the reduction scheduling method provided by the embodiments of the present disclosure further includes:
[0183] Step 710: If the number of vectors to be reduced is greater than the first product, determine the second number based on the ratio of the number of vectors to be reduced to the first product;
[0184] Step 720: If the second number is not greater than the number of execution units included in the computing unit, set multiple warps corresponding to each of the second number of execution units.
[0185] It should be noted that the computing unit is a unit for measuring computing resources in a computer system. Refer to Figure 3 , and multiple execution units are usually set in the computing unit.
[0186] The following describes steps 710 and 720 in detail.
[0187] In step 710, if the number of vectors to be reduced is greater than the first product, determine the second number based on the ratio of the number of vectors to be reduced to the first product.
[0188] If the number of vectors to be reduced is greater than the first product, then a single execution unit cannot accommodate the multiple vectors to be reduced corresponding to the target reduction task. At this time, calculate the ratio of the number of vectors to be reduced to the first product and round up to determine the second number. The second number can be regarded as the number of execution units that need to set warps corresponding to the target reduction task. Suppose the number of vectors to be reduced is 3MN - 1 and the first product is MN, then the second number is 3. If the number of vectors to be reduced is 3MN + 1 and the first product is MN, then the second number is 4.
[0189] In step 720, if the second number is not greater than the number of execution units included in the computing unit, set multiple warps corresponding to each of the second number of execution units.
[0190] If the second number is not greater than the number of execution units included in the computing unit, then the execution units included in a single computing unit are sufficient to accommodate multiple vectors to be reduced corresponding to the target reduction task. At this time, for the second number of execution units in the computing unit, a plurality of warps corresponding to each execution unit are set. Among them, the number of warps in the second number minus 1 execution units is the maximum number of warps, and the number of warps in the other execution unit is determined based on the number of vectors to be reduced, the first product, the second number, and the maximum number of warps. Assume that the number of vectors to be reduced is 3MN - 1, the first product is MN, and the second number is 3. Then the number of warps in two of the execution units is M, the number of vectors to be reduced corresponding to the other execution unit is MN - 1, calculate the ratio of MN - 1 to the maximum number of threads N and round up, and M warps are set in this execution unit.
[0191] Refer to Figure 6 , in the case where the number of vectors to be reduced is greater than the first product, determine the second number based on the ratio of the number of vectors to be reduced to the first product. Then, in the case where the second number is not greater than the number of execution units included in the computing unit, for the second number of execution units, set a plurality of warps corresponding to each execution unit.
[0192] In the embodiments of the above steps 710 and 720, when the number of vectors to be reduced is greater than the first product and the second number is not greater than the number of execution units included in the computing unit, for the second number of execution units, set a plurality of warps corresponding to each execution unit. The embodiments of the present disclosure can execute the target reduction task using as few execution units as possible, maximize the computing power of the execution units, reduce resource waste, and improve the reduction efficiency.
[0193] In one embodiment, the computing unit is a plurality of computing units in a computing component. Refer to Figure 8 , after step 710, the reduction scheduling method provided by the embodiments of the present disclosure further includes:
[0194] Step 810: If the second number is greater than the number of execution units included in the computing unit, determine a third number based on the ratio of the second number to the number of execution units;
[0195] Step 820: If the third number is not greater than the number of computing units included in the computing component, for the third number of computing units, set a plurality of warps corresponding to the execution units of each computing unit.
[0196] It should be noted that the computing component is a computer processing device, which has a certain computing power and can perform reduction scheduling for a given task to obtain a reduction result. The computing component can be a CPU, a GPU, etc. Refer to Figure 1 andFigure 9 , the computing component includes multiple computing units, and each computing unit includes multiple execution units.
[0197] The following will describe steps 810 and 820 in detail.
[0198] In step 810, if the second number is greater than the number of execution units included in the computing unit, based on the ratio of the second number to the number of execution units, determine the third number.
[0199] If the second number is greater than the number of execution units included in the computing unit, then the multiple execution units corresponding to a single computing unit cannot accommodate the multiple vectors to be reduced corresponding to the target reduction task. Therefore, calculate the ratio of the second number to the number of execution units, and round up to determine the third number. For example, assume that the number of execution units included in the computing unit is 4, the number of vectors to be reduced is (4M + 1)N, the first product is MN, then the second number is 5. At this time, the second number is greater than the number of execution units included in the computing unit. Calculate the ratio of the second number to the number of execution units, and round up to determine the third number as 2.
[0200] In step 820, if the third number is not greater than the number of computing units included in the computing component, for the third number of computing units, set multiple warps corresponding to the execution units of each computing unit.
[0201] If the third number is not greater than the number of computing units included in the computing component, then the multiple execution units corresponding to a single computing component are sufficient to accommodate the multiple vectors to be reduced corresponding to the target reduction task. Therefore, for the third number of computing units, set multiple warps corresponding to the execution units of each computing unit. For example, the third number is 2, the number of computing units included in the computing component is 4, the number of vectors to be reduced is (4M + 1)N, the first product is MN. Then, the execution units of one of the computing units are set with the maximum number of warps, and the execution units of the other computing unit accommodate N vectors to be reduced. Then, only one warp is set in one of the execution units of that computing unit.
[0202] Refer to Figure 6 , when the second number is greater than the number of execution units included in the computing unit, based on the ratio of the second number to the number of execution units, determine the third number. After that, if the third number is not greater than the number of computing units included in the computing component, for the third number of computing units, set multiple warps corresponding to the execution units of each computing unit.
[0203] It should be noted that when the third number is greater than the number of computing units included in the computing component, if the computer is a single-core, that is, only contains one computing component, then after waiting for the currently allocated warp to process multiple vectors to be reduced, corresponding warps can be allocated for the remaining vectors to be reduced. If the computer is multi-core, that is, contains multiple computing components, the fourth number can be determined based on the ratio of the third number to the number of computing components included in the computer, and multiple warps corresponding to the execution units of the computing units in the fourth number of computing components can be set.
[0204] In the embodiments of the above steps 810 and 820, when the second number is greater than the number of execution units included in the computing unit and the third number is not greater than the number of computing units included in the computing component, multiple warps corresponding to the execution units of the third number of computing units are set. The embodiments of the present disclosure can execute the target reduction task using as few computing units as possible, maximize the computing power of the computing units, reduce resource waste, and improve the reduction efficiency.
[0205] Detailed description of step 220
[0206] In step 220, each vector to be reduced of the target reduction task is stored in each vector register row.
[0207] In one embodiment, referring to Figure 10 , step 220 includes:
[0208] Step 1010, determining multiple vector register regions corresponding to the multiple warps set;
[0209] Step 1020, sequentially storing each vector to be reduced of the target reduction task in the vector register rows in each vector register region.
[0210] The following provides a detailed description of step 1010 and step 1020.
[0211] In step 1010, multiple vector register regions corresponding to the multiple warps set are determined.
[0212] After multiple warps are set and allocated to the execution units, the warps are put into one-to-one correspondence with the vector register regions.
[0213] In step 1020, each vector to be reduced of the target reduction task is sequentially stored in the vector register rows in each vector register region.
[0214] Since the threads in the warp are in one-to-one correspondence with the vectors to be reduced, after determining the vector register region corresponding to the warp, the vectors to be reduced corresponding to each thread in the warp are sequentially stored in the vector register rows in the vector register region.
[0215] Refer to Figure 4A and Figure 4B Assume that for the target reduction task to be reduced, there are thread bundles 1, 2, 3, …, M set up correspondingly, and the thread bundles are in one-to-one correspondence with the vector register area in the execution unit. For thread bundle 1, it is determined that the vectors to be reduced corresponding to thread bundle 1 are A1, A2, …, AN, and it corresponds to vector register area 1. Therefore, the vectors to be reduced A1, A2, …, AN are sequentially placed into the vector register rows in vector register area 1.
[0216] The embodiments of the above step 1010 and step 1020 first determine multiple vector register areas corresponding to multiple thread bundles, and then sequentially store each vector to be reduced in the target reduction task into the vector register rows in each vector register area, so that multiple vectors to be reduced are placed into the vector registers, facilitating the reduction calculation of multiple vectors to be reduced through the vector registers. Moreover, it divides a single target reduction task into the reduction of multiple vectors to be reduced, which can improve the reduction efficiency.
[0217] Detailed description of step 240
[0218] In step 240, through the target thread bundle among the multiple thread bundles corresponding to the execution unit, the second reduction instruction is executed to use the reduction unit to reduce each scalar register bit in the scalar register in the execution unit to obtain the second reduction result.
[0219] In one embodiment, refer to Figure 11 , step 240 includes:
[0220] Step 1110: Reduce each scalar register bit corresponding to each thread in a single thread bundle in the scalar register to obtain a first intermediate reduction result;
[0221] Step 1120: Reduce the first intermediate reduction results of each thread bundle in the scalar register to obtain the second reduction result.
[0222] The following gives a detailed description of step 1110 and step 1120.
[0223] In step 1110, each scalar register bit corresponding to each thread in a single thread bundle in the scalar register is reduced to obtain a first intermediate reduction result.
[0224] Since the scalar register area and the vector register area are in one-to-one correspondence, each scalar register bit in the scalar register area stores the first reduction result, which is the reduction result of the vector to be reduced in the vector register row corresponding to its corresponding thread.
[0225] Embodiments of the present disclosure can perform reduction on scalar register bits corresponding to each thread in a single warp in a single scalar register area to obtain a first intermediate reduction result. For example, referring to Figure 4B , S11 is the first reduction result corresponding to the vector A1 to be reduced, S1 N is the first reduction result corresponding to the vector AN to be reduced. The reduction is performed on S11 to S1 N in the scalar register area 1 to obtain the first intermediate reduction result X1. In addition, X2 is the first intermediate reduction result corresponding to the scalar register area 2, and XM is the first intermediate reduction result corresponding to the scalar register area M.
[0226] In step 1120, the first intermediate reduction results of each warp in the scalar register are reduced to obtain a second reduction result.
[0227] After determining the first intermediate reduction results of each warp, the first intermediate reduction results of each warp are reduced to obtain a second reduction result.
[0228] Referring to Figure 4B , for each warp, that is, the first intermediate reduction results X1, X2, X3,..., XM corresponding to each scalar register area are reduced to obtain the second reduction result XM.
[0229] In the embodiments of the above steps 1110 and 1120, first, the first intermediate reduction results of the scalar register bits corresponding to each thread in each warp are calculated, and then the first intermediate reduction results of each warp are reduced to obtain the second reduction result. When the number of warps in the execution unit is less than the number of threads in a warp, the method only needs to perform the reduction for the number of warps in the execution unit times to obtain the first intermediate reduction result, and the number of times of reduction required is significantly reduced. The first reduction result can be quickly reduced to obtain the second reduction result, improving the reduction efficiency.
[0230] In another embodiment, referring to Figure 12 , step 240 includes:
[0231] Step 1210: Assign thread indices to each thread in a single warp in the scalar register according to a predetermined index assignment rule;
[0232] Step 1220: Reduce the scalar register bits corresponding to the threads with the same thread index in each warp to obtain a second intermediate reduction result;
[0233] Step 1230: Reduce the second intermediate reduction results corresponding to each thread index in the scalar register to obtain a second reduction result.
[0234] The following will describe steps 1210 to 1230 in detail.
[0235] In step 1210, thread indices are assigned to each thread in a warp in the scalar register according to a predetermined index assignment rule.
[0236] A thread index refers to an index flag that distinguishes each thread in a single warp. If the index assignment rules for allocating thread indices in multiple warps are the same, then in the scalar register, the number of threads with the same thread index is equal to the number of warps.
[0237] The predetermined index assignment rule can be set as needed. For example, rules such as increment, decrement, square, etc. It only needs to be ensured that the thread indices corresponding to each thread in the same warp are different, and the number of threads with the same thread index is equal to the number of warps. Referring to Figure 4A , in the embodiment of the present disclosure, the thread index of the thread corresponding to S11 is set to 1, the thread index of the thread corresponding to S12 is set to 2, and the thread index of the thread corresponding to S1N is set to N. Then for warp M, the thread index corresponding to SM1 is 1, the thread index corresponding to SM2 is 2, and the thread index corresponding to SMN is N.
[0238] In step 1220, the scalar register bits corresponding to the threads with the same thread index in each warp are reduced to obtain a second intermediate reduction result.
[0239] For the threads with the same thread index in each warp, the first reduction results in their corresponding scalar register bits are reduced to obtain a second intermediate reduction result. For example, referring to Figure 4A , the thread indices of the threads corresponding to S11, S21, S31, ……, SM1 are all 1. Therefore, S11, S21, S31, ……, SM1 are reduced to obtain a second intermediate reduction result.
[0240] In step 1230, the second intermediate reduction results corresponding to each thread index in the scalar register are reduced to obtain a second reduction result.
[0241] After determining the second intermediate reduction results corresponding to each thread index, the multiple second intermediate reduction results are reduced to obtain a second reduction result. After determining the second intermediate reduction results corresponding to thread indices 1 to N, the second intermediate reduction results corresponding to thread indices 1 to N are reduced to obtain a second reduction result.
[0242] In the above steps 1120 to 1130, thread indices are assigned to each thread in a warp. After that, the scalar register bits with the same thread index are reduced to obtain a second intermediate reduction result. After determining the second intermediate reduction results corresponding to each thread index, multiple second intermediate reduction results are reduced to obtain a second reduction result. When the number of warps in the execution unit is greater than the number of threads in a warp, the method only needs to perform the reduction for the number of threads in the warp during the calculation of the second intermediate reduction result. Then, the number of reductions required to reduce the first reduction result to obtain the second reduction result is significantly reduced, further improving the reduction efficiency.
[0243] The above has explained steps 210 to 240 in detail. Below, some specific points or extended contents involved will be described in detail by topic. These topics include the determination and permission setting of the target warp, the scheduling and execution of the third reduction instruction, the scheduling and execution of the fourth reduction instruction, the synchronization of multiple threads in a warp, etc.
[0244] Determination and Permission Setting of Target Warp
[0245] In the embodiments of the present disclosure, the second reduction instruction is executed by the target warp among multiple warps corresponding to the execution unit. Among them, the target warp can be preset and can be any one of the multiple warps corresponding to the execution unit. For example Figure 4A warp 1, warp 2, warp 3, or warp M in
[0246] In addition, in order to improve the execution efficiency of the second reduction instruction, the target warp can also be set according to the performance of each warp. In this case,
[0247] In one embodiment, referring to Figure 13 , the method for selecting the target warp among multiple warps includes:
[0248] Step 1310, obtain the remaining processing capabilities of each warp;
[0249] Step 1320, obtain the capacity of the vector register area corresponding to each warp;
[0250] Step 1330, determine the target warp among multiple warps based on the remaining processing capabilities and the capacity of the vector register area.
[0251] The following will describe steps 1310 to 1330 in detail.
[0252] In step 1310, obtain the remaining processing capabilities of each warp.
[0253] The remaining processing capacity refers to how much data a warp can process, which can be measured based on the number of idle threads in the warp. The more idle warps there are in the warp, the stronger the remaining processing capacity of the warp.
[0254] In step 1320, obtain the capacity of the vector register area corresponding to each warp.
[0255] The capacity of the vector register area refers to how much data the vector register area can hold. The larger the capacity of the vector register area, the more data the vector register area can store, and thus the more data the warp can process.
[0256] In step 1330, based on the remaining processing capacity and the capacity of the vector register area, determine the target warp among multiple warps.
[0257] The target warp is determined among multiple warps based on the remaining processing capacity and the capacity of the vector register area. The stronger the remaining processing capacity of the warp and the larger the capacity of the register area, the more data the warp can process and the faster the processing speed, and the better the performance of the warp. Therefore, the embodiments of the present disclosure select the warp with more excellent performance among multiple warps based on the remaining processing capacity and the capacity of the register area, and use it as the target warp.
[0258] Refer to Figure 14 , the embodiments of the present disclosure determine the performance of the warp based on the remaining processing capacity of the warp and the capacity of the register area of its corresponding vector register area. After determining the remaining processing capacity and the capacity of the register area of multiple warps, compare the multiple warps to determine the target warp.
[0259] The above embodiments of steps 1310 to 1330 determine the target warp among multiple warps based on the remaining processing capacity and the capacity of the vector register area, which can improve the processing efficiency of the target warp for the first reduction result, and thus improve the efficiency of reduction scheduling.
[0260] The above is a detailed description of steps 1310 to 1330. Below, a detailed description of the specific implementation process of step 1330 will be given.
[0261] In step 1330, based on the remaining processing capacity and the capacity of the vector register area, determine the target warp among multiple warps.
[0262] In one embodiment, refer to Figure 15 , step 1330 includes:
[0263] Step 1510, determine the first score of the warp based on the remaining processing capacity;
[0264] Step 1520: Determine the second score of the warp based on the capacity of the vector register area;
[0265] Step 1530: Determine the total score of the warp based on the first score and the second score;
[0266] Step 1540: Determine the target warp among multiple warps based on the total scores of each warp.
[0267] The following provides a detailed description of Step 1510 and Step 1540.
[0268] In Step 1510, determine the first score of the warp based on the remaining processing capacity.
[0269] The first score is determined based on the remaining processing capacity. Refer to Figure 14 , the maximum processing capacity of warp X is 1000, and the remaining processing capacity is 840. Thus, its first score is determined to be 84.
[0270] In Step 1520, determine the second score of the warp based on the capacity of the vector register area.
[0271] The second score is determined based on the capacity of the vector register area. Refer to Figure 14 , the maximum capacity of the vector register area corresponding to warp X is 256 bytes, and the current register area capacity is 224 bytes. Therefore, the second score of the warp is 87.5.
[0272] In Step 1530, determine the total score of the warp based on the first score and the second score.
[0273] The total score is determined based on the first score and the second score. To ensure the accuracy of the total score, in the embodiments of the present disclosure, the first score and the second score are weighted and summed to obtain the total score of the warp. Specifically, set the first weight and the second weight, where the first weight corresponds to the first score and the second weight corresponds to the second score. Then, calculate the product of the first weight and the first score, the product of the second weight and the second score, and add the two obtained products to get the total score of the warp. For example, refer to 14, the first score of the warp is 84, the second score is 87.5, the first weight is set to 0.6, and the second weight is set to 0.4. Then, the total score of the warp can be obtained as 85.4.
[0274] In Step 1540, determine the target warp among multiple warps based on the total scores of each warp.
[0275] Since the stronger the remaining processing power of a warp and the larger the capacity of the vector register area, the better the performance of the warp, the higher the total score of the warp, the better the performance of the warp. Therefore, select the warp with the highest total score among multiple warps as the target warp.
[0276] The embodiments of the above steps 1510 to 1540 determine the first score of a warp based on the remaining processing power and determine the second score of the warp based on the capacity of the vector register area. Then, based on the first score and the second score, determine the total score of the warp to select the target warp through the total score. The determination of the total score quantifies the selection process of the target warp, making the performance of the obtained target warp better than that of other warps, thereby improving the reduction efficiency of the second reduction instruction.
[0277] In one embodiment, referring to Figure 16 , before step 240, the reduction scheduling method provided by the embodiments of the present disclosure further includes:
[0278] Step 1610, enable the target warp to have access rights to each scalar register area in the scalar register, while other warps among multiple warps only have access rights to the scalar register areas corresponding to the other warps.
[0279] It should be noted that the embodiments of the present disclosure execute the second reduction instruction through the target warp. During the process of the second reduction instruction, the reduction unit reduces each scalar register bit in the scalar register in the execution unit. To ensure that the second reduction instruction can be executed through the target warp, it is necessary to enable the target warp to have access rights to each scalar register area in the scalar register. In addition, since other warps except the target warp do not need to execute the second reduction instruction, to reduce errors, the access rights of other warps among multiple warps should only be set to the access rights to the scalar register areas corresponding to them.
[0280] Scheduling and execution of the third reduction instruction
[0281] In one embodiment, the execution unit is multiple execution units in the computing unit. Referring to Figure 17 , after step 420, the reduction scheduling method provided by the embodiments of the present disclosure further includes:
[0282] Step 1710, execute the third reduction instruction through the target warp of the target execution unit among multiple execution units to reduce the second reduction results of each execution unit to obtain the third reduction result.
[0283] It should be noted that the computing unit is a unit for measuring computing resources in a computer system. Referring to Figure 3 , multiple execution units are usually set in the computing unit.
[0284] The following provides a detailed description of step 1710.
[0285] In step 1710, the third reduction instruction is executed through the target warp of the target execution unit among multiple execution units to reduce the second reduction results of the respective execution units and obtain the third reduction result.
[0286] The target execution unit is one of the multiple execution units corresponding to the computing unit. Referring to Figure 3 , the target execution unit can be any one of execution unit EU0, execution unit EU1, execution unit EU2, and execution unit EU3. Additionally, the target execution unit is the execution unit in the computing unit where a warp is set, that is, the target execution unit is one of the multiple execution units for executing the target reduction task.
[0287] The third reduction instruction refers to an instruction for reducing the second reduction results of the respective execution units. The third reduction instruction is executed through the target warp to reduce the second reduction results of the respective execution units and obtain the third reduction result. Referring to Figure 3 , assuming that execution unit EU0, execution unit EU1, execution unit EU2, and execution unit EU3 are all used to execute the target reduction task, and execution unit EU0 is used as the target execution unit, the third reduction instruction is executed through the target warp in execution unit EU0 to reduce the second reduction results of execution unit EU0, execution unit EU1, execution unit EU2, and execution unit EU3 and obtain the third reduction result.
[0288] In the embodiment of the above step 1710, when the execution unit is multiple execution units in the computing unit, the third reduction instruction is executed through the target warp of the target execution unit to reduce the second reduction results of the respective execution units. This method is applicable to the situation where the execution unit cannot accommodate the multiple vectors to be reduced corresponding to the target reduction task and is applicable to the case of a large amount of data, thereby improving the applicability of the reduction scheduling method and the reduction efficiency.
[0289] The above is the overall description of step 1710. The following provides a detailed description of the specific implementation process of step 1710.
[0290] In step 1710, the third reduction instruction is executed through the target warp of the target execution unit among multiple execution units to reduce the second reduction results of the respective execution units and obtain the third reduction result.
[0291] In one embodiment, the computing unit further has a shared memory, and the second reduction results of the respective execution units are stored in the shared memory. Referring to Figure 18 , step 1710 includes:
[0292] Step 1810: Read the second reduction results of each execution unit in the shared memory into the reduction unit of the target execution unit;
[0293] Step 1820: Reduce the second reduction results of each execution unit in the reduction unit of the target execution unit to obtain a third reduction result.
[0294] It should be noted that the shared memory refers to the large-capacity memory in the computing unit that can be accessed by different execution units. During the execution of the second reduction instruction, after the second reduction result is obtained, the second reduction result is automatically transferred and stored in the shared memory of the computing unit. For example, referring to Figure 3 , after obtaining the second reduction results of each execution unit, execution unit EU0, execution unit EU1, execution unit EU2, and execution unit EU3 automatically transfer their corresponding second reduction results to the shared memory.
[0295] The following is a detailed description of Step 1810 and Step 1820.
[0296] In Step 1810, the second reduction results of each execution unit in the shared memory are read into the reduction unit of the target execution unit.
[0297] The scheduling unit reads the second reduction results of each execution unit from the shared memory and writes the read second reduction results into the reduction unit of the target execution unit.
[0298] Suppose Figure 4A is Figure 3 a schematic diagram of execution unit EU0 in
[0299] and execution unit EU0 is the target execution unit. Then, the second reduction results of execution unit EU0, execution unit EU1, execution unit EU2, and execution unit EU3 in the shared memory are read into the reduction unit of execution unit EU0.
[0300] In Step 1820, the second reduction results of each execution unit in the reduction unit of the target execution unit are reduced to obtain a third reduction result.
[0301] It should be noted that after obtaining the third reduction result, the third reduction result is stored in a register or the shared memory. The reduction unit is only a module that executes the reduction operation and a module that stores the data that needs to be cached during the reduction operation.
[0302] In the embodiment of the above steps 1810 to 1820, the second reduction result is read into the reduction unit of the target execution unit for reduction calculation to obtain a third reduction result. Performing reduction calculation on the second reduction result in the reduction unit of the target execution unit is applicable not only to array reduction (addition), but also to other reduction methods such as multiplication, improving the applicability of reduction scheduling.
[0303] In one embodiment, the computing unit further has a shared memory, and the second reduction results of each execution unit are stored in the shared memory. Referring to Figure 19 , step 1710 includes:
[0304] Step 1910, using a shared memory atomic operation, reduce the second reduction results of each execution unit in the shared memory to obtain a third reduction result.
[0305] It should be noted that in the embodiment of the present disclosure, the second reduction results of each execution unit can be reduced in the shared memory by using a shared memory atomic operation to obtain a third reduction result. The shared memory atomic operation can quickly complete the summation in array reduction (addition), is applicable to the scenario of array reduction, and improves the reduction efficiency during array reduction.
[0306] In one embodiment, referring to Figure 20 , the second reduction result is stored in the shared memory in the following manner:
[0307] Step 2010, through the reduction unit in each execution unit, obtain the first storage location of the second reduction result in the scalar register of the execution unit;
[0308] Step 2020, according to the first storage location, read the second reduction result from the scalar registers of each execution unit and write the read second reduction result into the shared memory.
[0309] The following describes steps 2010 and 2020 in detail.
[0310] In step 2010, through the reduction unit in each execution unit, obtain the first storage location of the second reduction result in the scalar register of the execution unit.
[0311] The first storage location corresponds to the execution unit, and the first storage location refers to the storage location of the second reduction result of the corresponding execution unit in the scalar register.
[0312] In step 2020, according to the first storage location, read the second reduction result from the scalar registers of each execution unit and write the read second reduction result into the shared memory.
[0313] The scheduling unit reads the second reduction result stored in the scalar register of each execution unit based on the obtained first storage location. Then, the read second reduction result is stored in the shared memory.
[0314] Referring to Figure 21 , the scheduling unit sequentially reads the second reduction result from the scalar register of each execution unit according to the first storage locations of execution unit 0, execution unit 1, execution unit 2, and execution unit 3 received, and writes the read second reduction result into the shared memory.
[0315] In the embodiments of step 2010 and step 2020 above, the reduction unit of each execution unit obtains the first storage location of the second reduction result in the scalar register of the execution unit, reducing the occurrence of the situation where the first storage location is obtained by other units, and improving the security of the first storage location during the propagation process.
[0316] In another embodiment, the shared memory has a first area and a second area. Referring to Figure 22 , the second reduction result is stored in the shared memory in the following manner:
[0317] Step 2210: Through the reduction unit in each execution unit, write the first storage location of the second reduction result in the scalar register of the execution unit into the first area;
[0318] Step 2220: Read the first storage location from the first area, read the second reduction result from the scalar register of each execution unit according to the first storage location, and write the read second reduction result into the second area.
[0319] It should be noted that the shared memory is provided with a first area and a second area. The first area is an area for storing the first storage location of each second reduction result in the scalar register of the execution unit, and the second area is an area for storing the second reduction result.
[0320] The following will describe step 2210 and step 2220 in detail.
[0321] In step 2210, through the reduction unit in each execution unit, write the first storage location of the second reduction result in the scalar register of the execution unit into the first area.
[0322] For each execution unit, after reducing to obtain the second reduction result, write the first storage location of the second reduction result in the scalar register of the execution unit into the first area of the shared memory. Referring to Figure 22 , after obtaining the second reduction result, the reduction units in execution unit 0, execution unit 1, execution unit 2, and execution unit 3 write the corresponding first storage location of the second reduction result into the first area.
[0323] In step 2220, the first storage location is read from the first region. According to the first storage location, the second reduction result is read from the scalar registers of each execution unit, and the read second reduction result is written into the second region.
[0324] After the writing to the first storage location is completed, the scheduling unit reads the first storage location from the first region, and according to the read first storage location, reads the second reduction result from the scalar registers of each execution unit. After that, the read second reduction result is stored in the second region in the shared memory.
[0325] It should be noted that the target execution unit stores the read second reduction result in the second region of the shared memory specifically to realize the partitioned storage of the first storage location and the second reduction result.
[0326] Refer to Figure 23 , the scheduling unit obtains the first storage locations of execution unit 0, execution unit 1, execution unit 2, and execution unit 3 from the first region of the shared memory, and successively reads the second reduction results from the scalar registers of each execution unit according to the first storage locations, and writes the second reduction results into the second region of the shared memory.
[0327] In the above embodiments of step 2210 and step 2220, a first region and a second region are set in the shared memory, where the first region is used to store the first storage locations of the second reduction results in the scalar registers of the execution units, realizing the broadcast of the second reduction results of each execution unit. The scheduling unit can directly obtain the first storage locations through the shared unit later, without repeatedly calling other execution units, improving the execution efficiency of the third reduction instruction and the reduction efficiency.
[0328] In one embodiment, refer to Figure 24 , the method for selecting a target execution unit among multiple execution units includes:
[0329] Step 2410, obtain the number of idle warps in each execution unit;
[0330] Step 2420, obtain the vector register capacity in the execution unit;
[0331] Step 2430, determine the target execution unit among multiple execution units based on the number of idle warps and the vector register capacity.
[0332] The following conducts a detailed operation on steps 2410 to 2430.
[0333] In step 2410, the number of idle warps in each execution unit is obtained.
[0334] The number of idle threads refers to the number of thread warps in an execution unit that are in an idle state.
[0335] In step 2420, obtain the vector register capacity in the execution unit.
[0336] The vector register capacity refers to the amount of data that the vector registers in the execution unit can hold. The larger the vector register capacity, the more data the vector registers can store, and thus the more data the execution unit can process.
[0337] In step 2430, based on the number of idle thread warps and the vector register capacity, determine a target execution unit among multiple execution units.
[0338] The target execution unit is determined based on the number of idle thread warps and the vector register capacity. The more idle thread warps an execution unit has and the larger its vector register capacity, the better its performance, and the greater the likelihood that this execution unit is the target execution unit.
[0339] Refer to Figure 25 , in the embodiments of the present disclosure, a third score of the execution unit is determined based on the number of idle thread warps, a fourth score of the execution unit is determined based on the vector register capacity, then a third weight corresponding to the third score and a fourth weight corresponding to the fourth score are set, and a weighted sum of the third score and the fourth score is performed based on the third weight and the fourth weight to obtain the total score of the execution unit. The target execution unit is determined among multiple confidence units based on the total score.
[0340] The embodiments of the above steps 2410 to 2430 determine a target execution unit among multiple execution units based on the number of idle thread warps and the vector register capacity, which can improve the processing efficiency of the target execution unit for the second reduction result, and thus improve the reduction scheduling efficiency.
[0341] Scheduling and execution of the fourth reduction instruction
[0342] In the embodiments of the above step 1710, the execution units are multiple execution units in a computing unit. Then, through the target thread warps of the target execution unit in the multiple execution units, the third reduction instruction is executed to use the reduction unit in the target execution unit to reduce the second reduction results of each execution unit to obtain the third reduction result. In this case, in one embodiment, the computing unit is multiple computing units in a computing component. Refer to Figure 26 , after step 1710, the reduction scheduling method provided by the embodiments of the present disclosure includes:
[0343] Step 2610, through the target thread warps of the target execution unit of the target computing unit among multiple computing units, execute the fourth reduction instruction to reduce the third reduction results of each computing unit to obtain the fourth reduction result.
[0344] It should be noted that the computing component is a computer processing device with certain computing capabilities, which can perform reduction scheduling for a given task to obtain a reduction result. The computing component can be a CPU, GPU, etc. Refer to Figure 1 and Figure 9 , the computing component includes multiple computing units, and each computing unit in turn includes multiple execution units.
[0345] The following is a detailed description of step 2610.
[0346] In step 2610, the fourth reduction instruction is executed through the target warp of the target execution unit of the target computing unit among the multiple computing units, so as to reduce the third reduction results of each computing unit to obtain the fourth reduction result.
[0347] The target computing unit is one of the multiple computing units corresponding to the computing component. Refer to Figure 9 , the target computing unit can be any one of CU0, CU1, CU2, CU3. In addition, the target computing unit is the computing unit in the computing component with a warp execution unit set, that is, one of the multiple computing units in which the target computing unit executes the target reduction task.
[0348] The fourth reduction instruction refers to the instruction for reducing the third reduction results of each computing unit. During the execution of the fourth reduction instruction, the scheduling unit reduces the third reduction results of each computing unit to obtain the fourth reduction result. Assume that the computing units CU0, CU1, CU2, CU3 are all used to execute the target reduction task, and the computing unit CU0 is used as the target execution unit. Through the target warp in the target execution unit EU0 in CU0, the fourth reduction instruction is executed to reduce the third reduction results of the computing units CU0, CU1, CU2, CU3 to obtain the fourth reduction result.
[0349] In the embodiment of the above step 2610, when the computing unit is multiple computing units in the computing component, the fourth reduction instruction is executed through the target warp of the target execution unit of the target computing unit, so as to reduce the third reduction results of each computing unit. This method is applicable to the case where a single computing unit cannot accommodate the multiple vectors to be reduced corresponding to the target reduction task, and is applicable to the case of a large amount of data, thereby improving the applicability of the reduction scheduling method and the reduction efficiency.
[0350] The above is the overall description of step 2610. The following is a detailed description of the specific implementation process of step 2610.
[0351] In step 2610, the fourth reduction instruction is executed by the target warp of the target execution unit of the target computing unit among multiple computing units to reduce the third reduction results of each computing unit and obtain the fourth reduction result.
[0352] In one embodiment, referring to Figure 27 , step 2610 includes:
[0353] Step 2710, notify each computing unit to store the third reduction result of the computing unit into the shared memory of the computing unit, and obtain the shared memory identifier of the computing unit;
[0354] Step 2720, search for the shared memory of the computing unit according to the shared memory identifier, obtain the third reduction results of each computing unit from the found shared memories of each computing unit, and reduce the third reduction results of each computing unit to obtain the fourth reduction result.
[0355] The following details step 2710 and step 2720.
[0356] In step 2710, notify each computing unit to store the third reduction result of the computing unit into the shared memory of the computing unit, and obtain the shared memory identifier of the computing unit.
[0357] The shared memory identifier refers to the identifier of the shared memory of the computing unit, which is used to distinguish the shared memories of each computing unit. During the execution of the fourth reduction instruction, notify each computing unit in the computing component to store its corresponding third reduction result into the shared memory of the computing unit. After that, obtain the shared memory identifier of the computing unit to determine the storage address of the third reduction result.
[0358] In step 2720, search for the shared memory of the computing unit according to the shared memory identifier, obtain the third reduction results of each computing unit from the found shared memories of each computing unit, and reduce the third reduction results of each computing unit to obtain the fourth reduction result.
[0359] After the third reduction result is stored in the shared memory of its corresponding computing unit, the scheduling unit searches for the shared memories of each computing unit based on the shared memory identifier and obtains the third reduction results of each computing unit from them to reduce the third reduction results and obtain the fourth reduction result.
[0360] Referring to Figure 28, each computing unit is provided with a corresponding shared memory. During the execution of the fourth reduction instruction, first, the notification computing units 0, 1, and 2 are utilized to store their corresponding third reduction results into the shared memory and obtain the shared memory identifiers of the computing units. Then, the scheduling unit searches for the shared memories of the respective computing units according to the shared memory identifiers and obtains the third reduction results of the respective computing units from them, so as to reduce the third reduction results to obtain the fourth reduction result.
[0361] In the embodiments of the above steps 2710 and 2720, the second reduction results of their corresponding computing units are stored in the shared memory, and the shared memory identifiers of the computing units are obtained. Then, based on the shared memory identifiers, the shared memories of the respective computing units are searched, and the third reduction results of the respective computing units are obtained from them, so as to reduce the third reduction results to obtain the fourth reduction result, ensuring that the fourth reduction instruction can be correctly executed and improving the efficiency of reduction scheduling.
[0362] In another embodiment, referring to Figure 9 , a global memory is provided in the computing component. The global memory refers to a large-capacity memory in the computing component that can be accessed by different computing units. Therefore, the third reduction results of the respective computing units can be stored in the global memory, and then the scheduling unit can reduce the third reduction results of the respective computing units in the global memory to obtain the fourth reduction result.
[0363] In step 2720, the shared memory of the computing unit is searched according to the shared memory identifier, the third reduction results of the respective computing units are obtained from the searched shared memories of the respective computing units, and the third reduction results of the respective computing units are reduced to obtain the fourth reduction result.
[0364] In one embodiment, referring to Figure 29 , step 2720 includes:
[0365] Step 2910: Read the third reduction results of the respective computing units into the reduction unit of the target execution unit of the target computing unit;
[0366] Step 2920: Reduce the third reduction results of the respective computing units in the reduction unit of the target execution unit of the target computing unit to obtain the fourth reduction result.
[0367] The following describes steps 2910 and 2920 in detail.
[0368] In step 2910, the third reduction results of the respective computing units are read into the reduction unit of the target execution unit of the target computing unit.
[0369] Locate the shared memory of the computing units according to the shared memory identifier, read the third reduction results of the respective computing units from the located shared memories of the computing units, and store the read third reduction results in the reduction unit of the target execution unit of the target computing unit.
[0370] In step 2920, reduce the third reduction results of the respective computing units in the reduction unit of the target execution unit of the target computing unit to obtain a fourth reduction result.
[0371] After reading the third reduction results into the reduction unit of the target execution unit of the target computing unit, reduce the third reduction results of the respective computing units in the reduction unit to obtain a fourth reduction result. Refer to Figure 28 , the scheduling unit reads the third reduction results of computing unit 0, computing unit 1, and computing unit 2 into computing unit 0, that is, the reduction unit of the target execution unit of the target computing unit, and reduces the third reduction results of the respective computing units in the reduction unit to obtain a fourth reduction result.
[0372] In the embodiments of the above steps 2910 and 2920, the third reduction results of the respective computing units are reduced in the reduction unit of the target execution unit of the target computing unit to obtain a fourth reduction result. Reducing the second reduction result in the reduction unit of the target execution unit is applicable to various reduction methods such as array reduction and multiplication reduction, improving the applicability of reduction scheduling.
[0373] Synchronization of multiple threads in a warp
[0374] In one embodiment, refer to Figure 30 , after step 230, the reduction scheduling method provided by the embodiments of the present disclosure includes:
[0375] Step 3010: If each thread in a single warp has obtained the first reduction result, broadcast a synchronization message to other warps;
[0376] Corresponding to step 3010, step 240 includes:
[0377] Step 3020: If the target warp among the multiple warps corresponding to the execution unit has received the synchronization messages of all other warps, execute a second reduction instruction.
[0378] The following describes steps 3010 and 3020 in detail.
[0379] In step 3010, if each thread in a single warp has obtained the first reduction result, broadcast a synchronization message to other warps.
[0380] The synchronization message is used to indicate that all the warps corresponding to the synchronization message have obtained the first reduction result, that is, the first reduction instruction part corresponding to the warp has been completed.
[0381] In step 3020, if the target warp in the multiple warps corresponding to the execution unit has received the synchronization messages of all the other warps, execute the second reduction instruction.
[0382] If the target warp in the multiple warps corresponding to the execution unit has received the synchronization messages of all the other warps, the first reduction instruction has been executed. At this time, the second reduction instruction can be started. Refer to Figure 4B , all the vectors A1 to AN to be reduced in warp 1 have obtained the corresponding first reduction results, then broadcast the synchronization message corresponding to warp 1 to other warps. In addition, the target warp has received the synchronization messages of warp 1 to warp M, the first reduction instruction has been executed, and the second reduction instruction is started.
[0383] The embodiments of the above steps 3010 and 3020 are based on executing the second reduction instruction after receiving the synchronization messages of all the other warps, which can ensure that the second reduction instruction is executed to reduce the first reduction results of multiple vectors to be reduced during the execution process, improving the accuracy of the execution of the second reduction instruction and the accuracy of reduction scheduling.
[0384] It can be understood that although the steps in the above various flowcharts are sequentially shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of the steps or stages in other steps or other steps.
[0385] Device description of the embodiments of the present disclosure
[0386] Figure 31It is a schematic structural diagram of a reduction scheduling device 3100 provided by an embodiment of the present disclosure. The reduction scheduling device 3100 includes a scheduling unit 3110 and at least one execution unit 3120. The execution unit 3120 includes a reduction unit 3121, a vector register 3122, and a scalar register 3123. Among them, the vector register 3122 includes a vector register area corresponding to each warp. The vector register area 3122 includes vector register rows corresponding to each thread in the warp; the scalar register 3123 includes a scalar register area corresponding to each warp. The scalar register area includes scalar register bits corresponding to each thread in the warp; the scheduling unit 3110 is configured to:
[0387] For a target reduction task, set a plurality of warps corresponding to the execution unit 3120;
[0388] Store each vector to be reduced of the target reduction task into each vector register row;
[0389] Through each warp, execute a first reduction instruction to, for each thread in the warp, use the reduction unit 3121 to reduce each vector element of the vector to be reduced in the vector register row corresponding to the thread, obtain a first reduction result, and store it in the scalar register bit corresponding to the thread;
[0390] Through a target warp among the multiple warps corresponding to the execution unit 3120, execute a second reduction instruction to use the reduction unit 3121 to reduce each scalar register bit in the scalar register 3123 in the execution unit 3120, and obtain a second reduction result.
[0391] Optionally, the reduction unit 3121 is specifically configured to:
[0392] Reduce the scalar register bits corresponding to each thread in a single warp in the scalar register 3123 to obtain a first intermediate reduction result:
[0393] Reduce the first intermediate reduction results of each warp in the scalar register 3123 to obtain a second reduction result.
[0394] Optionally, the reduction unit 3121 is further specifically configured to:
[0395] Allocate thread indices to each thread in a single warp in the scalar register 3123 according to a predetermined index allocation rule;
[0396] Reduce the scalar register bits corresponding to the threads with the same thread index in each warp to obtain a second intermediate reduction result:
[0397] Reduce the second intermediate reduction results corresponding to each thread index in the scalar register 3123 to obtain a second reduction result.
[0398] Optionally, the scheduling unit 3110 is further specifically configured to:
[0399] Enable the target warp to have access rights to each scalar register area in the scalar register 3123, while other warps in the multiple warps only have access rights to the scalar register areas corresponding to the other warps.
[0400] Optionally, referring to Figure 3 , the execution unit 3120 is multiple execution units 3120 in the computing unit;
[0401] The scheduling unit 3110 is further specifically configured to:
[0402] Execute a third reduction instruction through the target warp of the target execution unit among the multiple execution units 3120, so as to use the reduction unit 3121 in the target execution unit to reduce the second reduction results of each execution unit 3120 to obtain a third reduction result.
[0403] Optionally, the scheduling unit 3110 is further specifically configured to:
[0404] Obtain the number of idle warps in each execution unit 3120;
[0405] Obtain the vector register capacity in the execution unit 3120;
[0406] Determine a target execution unit among the multiple execution units based on the number of idle warps and the vector register capacity.
[0407] Optionally, the computing unit further has a shared memory, and the second reduction results of each execution unit 3120 are stored in the shared memory. The scheduling unit 3110 is further specifically configured to:
[0408] Read the second reduction results of each execution unit 3120 in the shared memory into the reduction unit 3121 of the target execution unit;
[0409] Reduce the second reduction results of each execution unit 3120 in the reduction unit 3121 of the target execution unit to obtain a third reduction result.
[0410] Optionally, the computing unit further has a shared memory, and the second reduction results of each execution unit 3120 are stored in the shared memory. The scheduling unit 3110 is further specifically configured to:
[0411] Use shared memory atomic operations to reduce the second reduction results of each execution unit 3120 in the shared memory to obtain a third reduction result.
[0412] Optionally, the scheduling unit 3110 is further specifically configured to:
[0413] Through the reduction unit 3121 in each execution unit 3120, obtain the first storage location of the second reduction result in the scalar register of the execution unit 3120;
[0414] According to the first storage location, read the second reduction result from the scalar registers of the respective execution units 3120, and write the read second reduction result into the shared memory.
[0415] Optionally, the shared memory has a first area and a second area, and the scheduling unit 3110 is further specifically configured to:
[0416] Through the reduction unit 3121 in each execution unit 3120, write the first storage location of the second reduction result in the scalar register of the execution unit 3120 into the first area;
[0417] Read the first storage location from the first area, and according to the first storage location, read the second reduction result from the scalar registers of the respective execution units 3120, and write the read second reduction result into the second area.
[0418] Optionally, the computing unit is a plurality of computing units in the computing component;
[0419] The scheduling unit 3110 is further specifically configured to:
[0420] Execute a fourth reduction instruction through the target warp of the target execution unit of the target computing unit among the plurality of computing units to reduce the third reduction results of the respective computing units to obtain a fourth reduction result.
[0421] Optionally, the scheduling unit 3110 is further specifically configured to:
[0422] Notify each computing unit to store the third reduction result of the computing unit in the shared memory of the computing unit, and obtain the shared memory identifier of the computing unit;
[0423] Find the shared memory of the computing unit according to the shared memory identifier, obtain the third reduction results of the respective computing units from the found shared memories of the respective computing units, and reduce the third reduction results of the respective computing units to obtain a fourth reduction result.
[0424] Optionally, the scheduling unit is further specifically configured to:
[0425] Read the third reduction results of the respective computing units into the reduction unit 3121 of the target execution unit of the target computing unit;
[0426] In the reduction unit 3121 of the target execution unit of the target computing unit, reduce the third reduction results of the respective computing units to obtain a fourth reduction result.
[0427] Optionally, the scheduling unit 3110 is further specifically configured to:
[0428] If each thread in a single warp has obtained the first reduction result, broadcast a synchronization message to other warps;
[0429] Execute a second reduction instruction through a target warp among multiple warps corresponding to the execution unit 3120, including: if the target warp among multiple warps corresponding to the execution unit 3120 has received the synchronization messages of each other warp, execute the second reduction instruction.
[0430] Optionally, the scheduling unit 3110 is further specifically configured to:
[0431] Obtain the number of vectors to be reduced of the target reduction task, the maximum number of threads in a warp, and the maximum number of warps in the execution unit 3120;
[0432] Calculate a first product of the maximum number of threads and the maximum number of warps;
[0433] If the number of vectors to be reduced is not greater than the first product, determine a first number based on the ratio of the number of vectors to be reduced to the maximum number of threads;
[0434] Set the first number of warps in the execution unit 3120.
[0435] Optionally, referring to Figure 3 , the execution unit 3120 is multiple execution units in a computing unit;
[0436] The scheduling unit 3110 is further specifically configured to:
[0437] If the number of vectors to be reduced is greater than the first product, determine a second number based on the ratio of the number of vectors to be reduced to the first product;
[0438] If the second number is not greater than the number of execution units included in the computing unit, for the second number of execution units, set the multiple warps corresponding to each execution unit 3120.
[0439] Optionally, referring to Figure 9 , the computing unit is multiple computing units in a computing component;
[0440] The scheduling unit 3110 is further specifically configured to:
[0441] If the second number is greater than the number of execution units included in the computing unit, determine a third number based on the ratio of the second number to the number of execution units;
[0442] If the third number is not greater than the number of computing units included in the computing component, for the third number of computing units, set the multiple warps corresponding to the execution units of each computing unit.
[0443] Optionally, the scheduling unit 3110 is further specifically configured to:
[0444] Determine a plurality of vector register regions corresponding to the set of thread warps;
[0445] Successively store each vector to be reduced of the target reduction task into the vector register rows in each vector register region.
[0446] Optionally, the scheduling unit 3110 is further specifically configured to:
[0447] Obtain the remaining processing capabilities of each thread warp;
[0448] Obtain the capacity of the vector register region corresponding to each thread warp;
[0449] Based on the remaining processing capabilities and the capacity of the vector register region, determine a target thread warp among the plurality of thread warps.
[0450] Optionally, the scheduling unit 3110 is further specifically configured to:
[0451] Based on the remaining processing capabilities, determine a first score of the thread warp;
[0452] Based on the capacity of the vector register region, determine a second score of the thread warp;
[0453] Based on the first score and the second score, determine the total score of the thread warp;
[0454] Based on the total scores of the respective thread warps, determine a target thread warp among the plurality of thread warps.
[0455] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar contents and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0456] It should be understood that in this disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated content and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates that the associated content before and after is an "or" relationship. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of a single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or plural.
[0457] It should be understood that in the description of the embodiments of this disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0458] It should also be understood that the various embodiments provided in the embodiments of this disclosure can be combined arbitrarily to achieve different technical effects.
[0459] The above is a specific description of the embodiments of this disclosure, but this disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of this disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this disclosure.
Claims
1. A reduction scheduling method, characterized in that: A scheduling unit is used for scheduling an execution unit, wherein the execution unit includes a reduction unit, a vector register and a scalar register, wherein the vector register includes a vector register area corresponding to each thread warp, and the vector register area includes a vector register row corresponding to each thread in the thread warp; the scalar register includes a scalar register area corresponding to each thread warp, and the scalar register area includes a scalar register bit corresponding to each thread in the thread warp; the reduction scheduling method includes: For a target reduction task, setting a plurality of the thread warps corresponding to the execution unit; storing each to-be-reduced vector of the target reduction task in each of the vector register rows, wherein the to-be-reduced vector includes a plurality of vector elements; Executing a first reduction instruction through each of the thread warps, so as to reduce each vector element of the to-be-reduced vector in the vector register row corresponding to the thread by using the reduction unit for each of the threads in the thread warp, to obtain a first reduction result, and store the result in the scalar register bit corresponding to the thread; The target warp is allowed to have access rights to each of the scalar register regions in the scalar register, while other warps in the plurality of warps only have access rights to the scalar register regions corresponding to the other warps; A second reduction instruction is executed by a target thread warp in the plurality of thread warps corresponding to the execution unit, so as to reduce each of the scalar register bits in the scalar register in the execution unit by using the reduction unit to obtain a second reduction result.
2. The reduction scheduling method according to claim 1, characterized in that: The reducing each of the scalar register bits in the scalar register in the execution unit to obtain a second reduction result includes: Reducing the scalar register bits corresponding to each of the threads in a single warp in the scalar register to obtain a first intermediate reduction result; The first intermediate reduction result of each of the thread warps in the scalar register is reduced to obtain the second reduction result.
3. The reduction scheduling method according to claim 1, characterized in that: The reducing each of the scalar register bits in the scalar register in the execution unit to obtain a second reduction result includes: Allocating a thread index for each of the threads in a single thread warp in the scalar register according to a predetermined index allocation rule; Reducing the scalar register bits corresponding to the threads with the same thread index in each of the thread warps to obtain a second intermediate reduction result; The second intermediate reduction result corresponding to each of the thread indexes in the scalar register is reduced to obtain the second reduction result.
4. The reduction scheduling method according to claim 1, characterized in that: The execution unit is a plurality of the execution units in the computing unit; After executing a second reduction instruction by a target thread warp in the plurality of thread warps corresponding to the execution unit to reduce each of the scalar register bits in the scalar register in the execution unit to obtain a second reduction result, the reduction scheduling method further includes: A third reduction instruction is executed by the target thread warp of the target execution unit among the plurality of execution units to reduce the second reduction results of each of the execution units to obtain a third reduction result.
5. The reduction scheduling method according to claim 4, characterized in that: The computing unit further comprises a shared memory, and the second reduction result of each of the execution units is stored in the shared memory; The reducing the second reduction results of each of the execution units to obtain a third reduction result includes: Reading the second reduction results of each of the execution units in the shared memory into the reduction unit of the target execution unit; The second reduction results of each of the execution units are reduced in the reduction unit of the target execution unit to obtain the third reduction result.
6. The reduction scheduling method according to claim 4, characterized in that: The computing unit further comprises a shared memory, and the second reduction result of each of the execution units is stored in the shared memory; The reducing the second reduction results of each of the execution units in the shared memory to obtain the third reduction result includes: The second reduction results of each of the execution units in the shared memory are reduced by using a shared memory atomic operation to obtain the third reduction result.
7. The reduction scheduling method according to claim 4, characterized in that: The computing unit is a plurality of computing units in a computing component; After executing a third reduction instruction through the target thread warp of the target execution unit among the plurality of execution units to reduce the second reduction results of each of the execution units to obtain a third reduction result, the reduction scheduling method further includes: A fourth reduction instruction is executed by the target warp of the target execution unit of the target computing unit among the plurality of computing units to reduce the third reduction results of each computing unit to obtain a fourth reduction result.
8. The reduction scheduling method according to claim 1, characterized in that: After executing a first reduction instruction through each of the thread warps to reduce, for each of the threads in the thread warps, each vector element of the to-be-reduced vector in the vector register row corresponding to the thread using the reduction unit to obtain a first reduction result and storing the result in the scalar register bit corresponding to the thread, the reduction scheduling method further includes: If all the threads in a single warp obtain the first reduction result, broadcast a synchronization message to other warps; The executing the second reduction instruction by the target warp among the plurality of warps corresponding to the execution unit includes: if the target warp among the plurality of warps corresponding to the execution unit has received the synchronization messages from the other warps, executing the second reduction instruction.
9. The reduction scheduling method according to claim 1, characterized in that: The step of setting the plurality of thread warps corresponding to the execution unit for the target reduction task includes: Obtaining the number of vectors to be reduced of the target reduction task, the maximum number of threads in the thread warp, and the maximum number of thread warps in the execution unit; Calculating a first product of the maximum number of threads and the maximum number of warps; If the number of vectors to be reduced is not greater than the first product, determining a first number based on a ratio of the number of vectors to be reduced to the maximum number of threads; A first number of the warps are arranged in the execution unit.
10. The reduction scheduling method according to claim 9, characterized in that: The execution unit is a plurality of the execution units in the computing unit; After calculating the first product of the maximum number of threads and the maximum number of warps, the reduction scheduling method further includes: If the number of vectors to be reduced is greater than the first product, determining a second number based on a ratio of the number of vectors to be reduced to the first product; If the second number is not greater than the number of execution units included in the computing unit, a plurality of the thread warps corresponding to each execution unit are set for the second number of execution units.
11. The reduction scheduling method according to claim 10, characterized in that: The computing unit is a plurality of computing units in a computing component; If the number of vectors to be reduced is greater than the first product, after determining a second number based on a ratio of the number of vectors to be reduced to the first product, the reduction scheduling method further includes: If the second number is greater than the number of execution units included in the computing unit, determining a third number based on a ratio of the second number to the number of execution units; If the third number is not greater than the number of computing units included in the computing component, a plurality of the thread warps corresponding to the execution unit of each computing unit are set for the third number of computing units.
12. The reduction scheduling method according to claim 1, characterized in that: The step of storing each to-be-reduced vector of the target reduction task into each of the vector register rows comprises: Determine a plurality of the vector register regions corresponding to the set plurality of the thread warps; Each to-be-reduced vector of the target reduction task is sequentially stored in each of the vector register rows in the vector register area.
13. A reduction scheduling device, characterized in that: The system comprises a scheduling unit and at least one execution unit, wherein the execution unit comprises a reduction unit, a vector register and a scalar register, wherein the vector register comprises a vector register area corresponding to each thread warp, the vector register area comprises a vector register row corresponding to each thread in the thread warp, and the vector register row comprises a plurality of vector elements; the scalar register comprises a scalar register area corresponding to each thread warp, the scalar register area comprises a scalar register bit corresponding to each thread in the thread warp; the scheduling unit is used to: For a target reduction task, setting a plurality of the thread warps corresponding to the execution unit; storing each to-be-reduced vector of the target reduction task into each of the vector register rows; Executing a first reduction instruction through each of the thread warps, so as to reduce each vector element of the to-be-reduced vector in the vector register row corresponding to the thread by using the reduction unit for each of the threads in the thread warp, to obtain a first reduction result, and store the result in the scalar register bit corresponding to the thread; The target warp is allowed to have access rights to each of the scalar register regions in the scalar register, while other warps in the plurality of warps only have access rights to the scalar register regions corresponding to the other warps; A second reduction instruction is executed by a target thread warp in the plurality of thread warps corresponding to the execution unit, so as to reduce each of the scalar register bits in the scalar register in the execution unit by using the reduction unit to obtain a second reduction result.
Citation Information
Patent Citations
Parallel multivalue reductions
CN111448545A