Vector calculation device
By designing instruction scheduling and decoding structures for only a single vector core in vector computing devices, the problems of logical complexity and chip area during multiple SIMT instructions are solved, and simpler logic and more efficient power consumption balance are achieved.
Patent Information
- Application Number
- CN202510712451.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
In GPU, when scheduling multiple SIMT instructions, the resource usage of the previous instructions needs to be considered, resulting in increased logic complexity and increased chip area, which may cause successful imbalance and heat dissipation problems.
The structure of a vector computing device is designed so that instruction scheduling and instruction decoding are only for a single vector core, and instructions are sent to the unique instruction decoding unit through the instruction scheduling unit, and operand addresses are sent to the single vector core execution module by the instruction decoding unit, simplifying instruction scheduling and decoding logic.
It reduces the logical complexity of instruction scheduling and decoding, reduces the chip area, improves power consumption balance, and reduces the difficulty of heat dissipation.
Smart Images

Figure CN120234045A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of processors, and more particularly, to a vector computing device. Background Art
[0002] Single Instruction Multiple Threads (SIMT) instructions are common programming instructions for Graphic Processing Units (GPUs) and have been widely used in large-scale parallel computing tasks such as graphics rendering and Artificial Intelligence (AI) operations. SIMT instructions perform parallel operations in units of thread groups (warps), and a thread group contains multiple threads. For example, SIMT32 instructions represent instructions containing 32 threads, and a single SIMT32 instruction can cause the threads in a warp containing 32 threads to execute the same operation in parallel. In this way, GPUs or AI chips can process multiple data points or computing tasks simultaneously, thereby improving computing efficiency.
[0003] In a GPU, SIMT instructions are executed through a Vector Core to perform corresponding vector operations, including but not limited to floating-point operations, fixed-point operations, logical operations, etc. The Vector Core is also called a Vector Execution Unit and is usually combined with a pipeline architecture to support large-scale parallel computing tasks.
[0004] The Vector Core may include multiple computing units, and each computing unit is used to process the data of one thread in a warp. Currently, usually only a part of the threads of a SIMT instruction, such as half of the threads, are instantiated in a single Vector Core, so that two Vector Cores can execute all operation stages of a SIMT instruction, such as instruction fetching, resource checking, instruction scheduling, decoding, operand fetching, operation execution, result write-back, etc., in a ping-pong manner.
[0005] However, in this ping-pong manner, when scheduling multiple SIMT instructions, it is necessary to consider the resource usage of previous instructions to avoid conflicts. At the same time, when the register is read and written by two Vector Cores, complex gating logic is also required for judgment. In addition, different scheduled instructions may also cause power consumption imbalance and heat dissipation problems. Summary of the Invention
[0006] To address at least one of the above problems, the solution of the present disclosure enables instruction scheduling and instruction decoding of a single instruction to be only for a single Vector Core by designing the structure of the vector computing device, thereby reducing the logical complexity of instruction scheduling and instruction decoding, and also being able to significantly reduce the chip area.
[0007] According to one aspect of the present disclosure, a vector computing device is provided. The vector computing device includes: an instruction scheduling unit, an instruction decoding unit connected to the instruction scheduling unit, and a plurality of vector core execution modules, wherein the plurality of vector core execution modules are connected in series and only one first vector core execution module is connected to the instruction decoding unit. Wherein, the instruction scheduling unit is configured to schedule single instruction multiple threads (SIMT) instructions and send the SIMT instructions to the instruction decoding unit, and the instruction decoding unit is configured to decode the SIMT instructions to determine the operand addresses of the operands of the SIMT instructions, and instruct the plurality of vector core execution modules to perform vector operations of the SIMT instructions on the operands respectively in a plurality of clock cycles.
[0008] In some embodiments, the instruction decoding unit is configured to enable the first vector core execution module to read the operand from the operand address in one clock cycle, and the first vector execution module is configured to enable the next vector execution module connected to the first vector execution module to read the operand from the operand address in the next clock cycle of the clock cycle.
[0009] In some embodiments, each vector core execution module includes: an operand acquisition unit, a register, and a vector core calculation unit, wherein the operand acquisition unit reads the operand from the register based on the operand address from the instruction decoding unit, and sends the read operand to the vector core calculation unit, and the vector core calculation unit performs vector operations of the SIMT instructions using the operand and writes the operation result back to the register.
[0010] In some embodiments, the vector core calculation units of each vector core execution module in the plurality of vector core execution modules respectively start the vector operations in the plurality of clock cycles in sequence based on shared control information from the instruction decoding unit.
[0011] In some embodiments, each calculation unit in the vector core calculation unit is used to perform vector operations of one thread of the SIMT instructions.
[0012] In some embodiments, the multiple vector core execution modules include a first vector core execution module and a second vector core execution module. The first vector core execution module includes a first operand acquisition unit, a first register, and a first vector core calculation unit. The first operand acquisition unit reads the operands of the upper half thread group of the SIMT instruction from the first register based on the operand address from the instruction decoding unit, and sends the read operands to the first vector core calculation unit. The first vector core calculation unit performs vector operations on the operands of the upper half thread group of the SIMT instruction using the operands of the upper half thread group, and writes the operation result back to the first register. The second vector core execution module includes a second operand acquisition unit, a second register, and a second vector core calculation unit. The second operand acquisition unit reads the operands of the lower half thread group of the SIMT instruction from the second register based on the operand address from the first operand acquisition unit, and sends the read operands to the second vector core calculation unit. The second vector core calculation unit performs vector operations on the operands of the lower half thread group of the SIMT instruction using the operands of the lower half thread group, and writes the operation result back to the second register.
[0013] In some embodiments, the vector calculation device is configured to: in a first clock cycle, the instruction scheduling unit sends the SIMT instruction to the instruction decoding unit; in the next clock cycle of the first clock cycle, a second clock cycle, the instruction decoding unit decodes the SIMT instruction to obtain the operand address of the operands of the SIMT instruction; in the next clock cycle of the second clock cycle, a third clock cycle, the first operand acquisition unit reads the operands of the upper half thread group of the SIMT instruction from the first register; In the next clock cycle of the third clock cycle, a fourth clock cycle, the first vector core calculation unit performs vector operations on the operands of the upper half thread group of the SIMT instruction. Meanwhile, the second operand acquisition unit reads the operands of the lower half thread group of the SIMT instruction from the second register under the read enable control of the first operand acquisition unit; in the next clock cycle of the fourth clock cycle, a fifth clock cycle, the first vector core calculation unit writes the operation result of performing vector operations on the operands of the upper half thread group of the SIMT instruction back to the first register. Meanwhile, the second vector core calculation unit performs vector operations on the operands of the lower half thread group of the SIMT instruction; in the next clock cycle of the fifth clock cycle, a sixth clock cycle, the second vector core calculation unit writes the operation result of performing vector operations on the operands of the lower half thread group of the SIMT instruction back to the second register.
[0014] In some embodiments, the multiple vector core execution modules include a first vector core execution module, a second vector core execution module, a third vector core execution module, and a fourth vector core execution module. The first vector core execution module includes a first operand acquisition unit, a first register, and a first vector core calculation unit. The first operand acquisition unit reads the operands of 1 / 4 thread groups of the SIMT instruction from the first register based on the operand address from the instruction decoding unit, and sends the read operands to the first vector core calculation unit. The first vector core calculation unit performs vector operations on 1 / 4 thread groups of the SIMT instruction using the operands of 1 / 4 thread groups, and writes the operation result back to the first register. The second vector core execution module includes a second operand acquisition unit, a second register, and a second vector core calculation unit. The second operand acquisition unit reads the operands of 2 / 4 thread groups of the SIMT instruction from the second register based on the operand address from the first operand acquisition unit, and sends the read operands to the second vector core calculation unit. The second vector core calculation unit performs vector operations on 2 / 4 thread groups of the SIMT instruction using the operands of 2 / 4 thread groups, and writes the operation result back to the second register. The third vector core execution module includes a third operand acquisition unit, a third register, and a third vector core calculation unit. The third operand acquisition unit reads the operands of 3 / 4 thread groups of the SIMT instruction from the third register based on the operand address from the second operand acquisition unit, and sends the read operands to the third vector core calculation unit. The third vector core calculation unit performs vector operations on 3 / 4 thread groups of the SIMT instruction using the operands of 3 / 4 thread groups, and writes the operation result back to the third register. And the fourth vector core execution module includes a fourth operand acquisition unit, a fourth register, and a fourth vector core calculation unit. The fourth operand acquisition unit reads the operands of 4 / 4 thread groups of the SIMT instruction from the fourth register based on the operand address from the third operand acquisition unit, and sends the read operands to the fourth vector core calculation unit. The fourth vector core calculation unit performs vector operations on 4 / 4 thread groups of the SIMT instruction using the operands of 4 / 4 thread groups, and writes the operation result back to the fourth register.
[0015] In some embodiments, the vector computing device is configured to: in a first clock cycle, the instruction scheduling unit sends the SIMT instruction to the instruction decoding unit; in the next clock cycle of the first clock cycle, i.e., the second clock cycle, the instruction decoding unit decodes the SIMT instruction to obtain the operand addresses of the operands of the SIMT instruction; in the next clock cycle of the second clock cycle, i.e., the third clock cycle, the first operand acquisition unit reads the operands of 1 / 4 thread groups of the SIMT instruction from the first register; in the next clock cycle of the third clock cycle, i.e., the fourth clock cycle, the first vector core calculation unit performs vector operations on the operands of 1 / 4 thread groups of the SIMT instruction. Meanwhile, the second operand acquisition unit, under the read enable control of the first operand acquisition unit, reads the operands of 2 / 4 thread groups of the SIMT instruction from the second register; in the next clock cycle of the fourth clock cycle, i.e., the fifth clock cycle, the first vector core calculation unit writes the operation result of performing vector operations on the operands of 1 / 4 thread groups of the SIMT instruction back to the first register. Meanwhile, the second vector core calculation unit performs vector operations on the operands of 2 / 4 thread groups of the SIMT instruction. Meanwhile, the third operand acquisition unit, under the read enable control of the second operand acquisition unit, reads the operands of 3 / 4 thread groups of the SIMT instruction from the third register; in the next clock cycle of the fifth clock cycle, i.e., the sixth clock cycle, the second vector core calculation unit writes the operation result of performing vector operations on the operands of 2 / 4 thread groups of the SIMT instruction back to the second register. Meanwhile, the third vector core calculation unit performs vector operations on the operands of 3 / 4 thread groups of the SIMT instruction. Meanwhile, the fourth operand acquisition unit, under the read enable control of the third operand acquisition unit, reads the operands of 4 / 4 thread groups of the SIMT instruction from the fourth register; in the next clock cycle of the sixth clock cycle, i.e., the seventh clock cycle, the third vector core calculation unit writes the operation result of performing vector operations on the operands of 3 / 4 thread groups of the SIMT instruction back to the third register. Meanwhile, the fourth vector core calculation unit performs vector operations on the operands of 4 / 4 thread groups of the SIMT instruction; in the next clock cycle of the seventh clock cycle, i.e., the eighth clock cycle, the fourth vector core calculation unit writes the operation result of performing vector operations on the operands of 4 / 4 thread groups of the SIMT instruction back to the fourth register. Description of the Drawings
[0016] The present disclosure will be better understood by referring to the following detailed description of the specific embodiments of the present disclosure given in the accompanying drawings, and other objects, details, features and advantages of the present disclosure will become more apparent.
[0017] Figure 1 Shows a schematic structural diagram of a vector computing device.
[0018] Figure 2 Shows Figure 1 An exemplary operation timing diagram of the vector computing device shown.
[0019] Figure 3 Shows a schematic structural diagram of a vector computing device according to an embodiment of the present invention.
[0020] Figure 4 Shows Figure 3 A schematic diagram showing a further detailed structure of the vector kernel execution module of the vector computing device shown.
[0021] Figure 5 Shows a schematic diagram of an exemplary vector computing device according to an embodiment of the present invention.
[0022] Figure 6 Shows Figure 5 An exemplary operation timing diagram of the vector computing device shown.
[0023] Figure 7 Shows a schematic diagram of an exemplary vector computing device according to some other embodiments of the present invention.
[0024] Figure 8 Shows Figure 7 An exemplary operation timing diagram of the vector computing device shown. Detailed implementation manners
[0025] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0026] As used herein, the term "comprising" and its variants mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one embodiment" and "some embodiments" mean "at least one exemplary embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0027] Figure 1 Shows a schematic structural diagram of a vector computing device 100. Figure 2 Shows Figure 1Exemplary operation timing diagram of the vector computing device 100 shown
[0028] As Figure 1 shown, the vector computing device 100 may include an instruction scheduling unit 110, and the instruction scheduling unit 110 may be connected to a plurality of vector core execution modules 120 ( Figure 1 two vector core execution modules 120-1 and 120-2 are exemplarily shown therein) to schedule instructions to the plurality of vector core execution modules 120. Here, it is assumed that the vector computing device 100 uses two vector core execution modules 120 to complete a SIMT instruction in a ping-pong manner. That is, each vector core execution module 120 instantiates N / 2 threads each time (where N is the warp required for the SIMT instruction. For example, for a SIMT32 instruction, N = 32).
[0029] Figure 1 In
[0030] As Figure 1 shown, the first vector core execution module 120-1 may include an instruction decoding unit 122-1, an upper half warp operand fetching unit 124-1, a lower half warp operand fetching unit 126-1, and a first vector core computing unit 128-1, where the first vector core computing unit 128-1 may include a plurality of computing units, and each computing unit is used to execute the operation of one thread. The second vector core execution module 120-2 may include an instruction decoding unit 122-2, an upper half warp operand fetching unit 124-2, a lower half warp operand fetching unit 126-2, and a second vector core computing unit 128-2, where the second vector core computing unit 128-2 may include a plurality of computing units, and each computing unit is used to execute the operation of one thread.
[0031] In addition, the vector computing device 100 further includes an upper half warp register 130-1 and a lower half warp register 130-2 to store upper half warp data and lower half warp data respectively.
[0032] In Figure 1When executing SIMT instructions for N threads in the vector computing device 100 shown, usually only N / 2 threads are instantiated in a single vector core execution module 120 each time. That is, the first vector core computing unit 128-1 and the second vector core computing unit 128-2 each contain N / 2 computing units, and each computing unit is used to run a thread. For example, for a SIMT32 instruction, the first vector core computing unit 128-1 with 16 computing units can instantiate the first 16 threads (referred to here as the upper half thread group) and the last 16 threads (referred to here as the lower half thread group) in different clock cycles. For another SIMT instruction, the second vector core computing unit 128-2 with 16 computing units can instantiate the first 16 threads (referred to here as the upper half thread group) and the last 16 threads (referred to here as the lower half thread group) in different clock cycles. Correspondingly, the upper half thread group register 130-1 and the lower half thread group register 130-2 can store the upper half thread group data and the lower half thread group data of these two SIMT32 instructions respectively.
[0033] The complete pipeline process of an instruction includes instruction issue, instruction decoding, operand fetching, operation execution, and result write-back performed in several consecutive clock cycles. Assume that the first vector core execution module 120-1 and the second vector core execution module 120-2 execute multiple SIMT32 instructions (such as Instruction 1, Instruction 2, Instruction 3, Instruction 4) in a ping-pong manner, then as Figure 2 shown, the execution process of these instructions can be described as: For Instruction 1: In the first clock cycle T1, the instruction scheduling unit 110 sends Instruction 1 to the instruction decoding unit 122-1 of the first vector core execution module 120-1; In the second clock cycle T2, the instruction decoding unit 122-1 decodes Instruction 1 to obtain the operand addresses of the operands of Instruction 1; In the third clock cycle T3, the upper half thread group operand fetching unit 124-1 reads the operands of the upper half thread group of Instruction 1 from the upper half thread group register 130-1; In the fourth clock cycle T4, the first vector core computing unit 128-1 performs a vector operation on the operands of the upper half thread group of Instruction 1. At the same time, the lower half thread group operand fetching unit 126-1 reads the operands of the lower half thread group of Instruction 1 from the lower half thread group register 130-2; In the fifth clock cycle T5, the first vector core computing unit 128-1 writes back the operation result of performing the vector operation on the operands of the upper half thread group of Instruction 1 to the upper half thread group register 130-1. At the same time, the first vector core computing unit 128-1 performs a vector operation on the operands of the lower half thread group of Instruction 1; In the sixth clock cycle T6, the first vector core calculation unit 128-1 writes back the operation result of performing a vector operation on the operands of the lower half thread group of instruction 1 to the lower half thread group register 130-2.
[0034] For instruction 2: In the second clock cycle T2, the instruction scheduling unit 110 sends instruction 2 to the instruction decoding unit 122-2 of the second vector core execution module 120-2; In the third clock cycle T3, the instruction decoding unit 122-2 decodes instruction 2 to obtain the operand addresses of the operands of instruction 2; In the fourth clock cycle T4, the upper half thread group operand acquisition unit 124-2 reads the operands of the upper half thread group of instruction 2 from the upper half thread group register 130-1; In the fifth clock cycle T5, the second vector core calculation unit 128-2 performs a vector operation on the operands of the upper half thread group of instruction 2. Meanwhile, the lower half thread group operand acquisition unit 126-2 reads the operands of the lower half thread group of instruction 2 from the lower half thread group register 130-2; In the sixth clock cycle T6, the second vector core calculation unit 128-2 writes back the operation result of performing a vector operation on the operands of the upper half thread group of instruction 2 to the upper half thread group register 130-1. Meanwhile, the second vector core calculation unit 128-2 performs a vector operation on the operands of the lower half thread group of instruction 2; In the seventh clock cycle T7, the second vector core calculation unit 128-2 writes back the operation result of performing a vector operation on the operands of the lower half thread group of instruction 2 to the lower half thread group register 130-2.
[0035] For instruction 3: In the third clock cycle T3, the instruction scheduling unit 110 sends instruction 3 to the instruction decoding unit 122-1 of the first vector core execution module 120-1; In the fourth clock cycle T4, the instruction decoding unit 122-1 decodes instruction 3 to obtain the operand addresses of the operands of instruction 3; In the fifth clock cycle T5, the upper half thread group operand acquisition unit 124-1 reads the operands of the upper half thread group of instruction 3 from the upper half thread group register 130-1; In the sixth clock cycle T6, the first vector core calculation unit 128-1 performs a vector operation on the operands of the upper half thread group of instruction 3. Meanwhile, the lower half thread group operand acquisition unit 126-1 reads the operands of the lower half thread group of instruction 3 from the lower half thread group register 130-2; In the seventh clock cycle T7, the first vector core computing unit 128-1 writes the operation result of performing a vector operation on the operands of the upper half thread group of instruction 3 back to the upper half thread group register 130-1. Meanwhile, the first vector core computing unit 128-1 performs a vector operation on the operands of the lower half thread group of instruction 3; In the eighth clock cycle T8, the first vector core computing unit 128-1 writes the operation result of performing a vector operation on the operands of the lower half thread group of instruction 3 back to the lower half thread group register 130-2.
[0036] For instruction 4: In the fourth clock cycle T4, the instruction scheduling unit 110 sends instruction 4 to the instruction decoding unit 122-2 of the second vector core execution module 120-2; In the fifth clock cycle T5, the instruction decoding unit 122-2 decodes instruction 4 to obtain the operand addresses of the operands of instruction 4; In the sixth clock cycle T6, the upper half thread group operand acquisition unit 124-2 reads the operands of the upper half thread group of instruction 4 from the upper half thread group register 130-1; In the seventh clock cycle T7, the second vector core computing unit 128-2 performs a vector operation on the operands of the upper half thread group of instruction 4. Meanwhile, the lower half thread group operand acquisition unit 126-2 reads the operands of the lower half thread group of instruction 4 from the lower half thread group register 130-2; In the eighth clock cycle T8, the second vector core computing unit 128-2 writes the operation result of performing a vector operation on the operands of the upper half thread group of instruction 4 back to the upper half thread group register 130-1. Meanwhile, the second vector core computing unit 128-2 performs a vector operation on the operands of the lower half thread group of instruction 4; In the ninth clock cycle T9, the second vector core computing unit 128-2 writes the operation result of performing a vector operation on the operands of the lower half thread group of instruction 4 back to the lower half thread group register 130-2.
[0037] It can be seen that the vector computing device 100 realizes the ping-pong execution of multiple instructions by sequentially performing each operation of the pipeline process of SIMT instructions in different clock cycles. This method can save resource area, reduce the risk of register memory access conflicts, and at the same time improve the overall computing power of the device.
[0038] However, in the above manner, when the instruction scheduling unit 110 performs instruction scheduling, it needs to pre-consider the resource usage status of each vector core execution module 120 and avoid register conflicts. For example, the upper half thread group operand fetching unit 124-1 and the lower half thread group operand fetching unit 126-1 of the first vector core execution module 120-1 need to execute the upper half thread group and lower half thread group operand fetching operations of instruction 1 in clock cycles T3 and T4 respectively, and the first vector core calculation unit 128-1 needs to write the results back to the upper half thread group register 130-1 and the lower half thread group register 130-2 in clock cycles T5 and T6 respectively; the upper half thread group operand fetching unit 124-2 and the lower half thread group operand fetching unit 126-2 of the second vector core execution module 120-2 need to execute the upper half thread group and lower half thread group operand fetching operations of instruction 2 in clock cycles T4 and T5 respectively, and the second vector core calculation unit 128-2 needs to write the results back to the upper half thread group register 130-1 and the lower half thread group register 130-2 in clock cycles T6 and T7 respectively. In this case, the upper half thread group register 130-1 and the lower half thread group register 130-2 will be accessed by the two vector core execution modules 120-1 and 120-2 for memory access, including read operations and write operations. Therefore, a multiplexer is needed to distinguish which vector core execution module the operation comes from. In addition, when the instruction scheduling unit 110 allocates instructions to the two vector core execution modules 120-1 and 120-2, a multiplexer is also needed to distinguish which vector core execution module the instruction should be scheduled to. The two vector core calculation units also need data multiplexers to distinguish whether the operand is the read return data from the upper half thread group register or the lower half thread group register, and the upper half thread group register and the lower half thread group register also need multiplexers respectively to distinguish which vector core calculation unit the enable signal and the write data come from. These multiplexers must occupy chip area, thus increasing the chip area.
[0039] In addition, the instruction scheduling unit schedules instructions to the two vector core execution modules in a pipeline form. Since the power consumption of different instructions is different, it is difficult for the instruction scheduling unit to ensure the balance of the instructions scheduled to the two vector core execution modules. Therefore, there may be some vector core execution modules executing instructions with high power consumption, such as executing floating-point multiply-add instructions, and some vector core execution modules executing instructions with low power consumption, such as executing logical operations. This may cause different power density between different vector core execution modules, resulting in local heating of the hardware and increasing the difficulty of heat dissipation.
[0040] To address at least one of the above problems, the present disclosure designs the structure of the vector computing device such that the instruction scheduling unit only sends instructions to a unified instruction decoding unit, and the instruction decoding unit only sends the operand addresses of the decoded operands to a single vector core execution module, so that the instruction scheduling unit does not need to perform register conflict checking in advance.
[0041] Figure 3 Shows a schematic structural diagram of a vector computing device 200 according to an embodiment of the present invention.
[0042] As Figure 3 shown, the vector computing device 200 may include an instruction scheduling unit 210, an instruction decoding unit 220 connected to the instruction scheduling unit 210, and a plurality of vector core execution modules 230 ( Figure 3 exemplarily shows 4 vector core execution modules 230-1, 230-2, 230-3, and 230-4). The plurality of vector core execution modules 230 are connected in series, and only one vector core execution module 230 (such as the first vector core execution module 230-1) is connected to the instruction decoding unit 220.
[0043] The instruction scheduling unit 210 is configured to schedule SIMT instructions and send the SIMT instructions to the instruction decoding unit 220.
[0044] The instruction decoding unit 220 is configured to decode the SIMT instruction to determine the operand addresses of the operands of the SIMT instruction, and instruct the plurality of vector core execution modules 230 to perform vector operations of the SIMT instruction on the operands respectively in a plurality of clock cycles.
[0045] It can be seen that different from the instruction scheduling unit 110 of the vector computing device 100, the instruction scheduling unit 210 of the vector computing device 200 schedules and sends a single SIMT instruction to a single instruction decoding unit 220 each time, without having to determine which of the plurality of instruction decoding units 120 to send the instruction to.
[0046] The instruction decoding unit 220 is configured to enable the vector core execution module 230 (such as the first vector core execution module 230-1) connected thereto to read the operand from the decoded operand address in one clock cycle, and the vector core execution module 230 (such as the first vector core execution module 230-1) is configured to enable the next vector execution module 230 (such as the second vector core execution module 230-2) connected to the vector core execution module 230 to read the operand from the operand address in the next clock cycle of the clock cycle.
[0047] Further, according to the relationship (described below) between the thread number requirement of the SIMT instruction and the number of computing units of each vector core execution module 230, in some cases (such as the case where a SIMT instruction requires more than two vector core execution modules 230 to cooperate in execution), the next vector execution module 230 (such as the second vector core execution module 230-2) is further configured to enable the vector execution module 230 (such as the third vector core execution module 230-3) connected thereto to read the operand from the operand address in the next clock cycle, and so on.
[0048] That is to say, among multiple vector core execution modules 230, only the vector core execution module 230 connected to the instruction decoding unit 220 is enabled by the instruction decoding unit 220 to obtain the operand, while other vector core execution modules 230 are sequentially enabled by the vector core execution module 230 connected thereto to obtain the operand.
[0049] Figure 4 Shows Figure 3 A schematic diagram further showing the detailed structure of the vector core execution module 230 of the vector computing device 200 shown.
[0050] Such as Figure 4 As shown in, the vector core execution module 230 may include an operand acquisition unit 232, a register 234, and a vector core computing unit 236. Among them, the operand acquisition unit 232 may read the operand from the register 234 based on the operand address from the instruction decoding unit 220, and send the read operand to the vector core computing unit 236. Here, when the vector core execution module 230 is directly connected to the instruction decoding unit 220, the vector core execution module 230 may directly obtain the operand address from the instruction decoding unit 220. When the vector core execution module 230 is not directly connected to the instruction decoding unit 220 (that is, when it is connected to another vector core execution module 230), the operand address is passed by another vector core execution module 230 through read enable.
[0051] The vector core computing unit 236 may use the operand to perform the vector operation of the SIMT instruction scheduled by the instruction scheduling unit 210, and write the operation result back to the register 234.
[0052] The vector core computing unit 236 may include multiple computing units, and each computing unit is used to perform the vector operation of one thread of the SIMT instruction.
[0053] Figure 5 Shows a schematic diagram of an exemplary vector computing device 200 according to an embodiment of the present invention. Figure 6 Shows Figure 5 An exemplary operation timing diagram of the vector computing device 200 shown. InFigure 5 In the illustrated example, a SIMT instruction is executed by two vector core execution modules 230 (i.e., the first vector core execution module 230-1 and the second vector core execution module 230-2). Similar to Figure 1 the case in, it is still assumed that for a SIMT instruction of N threads, only N / 2 threads are instantiated in a single vector core execution module 230 each time. For example, for a SIMT32 instruction, the first vector core computing unit 236-1 with 16 computing units can instantiate the first 16 threads (referred to as the upper half thread group here) in the first clock cycle, and the second vector core computing unit 236-2 with 16 computing units can instantiate the last 16 threads (referred to as the lower half thread group here) in the next clock cycle of the first clock cycle.
[0054] As Figure 5 shown, the first vector core execution module 230-1 includes a first operand acquisition unit 232-1, a first register 234-1, and a first vector core computing unit 236-1. The first operand acquisition unit 232-1 reads the operands of the upper half thread group of the SIMT instruction from the first register 234-1 based on the operand address from the instruction decoding unit 220, and sends the read operands to the first vector core computing unit 236-1. The first vector core computing unit 236-1 performs vector operations on the upper half thread group of the SIMT instruction using the operands of the upper half thread group, and writes the operation results back to the first register 234-1.
[0055] The second vector core execution module 230-2 includes a second operand acquisition unit 232-2, a second register 234-2, and a second vector core computing unit 236-2. The second operand acquisition unit 232-2 reads the operands of the lower half thread group of the SIMT instruction from the second register 234-2 based on the operand address from the first operand acquisition unit 232-1, and sends the read operands to the second vector core computing unit 236-2. The second vector core computing unit 236-2 performs vector operations on the lower half thread group of the SIMT instruction using the operands of the lower half thread group, and writes the operation results back to the second register 234-2.
[0056] In addition, the vector core computing units 236 of each vector core execution module 230 respectively start vector operations in multiple clock cycles in sequence based on the shared control information from the instruction decoding unit 220. For example, in Figure 5 the example of, the first vector core computing unit 236-1 and the second vector core computing unit 236-2 can start vector operations on the upper half thread group and the lower half thread group respectively in two consecutive clock cycles (such as Figure 6 the clock cycles T4 and T5 shown).
[0057] As Figure 6 shown, still taking the execution of multiple SIMT 32 instructions (such as Instruction 1, Instruction 2, Instruction 3, Instruction 4) shown in Figure 2 as an example, the execution process of these instructions can be described as follows: For Instruction 1: In the first clock cycle T1, the instruction scheduling unit 210 sends Instruction 1 to the instruction decoding unit 220; In the second clock cycle T2, the instruction decoding unit 220 decodes Instruction 1 to obtain the operand addresses of the operands of Instruction 1; In the third clock cycle T3, the first operand acquisition unit 232-1 reads the operands of the upper half thread group of Instruction 1 from the first register 234-1; In the fourth clock cycle T4, the first vector core calculation unit 236-1 performs a vector operation on the operands of the upper half thread group of Instruction 1. At the same time, the second operand acquisition unit 232-2, under the read enable control of the first operand acquisition unit 232-1, reads the operands of the lower half thread group of Instruction 1 from the second register 234-2; In the fifth clock cycle T5, the first vector core calculation unit 236-1 writes the operation result of performing the vector operation on the operands of the upper half thread group of Instruction 1 back to the first register 234-1. At the same time, the second vector core calculation unit 236-2 performs a vector operation on the operands of the lower half thread group of Instruction 1; In the sixth clock cycle T6, the second vector core calculation unit 236-2 writes the operation result of performing the vector operation on the operands of the lower half thread group of Instruction 1 back to the second register 234-2.
[0058] For Instruction 2: In the second clock cycle T2, the instruction scheduling unit 210 sends Instruction 2 to the instruction decoding unit 220; In the third clock cycle T3, the instruction decoding unit 220 decodes Instruction 2 to obtain the operand addresses of the operands of Instruction 2; In the fourth clock cycle T4, the first operand acquisition unit 232-1 reads the operands of the upper half thread group of Instruction 2 from the first register 234-1; In the fifth clock cycle T5, the first vector core calculation unit 236-1 performs a vector operation on the operands of the upper half thread group of Instruction 2. At the same time, the second operand acquisition unit 232-2, under the read enable control of the first operand acquisition unit 232-1, reads the operands of the lower half thread group of Instruction 2 from the second register 234-2; In the sixth clock cycle T6, the first vector core computing unit 236-1 writes back the operation result of performing a vector operation on the operands of the upper half thread group of instruction 2 to the first register 234-1. Meanwhile, the second vector core computing unit 236-2 performs a vector operation on the operands of the lower half thread group of instruction 2; In the seventh clock cycle T7, the second vector core computing unit 236-2 writes back the operation result of performing a vector operation on the operands of the lower half thread group of instruction 2 to the second register 234-2.
[0059] For instruction 3: In the third clock cycle T3, the instruction scheduling unit 210 sends instruction 3 to the instruction decoding unit 220; In the fourth clock cycle T4, the instruction decoding unit 220 decodes instruction 3 to obtain the operand addresses of the operands of instruction 3; In the fifth clock cycle T5, the first operand acquisition unit 232-1 reads the operands of the upper half thread group of instruction 3 from the first register 234-1; In the sixth clock cycle T6, the first vector core computing unit 236-1 performs a vector operation on the operands of the upper half thread group of instruction 3. Meanwhile, the second operand acquisition unit 232-2, under the read enable control of the first operand acquisition unit 232-1, reads the operands of the lower half thread group of instruction 3 from the second register 234-2; In the seventh clock cycle T7, the first vector core computing unit 236-1 writes back the operation result of performing a vector operation on the operands of the upper half thread group of instruction 3 to the first register 234-1. Meanwhile, the second vector core computing unit 236-2 performs a vector operation on the operands of the lower half thread group of instruction 3; In the eighth clock cycle T8, the second vector core computing unit 236-2 writes back the operation result of performing a vector operation on the operands of the lower half thread group of instruction 3 to the second register 234-2.
[0060] For instruction 4: In the fourth clock cycle T4, the instruction scheduling unit 210 sends instruction 4 to the instruction decoding unit 220; In the fifth clock cycle T5, the instruction decoding unit 220 decodes instruction 4 to obtain the operand addresses of the operands of instruction 4; In the sixth clock cycle T6, the first operand acquisition unit 232-1 reads the operands of the upper half thread group of instruction 4 from the first register 234-1; In the seventh clock cycle T7, the first vector core computing unit 236-1 performs a vector operation on the operands of the upper half thread group of instruction 4. Meanwhile, under the read enable control of the first operand fetching unit 232-1, the second operand fetching unit 232-2 fetches the operands of the lower half thread group of instruction 4 from the second register 234-2; In the eighth clock cycle T8, the first vector core computing unit 236-1 writes the operation result of the vector operation performed on the operands of the upper half thread group of instruction 4 back to the first register 234-1. Meanwhile, the second vector core computing unit 236-2 performs a vector operation on the operands of the lower half thread group of instruction 4; In the ninth clock cycle T9, the second vector core computing unit 236-2 writes the operation result of the vector operation performed on the operands of the lower half thread group of instruction 4 back to the second register 234-2.
[0061] Here, the timing of the vector operations of the first vector core computing unit 236-1 and the second vector core computing unit 236-2 is controlled by the shared control information of the instruction decoding unit 220.
[0062] It can be seen that in the pipelined operation of such SIMT instructions, the instruction scheduling unit 210 only needs to send instructions to the unique instruction decoding unit 220, and the instruction decoding unit 220 only needs to send operand addresses to the vector core execution module 230 connected to it, and the above operations can be repeated cycle by cycle to send and decode different instructions.
[0063] Figure 7 FIG. shows a schematic diagram of an exemplary vector computing device 200 according to some other embodiments of the present invention. Figure 8 FIG. shows Figure 7 an exemplary operation timing diagram of the vector computing device 200 shown. In Figure 7 the example shown, one SIMT instruction is executed by four vector core execution modules 230 (i.e., the first vector core execution module 230-1, the second vector core execution module 230-2, the third vector core execution module 230-3, and the fourth vector core execution module 230-4). And Figure 1 and Figure 5Differently here, it is assumed that for SIMT instructions of N threads, only N / 4 threads are instantiated in the single vector core execution module 230 each time. For example, for a SIMT32 instruction, the first vector core computing unit 236-1 with 8 computing units can instantiate the first N / 4 threads (referred to as the 1 / 4 thread group) in the first clock cycle, the second vector core computing unit 236-2 with 8 computing units can instantiate the second N / 4 threads (referred to as the 2 / 4 thread group) in the next clock cycle (the second clock cycle) of the first clock cycle, the third vector core computing unit 236-3 with 8 computing units can instantiate the third N / 4 threads (referred to as the 3 / 4 thread group) in the next clock cycle (the third clock cycle), and the fourth vector core computing unit 236-4 with 8 computing units can instantiate the fourth N / 4 threads (referred to as the 4 / 4 thread group) in the next clock cycle (the fourth clock cycle).
[0064] As Figure 7 As shown in the figure, the first vector core execution module 230-1 includes a first operand acquisition unit 232-1, a first register 234-1, and a first vector core computing unit 236-1. The first operand acquisition unit 232-1 reads the operands of the 1 / 4 thread group of the SIMT instruction from the first register 234-1 based on the operand address from the instruction decoding unit 220, and sends the read operands to the first vector core computing unit 236-1. The first vector core computing unit 236-1 performs vector operations on the 1 / 4 thread group of the SIMT instruction using the operands of the 1 / 4 thread group, and writes the operation result back to the first register 234-1.
[0065] The second vector core execution module 230-2 includes a second operand acquisition unit 232-2, a second register 234-2, and a second vector core computing unit 236-2. The second operand acquisition unit 232-2 reads the operands of the 2 / 4 thread group of the SIMT instruction from the second register 234-2 based on the operand address from the first operand acquisition unit 232-1, and sends the read operands to the second vector core computing unit 236-2. The second vector core computing unit 236-2 performs vector operations on the 2 / 4 thread group of the SIMT instruction using the operands of the 2 / 4 thread group, and writes the operation result back to the second register 234-2.
[0066] The third vector core execution module 230-3 includes a third operand acquisition unit 232-3, a third register 234-3, and a third vector core calculation unit 236-3. The third operand acquisition unit 232-3 reads the operands of the 3 / 4 thread groups of the SIMT instruction from the third register 234-3 based on the operand address from the second operand acquisition unit 232-2, and sends the read operands to the third vector core calculation unit 236-3. The third vector core calculation unit 236-3 performs vector operations on the 3 / 4 thread groups of the SIMT instruction using the operands of the 3 / 4 thread groups, and writes the operation result back to the second register 234-2.
[0067] The fourth vector core execution module 230-4 includes a fourth operand acquisition unit 232-4, a fourth register 234-4, and a fourth vector core calculation unit 236-4. The fourth operand acquisition unit 232-4 reads the operands of the 4 / 4 thread groups of the SIMT instruction from the fourth register 234-4 based on the operand address from the third operand acquisition unit 232-3, and sends the read operands to the fourth vector core calculation unit 236-4. The fourth vector core calculation unit 236-4 performs vector operations on the 4 / 4 thread groups of the SIMT instruction using the operands of the 4 / 4 thread groups, and writes the operation result back to the fourth register 234-4.
[0068] In addition, the vector core calculation units 236 of each vector core execution module 230 respectively start vector operations in multiple clock cycles based on the shared control information from the instruction decoding unit 220. For example, in Figure 7 the instance of, the first vector core calculation unit 236-1, the second vector core calculation unit 236-2, the third vector core calculation unit 236-3, and the fourth vector core calculation unit 236-4 can start vector operations on the 1 / 4 thread group, 2 / 4 thread group, 3 / 4 thread group, and 4 / 4 thread group respectively in four consecutive clock cycles (such as the clock cycles T4 - T7 shown in Figure 8 .
[0069] Still taking the execution of multiple SIMT32 instructions (such as instruction 1, instruction 2, instruction 3, instruction 4) shown in Figure 2 and Figure 6 as an example, the execution process of these instructions can be described as: For instruction 1: In the first clock cycle T1, the instruction scheduling unit 210 sends instruction 1 to the instruction decoding unit 220; In the second clock cycle T2, the instruction decoding unit 220 decodes instruction 1 to obtain the operand addresses of the operands of instruction 1; In the third clock cycle T3, the first operand acquisition unit 232-1 reads the operands of 1 / 4 thread group of instruction 1 from the first register 234-1; In the fourth clock cycle T4, the first vector core calculation unit 236-1 performs a vector operation on the operands of 1 / 4 thread group of instruction 1. Meanwhile, under the read enable control of the first operand acquisition unit 232-1, the second operand acquisition unit 232-2 reads the operands of 2 / 4 thread group of instruction 1 from the second register 234-2; In the fifth clock cycle T5, the first vector core calculation unit 236-1 writes the operation result of performing the vector operation on the operands of 1 / 4 thread group of instruction 1 back to the first register 234-1. Meanwhile, the second vector core calculation unit 236-2 performs a vector operation on the operands of 2 / 4 thread group of instruction 1. Meanwhile, under the read enable control of the second operand acquisition unit 232-2, the third operand acquisition unit 232-3 reads the operands of 3 / 4 thread group of instruction 1 from the third register 234-3; In the sixth clock cycle T6, the second vector core calculation unit 236-2 writes the operation result of performing the vector operation on the operands of 2 / 4 thread group of instruction 1 back to the second register 234-2. Meanwhile, the third vector core calculation unit 236-3 performs a vector operation on the operands of 3 / 4 thread group of instruction 1. Meanwhile, under the read enable control of the third operand acquisition unit 232-3, the fourth operand acquisition unit 232-4 reads the operands of 4 / 4 thread group of instruction 1 from the fourth register 234-4; In the seventh clock cycle T7, the third vector core calculation unit 236-3 writes the operation result of performing the vector operation on the operands of 3 / 4 thread group of instruction 1 back to the third register 234-3. Meanwhile, the fourth vector core calculation unit 236-4 performs a vector operation on the operands of 4 / 4 thread group of instruction 1; In the eighth clock cycle T8, the fourth vector core calculation unit 236-4 writes the operation result of performing the vector operation on the operands of 4 / 4 thread group of instruction 1 back to the fourth register 234-4.
[0070] For instruction 2: In the second clock cycle T2, the instruction scheduling unit 210 sends instruction 2 to the instruction decoding unit 220; In the third clock cycle T3, the instruction decoding unit 220 decodes instruction 2 to obtain the operand addresses of the operands of instruction 2; In the fourth clock cycle T4, the first operand acquisition unit 232-1 reads the operands of 1 / 4 thread group of instruction 2 from the first register 234-1; In the fifth clock cycle T5, the first vector core computing unit 236-1 performs a vector operation on the operands of the 1 / 4 thread group of instruction 2. Meanwhile, under the read enable control of the first operand fetching unit 232-1, the second operand fetching unit 232-2 reads the operands of the 2 / 4 thread group of instruction 2 from the second register 234-2; In the sixth clock cycle T6, the first vector core computing unit 236-1 writes the operation result of performing the vector operation on the operands of the 1 / 4 thread group of instruction 2 back to the first register 234-1. Meanwhile, the second vector core computing unit 236-2 performs a vector operation on the operands of the 2 / 4 thread group of instruction 2. Meanwhile, under the read enable control of the second operand fetching unit 232-2, the third operand fetching unit 232-3 reads the operands of the 3 / 4 thread group of instruction 2 from the third register 234-3; In the seventh clock cycle T7, the second vector core computing unit 236-2 writes the operation result of performing the vector operation on the operands of the 2 / 4 thread group of instruction 2 back to the second register 234-2. Meanwhile, the third vector core computing unit 236-3 performs a vector operation on the operands of the 3 / 4 thread group of instruction 2. Meanwhile, under the read enable control of the third operand fetching unit 232-3, the fourth operand fetching unit 232-4 reads the operands of the 4 / 4 thread group of instruction 2 from the fourth register 234-4; In the eighth clock cycle T8, the third vector core computing unit 236-3 writes the operation result of performing the vector operation on the operands of the 3 / 4 thread group of instruction 2 back to the third register 234-3. Meanwhile, the fourth vector core computing unit 236-4 performs a vector operation on the operands of the 4 / 4 thread group of instruction 2; In the ninth clock cycle T9, the fourth vector core computing unit 236-4 writes the operation result of performing the vector operation on the operands of the 4 / 4 thread group of instruction 2 back to the fourth register 234-4.
[0071] For instruction 3: In the third clock cycle T3, the instruction scheduling unit 210 sends instruction 3 to the instruction decoding unit 220; In the fourth clock cycle T4, the instruction decoding unit 220 decodes instruction 3 to obtain the operand addresses of the operands of instruction 3; In the fifth clock cycle T5, the first operand fetching unit 232-1 reads the operands of the 1 / 4 thread group of instruction 3 from the first register 234-1; In the sixth clock cycle T6, the first vector core calculation unit 236-1 performs a vector operation on the operands of 1 / 4 thread group of instruction 3. Meanwhile, under the read enable control of the first operand acquisition unit 232-1, the second operand acquisition unit 232-2 reads the operands of 2 / 4 thread group of instruction 3 from the second register 234-2; In the seventh clock cycle T7, the first vector core calculation unit 236-1 writes the operation result of the vector operation on the operands of 1 / 4 thread group of instruction 3 back to the first register 234-1. Meanwhile, the second vector core calculation unit 236-2 performs a vector operation on the operands of 2 / 4 thread group of instruction 3. Meanwhile, under the read enable control of the second operand acquisition unit 232-2, the third operand acquisition unit 232-3 reads the operands of 3 / 4 thread group of instruction 3 from the third register 234-3; In the eighth clock cycle T8, the second vector core calculation unit 236-2 writes the operation result of the vector operation on the operands of 2 / 4 thread group of instruction 3 back to the second register 234-2. Meanwhile, the third vector core calculation unit 236-3 performs a vector operation on the operands of 3 / 4 thread group of instruction 3. Meanwhile, under the read enable control of the third operand acquisition unit 232-3, the fourth operand acquisition unit 232-4 reads the operands of 4 / 4 thread group of instruction 3 from the fourth register 234-4; In the ninth clock cycle T9, the third vector core calculation unit 236-3 writes the operation result of the vector operation on the operands of 3 / 4 thread group of instruction 3 back to the third register 234-3. Meanwhile, the fourth vector core calculation unit 236-4 performs a vector operation on the operands of 4 / 4 thread group of instruction 3; In the tenth clock cycle T10, the fourth vector core calculation unit 236-4 writes the operation result of the vector operation on the operands of 4 / 4 thread group of instruction 3 back to the fourth register 234-4.
[0072] For instruction 4: In the fourth clock cycle T4, the instruction scheduling unit 210 sends instruction 4 to the instruction decoding unit 220; In the fifth clock cycle T5, the instruction decoding unit 220 decodes instruction 4 to obtain the operand addresses of the operands of instruction 4; In the sixth clock cycle T6, the first operand acquisition unit 232-1 reads the operands of 1 / 4 thread group of instruction 4 from the first register 234-1; In the seventh clock cycle T7, the first vector core computing unit 236-1 performs a vector operation on the operands of 1 / 4 thread groups of instruction 4. Meanwhile, under the read enable control of the first operand fetching unit 232-1, the second operand fetching unit 232-2 reads the operands of 2 / 4 thread groups of instruction 4 from the second register 234-2; In the eighth clock cycle T8, the first vector core computing unit 236-1 writes the operation result of performing the vector operation on the operands of 1 / 4 thread groups of instruction 4 back to the first register 234-1. Meanwhile, the second vector core computing unit 236-2 performs a vector operation on the operands of 2 / 4 thread groups of instruction 4. Meanwhile, under the read enable control of the second operand fetching unit 232-2, the third operand fetching unit 232-3 reads the operands of 3 / 4 thread groups of instruction 4 from the third register 234-3; In the ninth clock cycle T9, the second vector core computing unit 236-2 writes the operation result of performing the vector operation on the operands of 2 / 4 thread groups of instruction 4 back to the second register 234-2. Meanwhile, the third vector core computing unit 236-3 performs a vector operation on the operands of 3 / 4 thread groups of instruction 4. Meanwhile, under the read enable control of the third operand fetching unit 232-3, the fourth operand fetching unit 232-4 reads the operands of 4 / 4 thread groups of instruction 4 from the fourth register 234-4; In the tenth clock cycle T10, the third vector core computing unit 236-3 writes the operation result of performing the vector operation on the operands of 3 / 4 thread groups of instruction 4 back to the third register 234-3. Meanwhile, the fourth vector core computing unit 236-4 performs a vector operation on the operands of 4 / 4 thread groups of instruction 4; In the eleventh clock cycle T11, the fourth vector core computing unit 236-4 writes the operation result of performing the vector operation on the operands of 4 / 4 thread groups of instruction 4 back to the fourth register 234-4.
[0073] Here, the timing of the vector operations of the first vector core computing unit 236-1, the second vector core computing unit 236-2, the third vector core computing unit 236-3, and the fourth vector core computing unit 236-4 is controlled by the shared control information of the instruction decoding unit 220.
[0074] It can be seen that in the pipelined operation of such SIMT instructions, the instruction scheduling unit 210 only needs to send instructions to the only instruction decoding unit 220, and the instruction decoding unit 220 only needs to send the operand addresses to the vector core execution module 230 connected thereto, and the above operations can be repeated periodically to send and decode different instructions.
[0075] Using the solution of the present invention, the instruction scheduling unit of the vector computing device only needs to send instructions to a single instruction decoding unit, without the need to determine which vector core to schedule the instructions to, making the logic simpler. Moreover, by controlling the pipelining operations of multiple vector core execution modules through a single instruction decoding unit, the chip area can be significantly reduced. In addition, compared with existing vector computing devices, fewer selectors are required, which can further reduce the chip area. Multiple vector core execution modules fixedly execute different thread group parts of the same instruction, so that the power consumption is more balanced, avoiding the concentration of power consumption or heat, which is beneficial to heat dissipation.
[0076] Those of ordinary skill in the art should also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments of the present disclosure can be implemented as electronic hardware, computer software, or a combination of the two.
[0077] The foregoing description of the present disclosure is provided to enable any ordinary skill in the art to make or use the present disclosure. Various modifications to the present disclosure will be apparent to those of ordinary skill in the art, and the general principles defined herein can be applied to other variations without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope of the principles and novel features disclosed herein.
Claims
1. A vector computing device, comprising: an instruction scheduling unit, an instruction decoding unit connected to the instruction scheduling unit, and a plurality of vector core execution modules, wherein the plurality of vector core execution modules are connected in series and only a first vector core execution module is connected to the instruction decoding unit, wherein the instruction scheduling unit is configured to schedule single instruction multiple thread (SIMT) instructions and send the SIMT instructions to the instruction decoding unit, and the instruction decoding unit is configured to decode the SIMT instructions to determine the operand addresses of the operands of the SIMT instructions, and instruct the plurality of vector core execution modules to perform vector operations of the SIMT instructions on the operands respectively in a plurality of clock cycles.
2. The vector computing device according to claim 1, wherein the instruction decoding unit is configured to enable the first vector core execution module to read the operand from the operand address in one clock cycle, and the first vector core execution module is configured to enable the next vector execution module connected to the first vector core execution module to read the operand from the operand address in the next clock cycle of the clock cycle.
3. The vector computing device according to claim 1, wherein each vector core execution module comprises: an operand acquisition unit, a register, and a vector core calculation unit, wherein the operand acquisition unit reads the operand from the register based on the operand address from the instruction decoding unit, and sends the read operand to the vector core calculation unit, and the vector core calculation unit performs vector operations of the SIMT instructions using the operand and writes the operation result back to the register.
4. The vector computing device according to claim 3, wherein the vector core calculation units of each vector core execution module in the plurality of vector core execution modules respectively start the vector operations in the plurality of clock cycles in sequence based on shared control information from the instruction decoding unit.
5. The vector computing device according to claim 3, wherein each calculation unit in the vector core calculation unit is used to perform vector operations of one thread of the SIMT instructions.
6. The vector computing device according to claim 2, wherein the plurality of vector core execution modules comprise a first vector core execution module and a second vector core execution module, and the first vector core execution module comprises a first operand acquisition unit, a first register, and a first vector core calculation unit, the first operand acquisition unit reads the operands of the upper half thread group of the SIMT instructions from the first register based on the operand address from the instruction decoding unit, and sends the read operands to the first vector core calculation unit, and the first vector core calculation unit performs vector operations of the upper half thread group of the SIMT instructions using the operands of the upper half thread group and writes the operation result back to the first register; The second vector core execution module includes a second operand acquisition unit, a second register, and a second vector core calculation unit. The second operand acquisition unit reads the operands of the lower half thread group of the SIMT instruction from the second register based on the operand address from the first operand acquisition unit, and sends the read operands to the second vector core calculation unit. The second vector core calculation unit performs vector operations on the operands of the lower half thread group of the SIMT instruction by using the operands of the lower half thread group, and writes the operation result back to the second register.
7. The vector computing device according to claim 6, wherein the vector computing device is configured to: In the first clock cycle, the instruction scheduling unit sends the SIMT instruction to the instruction decoding unit; In the next clock cycle of the first clock cycle, i.e., the second clock cycle, the instruction decoding unit decodes the SIMT instruction to obtain the operand addresses of the operands of the SIMT instruction; In the next clock cycle of the second clock cycle, i.e., the third clock cycle, the first operand acquisition unit reads the operands of the upper half thread group of the SIMT instruction from the first register; In the next clock cycle of the third clock cycle, i.e., the fourth clock cycle, the first vector core calculation unit performs vector operations on the operands of the upper half thread group of the SIMT instruction. Meanwhile, under the read enable control of the first operand acquisition unit, the second operand acquisition unit reads the operands of the lower half thread group of the SIMT instruction from the second register; In the next clock cycle of the fourth clock cycle, i.e., the fifth clock cycle, the first vector core calculation unit writes the operation result of performing vector operations on the operands of the upper half thread group of the SIMT instruction back to the first register. Meanwhile, the second vector core calculation unit performs vector operations on the operands of the lower half thread group of the SIMT instruction; In the next clock cycle of the fifth clock cycle, i.e., the sixth clock cycle, the second vector core calculation unit writes the operation result of performing vector operations on the operands of the lower half thread group of the SIMT instruction back to the second register.
8. The vector computing device according to claim 2, wherein the multiple vector core execution modules include a first vector core execution module, a second vector core execution module, a third vector core execution module, and a fourth vector core execution module, and The first vector core execution module includes a first operand acquisition unit, a first register, and a first vector core calculation unit. The first operand acquisition unit reads the operands of the 1 / 4 thread group of the SIMT instruction from the first register based on the operand address from the instruction decoding unit, and sends the read operands to the first vector core calculation unit. The first vector core calculation unit performs vector operations on the operands of the 1 / 4 thread group of the SIMT instruction by using the operands of the 1 / 4 thread group, and writes the operation result back to the first register; The second vector core execution module includes a second operand acquisition unit, a second register, and a second vector core calculation unit. The second operand acquisition unit reads the operands of the 2 / 4 thread groups of the SIMT instruction from the second register based on the operand address from the first operand acquisition unit, and sends the read operands to the second vector core calculation unit. The second vector core calculation unit performs vector operations on the operands of the 2 / 4 thread groups of the SIMT instruction and writes the operation results back to the second register; The third vector core execution module includes a third operand acquisition unit, a third register, and a third vector core calculation unit. The third operand acquisition unit reads the operands of the 3 / 4 thread groups of the SIMT instruction from the third register based on the operand address from the second operand acquisition unit, and sends the read operands to the third vector core calculation unit. The third vector core calculation unit performs vector operations on the operands of the 3 / 4 thread groups of the SIMT instruction and writes the operation results back to the third register; and The fourth vector core execution module includes a fourth operand acquisition unit, a fourth register, and a fourth vector core calculation unit. The fourth operand acquisition unit reads the operands of the 4 / 4 thread groups of the SIMT instruction from the fourth register based on the operand address from the third operand acquisition unit, and sends the read operands to the fourth vector core calculation unit. The fourth vector core calculation unit performs vector operations on the operands of the 4 / 4 thread groups of the SIMT instruction and writes the operation results back to the fourth register.
9. The vector computing device according to claim 8, wherein the vector computing device is configured to: In a first clock cycle, the instruction scheduling unit sends the SIMT instruction to the instruction decoding unit; In the next clock cycle of the first clock cycle, i.e., the second clock cycle, the instruction decoding unit decodes the SIMT instruction to obtain the operand addresses of the operands of the SIMT instruction; In the next clock cycle of the second clock cycle, i.e., the third clock cycle, the first operand acquisition unit reads the operands of the 1 / 4 thread groups of the SIMT instruction from the first register; In the next clock cycle of the third clock cycle, i.e., the fourth clock cycle, the first vector core calculation unit performs vector operations on the operands of the 1 / 4 thread groups of the SIMT instruction. Meanwhile, under the read enable control of the first operand acquisition unit, the second operand acquisition unit reads the operands of the 2 / 4 thread groups of the SIMT instruction from the second register; In the next clock cycle of the fourth clock cycle, i.e., the fifth clock cycle, the first vector core computing unit writes back the operation result of performing vector operations on the operands of 1 / 4 thread groups of the SIMT instruction to the first register. Meanwhile, the second vector core computing unit performs vector operations on the operands of 2 / 4 thread groups of the SIMT instruction. Meanwhile, under the read enable control of the second operand acquisition unit, the third operand acquisition unit reads the operands of 3 / 4 thread groups of the SIMT instruction from the third register; In the next clock cycle of the fifth clock cycle, i.e., the sixth clock cycle, the second vector core computing unit writes back the operation result of performing vector operations on the operands of 2 / 4 thread groups of the SIMT instruction to the second register. Meanwhile, the third vector core computing unit performs vector operations on the operands of 3 / 4 thread groups of the SIMT instruction. Meanwhile, under the read enable control of the third operand acquisition unit, the fourth operand acquisition unit reads the operands of 4 / 4 thread groups of the SIMT instruction from the fourth register; In the next clock cycle of the sixth clock cycle, i.e., the seventh clock cycle, the third vector core computing unit writes back the operation result of performing vector operations on the operands of 3 / 4 thread groups of the SIMT instruction to the third register. Meanwhile, the fourth vector core computing unit performs vector operations on the operands of 4 / 4 thread groups of the SIMT instruction; In the next clock cycle of the seventh clock cycle, i.e., the eighth clock cycle, the fourth vector core computing unit writes back the operation result of performing vector operations on the operands of 4 / 4 thread groups of the SIMT instruction to the fourth register.
Citation Information
Patent Citations
Vector crossing multithread processing method and vector crossing multithread microprocessor
CN102156637A
Floating point processing method and system applied to vector operation, medium and equipment
CN116414460A
Instruction processing device and method, processor, electronic equipment and storage medium
CN118733115A
GPGPU (General Purpose Graphics Processing Unit)-based instruction execution acceleration method and GPGPU architecture
CN119537041A
Apparatus and method of optimising divergent processing in thread groups
US20240036874A1
Cited By
Vector kernel module of artificial intelligence chip and operation method thereof
CN120469721A