SIMD and SIMT Cooperative Computing Method, System and Device

By combining SIMD and SIMT technologies in the collaborative computing methods of the SIMT front-end and the SIMD back-end, the problems of the respective limitations of SIMD and SIMT are solved, and more efficient computing performance and more flexible development capabilities are achieved.

CN119536820BActive Publication Date: 2025-06-17SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510088871.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-17
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The respective limitations of SIMD and SIMT seriously limit further improvements in application performance, and also restrict developers' programming flexibility, making it difficult to choose the appropriate computing method according to actual needs.

Method used

By introducing a collaborative computing method between SIMD and SIMT in the thread bundle Warp composed of multiple stream processors SPs, the combination of the SIMT front-end and the SIMD back-end can be used to achieve efficient processing of computing tasks and flexible register allocation.

Benefits of technology

It significantly improves computing efficiency and performance, especially suitable for tasks such as graphics processing, scientific computing and deep learning, and improves application performance and development flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119536820B_ABST
    Figure CN119536820B_ABST
Patent Text Reader

Abstract

The present application provides a collaborative computing method, system and device based on SIMD and SIMT, belonging to the technical field of computer processing. The method is applied to a warp composed of multiple SPs, and the warp includes a SIMT front end and a SIMD back end; each SP corresponds to a vector register composed of multiple short registers; when the warp scheduler at the SIMT front end determines that the first warp corresponding to the computing task is in an active state, it generates fetch instruction information; the fetch unit obtains the first cache instruction corresponding to the first warp and sends it to the decoding unit at the SIMT front end for parsing; through the operand collection unit at the SIMD back end, one or more required operands are obtained; through the SIMD back end, each required operand is distributed to the corresponding computing management unit to execute the corresponding computing task, and the computing result corresponding to the computing task is written back to the storage unit through the write-back unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer processing technologies, and in particular, to a collaborative computing method, system, and device based on Single Instruction Multiple Data (SIMD) and Single Instruction Multiple Threads (SIMT). Background Art

[0002] In today's computing field, with the explosive growth of data volume and the continuous emergence of various complex application scenarios, the demand for computing performance is becoming increasingly urgent. Currently, the Central Processing Unit (CPU) and DSP chips generally adopt SIMD technology to achieve acceleration. By providing an Intrinsic programming interface, developers can directly operate the underlying registers, achieving performance improvement and power consumption reduction to a certain extent.

[0003] However, this SIMD-based acceleration method has significant limitations. It only supports being written with SIMD instructions and is unable to cope when facing tasks such as large-scale matrix operations that require highly parallel processing and complex calculation logics. Large-scale matrix operations often involve parallel processing of massive amounts of data and diverse calculation operations, and the limitations of SIMD instructions make it unable to efficiently meet these complex requirements.

[0004] On the other hand, Graphics Processing Unit (GPU), ASIC chips, etc. adopt SIMT technology to achieve acceleration, which can improve the program parallelism. However, its single matrix calculation unit cannot utilize wide registers to merge similar calculations like SIMD, which leads to the situation that in some scenarios with extremely high requirements for computing efficiency, the single calculation unit of SIMT cannot fully exert the potential of the hardware, thus affecting the overall computing performance.

[0005] The inventors found that if the overall processing ability of the SIMT matrix calculation unit can be utilized while introducing SIMD technology to give full play to its advantages in vectorized operations and achieve an organic combination of the two, the processing ability of the hardware can be effectively improved, and further the performance and development flexibility of application programs can be significantly improved.

[0006] Based on this, there is an urgent need for a technical solution that can combine SIMD and SIMT to improve the performance and development flexibility of application programs. Summary of the Invention

[0007] The embodiments of the present application provide a collaborative computing method, system and device based on SIMD and SIMT, which are used to solve the technical problems that the respective limitations of SIMD and SIMT severely limit the further improvement of application program performance, and also greatly restrict the flexibility of developers in the programming process, making it difficult for them to flexibly select appropriate computing methods according to actual needs.

[0008] On the one hand, the embodiments of the present application provide a collaborative computing method based on SIMD and SIMT. The method is applied to a thread bundle Warp composed of multiple stream processors SP. The Warp includes a SIMT front end and a SIMD back end; each SP corresponds to a vector register composed of multiple short registers; the method includes:

[0009] When the thread bundle scheduler at the SIMT front end determines that the first thread bundle corresponding to the computing task is in an active state, fetch instruction indication information is generated; wherein, the active state is determined based on the execution dependent resources of the computing task in the task queue state;

[0010] In response to the fetch instruction indication information, a first cache instruction corresponding to the first thread bundle is obtained through a fetch unit and sent to a decoding unit at the SIMT front end for parsing, so as to dispatch the parsed instruction to the SIMD back end;

[0011] One or more required operands corresponding to the parsed instruction are obtained through an operand collection unit at the SIMD back end;

[0012] Through the SIMD back end, each of the required operands is distributed to a corresponding computing management unit to execute the corresponding computing task, and the computing result corresponding to the computing task is written back to a storage unit through a write-back unit.

[0013] In an implementation manner of the present application, the multiple short register combinations are obtained based on the bit width of the computing task; specifically including:

[0014] According to the bit width of the computing task, the corresponding required register length is determined;

[0015] According to the required register length and the current available state of the register, a candidate register combination in a preset available register set is determined; wherein, the preset available register set contains several 32-bit registers;

[0016] According to each of the candidate register combinations and a preset register cost function, a selected register combination is determined, so as to obtain the vector register according to the selected register combination.

[0017] In an implementation of the present application, determining a selected register combination according to each of the to-be-selected register combinations and a preset register cost function specifically includes:

[0018] Performing an optimization operation on the preset register cost function through a gradient descent algorithm; wherein, the preset register cost function at least includes a time delay cost and an energy consumption cost corresponding to the corresponding register combination;

[0019] When the gradient norm corresponding to the gradient descent algorithm is less than a preset threshold, stop the iteration and determine the corresponding to-be-selected register combination as the selected register combination.

[0020] In an implementation of the present application, the instruction cache unit of the first cache instruction includes vector instructions added by RISC-V; the vector instructions include: configuration setting instructions, vector load and store instructions, vector integer instructions, vector fixed-point instructions, and vector floating-point instructions.

[0021] In an implementation of the present application, the Warp corresponds to 7 CSR registers, including: vector start position register, fixed-point saturation flag register, fixed-point rounding mode register, vector control and status register, vector length register, vector data type register, and vector register length register.

[0022] In an implementation of the present application, the method further includes:

[0023] Before adding vector instructions by RISC-V, define a front-end vector operation interface for front-end instructions in the vector instruction description file of RISC-V through a tablegen tool; the front-end vector operation interface is at least used for vector addition, vector multiplication, and vector loading;

[0024] Store the interface header file corresponding to the front-end vector operation interface into a preset tool chain, so as to call the encapsulated vector instructions through the preset tool chain for RISC-V vector programming.

[0025] In an implementation of the present application, provide an Intrinsic interface; the Intrinsic interface is used to perform vectorization acceleration by using hardware register resources when writing device-side code at the upper layer.

[0026] In an implementation of the present application, the method further includes:

[0027] According to the functional category and operation complexity of the vector instruction and the preset opcode subspace comparison table, map the opcode of the corresponding vector instruction to the opcode space of the RISC-V basic instruction format; wherein, the preset opcode subspace comparison table includes the correspondence between the opcodes of different functional categories and different operation complexities and different subspaces.

[0028] On the other hand, an embodiment of the present application further provides a collaborative computing system based on SIMD and SIMT. The system is applied to a warp composed of multiple stream processors SP. The warp includes a SIMT front end and a SIMD back end; each SP corresponds to a vector register composed of multiple short register combinations; the system includes:

[0029] A generation module, configured to generate fetch indication information when the thread bundle scheduler at the SIMT front end determines that the first thread bundle corresponding to the computing task is in an active state; wherein, the active state is determined based on the execution dependent resources of the computing task in the task queue state.

[0030] An acquisition and sending module, configured to, in response to the fetch indication information, obtain the first cache instruction corresponding to the first thread bundle through the fetch unit and send it to the decoding unit at the SIMT front end for parsing, so as to dispatch the parsed instruction to the SIMD back end.

[0031] An acquisition module, configured to obtain one or more required operands corresponding to the parsed instruction through the operand collection unit at the SIMD back end.

[0032] A distribution module, configured to distribute each of the required operands to the corresponding computing management unit through the SIMD back end to execute the corresponding computing task, and write back the computing result corresponding to the computing task to the storage unit through the write-back unit.

[0033] On yet another aspect, an embodiment of the present application further provides a collaborative computing device based on SIMD and SIMT. The device includes:

[0034] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a collaborative computing method based on SIMD and SIMT as described above.

[0035] Compared with the prior art, the present application has the following remarkable effects:

[0036] (1) Combine SIMD and SIMT, integrating the single-core vectorization processing ability of SIMD and the advantages of multi-core parallel execution of SIMT. It can make full use of the multi-thread management of SIMT and the data parallelism of SIMD to improve computing efficiency and performance, especially suitable for tasks with a large amount of parallelism and data-level parallelism such as graphics processing, scientific computing, and deep learning, significantly enhancing the performance of application programs. At the same time, configure a flexibly combinable vector register for each SP, optimizing register allocation and improving resource utilization.

[0037] (2) This application also adds vector instructions based on RISC-V, updates the compiler and tool chain, and provides an Intrinsic interface to facilitate developers to utilize the hardware vectorization ability and enhance development flexibility. Description of the Drawings

[0038] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0039] Figure 1 It is a schematic flow diagram of a collaborative computing method based on SIMD and SIMT in an embodiment of the present application;

[0040] Figure 2 It is a schematic diagram of the Warp architecture in a collaborative computing method based on SIMD and SIMT in an embodiment of the present application;

[0041] Figure 3 It is a schematic diagram of register combination in a collaborative computing method based on SIMD and SIMT in an embodiment of the present application;

[0042] Figure 4 It is a schematic flow diagram of the modified setting process for adding vector instructions in a collaborative computing method based on SIMD and SIMT in an embodiment of the present application;

[0043] Figure 5 It is another schematic flow diagram of a collaborative computing method based on SIMD and SIMT in an embodiment of the present application;

[0044] Figure 6 It is a schematic structural diagram of a collaborative computing system based on SIMD and SIMT in an embodiment of the present application;

[0045] Figure 7 It is a schematic structural diagram of a collaborative computing device based on SIMD and SIMT in an embodiment of the present application. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments and corresponding drawings of this application. Apparently, the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0047] The embodiments of this application provide a collaborative computing method, system, and device based on SIMD and SIMT to solve the technical problem that the respective limitations of SIMD and SIMT severely restrict the further improvement of application program performance, and at the same time greatly restrict the flexibility of developers in the programming process, making it difficult for them to flexibly select appropriate computing methods according to actual needs.

[0048] The following will describe each embodiment of this application in detail with reference to the drawings.

[0049] The embodiments of this application provide a collaborative computing method based on SIMD and SIMT. This method is applied to a warp composed of multiple stream processors (SPs). The warp includes a SIMT front end and a SIMD back end. Each SP corresponds to a vector register composed of multiple short registers. As Figure 1 shown, this method may include steps S101 - S104:

[0050] S101, when the thread block scheduler at the SIMT front end determines that the first thread block corresponding to the computing task is in an active state, generate instruction fetch indication information.

[0051] Among them, the active state is determined based on the execution - dependent resources of the computing tasks in the task queue state.

[0052] It should be noted that the GPU can be used as the execution entity of the collaborative computing method based on SIMD and SIMT, which is only an exemplary existence. The execution entity is not limited to the GPU, and this application does not make specific limitations in this regard.

[0053] In the embodiments of the present application, when the GPU executes a computing task, the warp scheduler in the SIMT front-end of a warp can constantly monitor the status of each warp. Among them, the warp scheduler maintains a task queue, which stores the computing tasks waiting to be executed. When a new computing task is submitted to the task queue, the warp scheduler will first check the status of the task queue. If there are unprocessed tasks in the task queue and the currently executed dependent resources (such as a sufficient number of free registers for storing operands) can meet the basic requirements for task execution, the warp scheduler will mark the first warp related to the task as activatable. For example, in a task of parallel computing matrix multiplication, after the task queue receives the matrix multiplication task, the scheduler will evaluate whether there are sufficient free computing units and memory bandwidth to support the execution of the warp. If so, the warp responsible for matrix multiplication calculation will be set to the active state. Subsequently, the warp scheduler will send an instruction to the fetch unit, and the fetch unit will fetch instructions from the instruction cache according to the fetch instruction information.

[0054] Among them, the SIMT front-end of the present application includes three stages: fetching, decoding, and dispatching, and the SIMD includes three stages: operand collection, execution, and data write-back. The architecture diagram corresponding to the warp is as Figure 2 shown. The SIMT front-end includes Fetch, Instruction buffer, Decode, Score board, Select, and SIMT Engine; the SIMD back-end includes Operand Collector, Register, Special Function Unit (SFU), Arithmetic Logic Unit (ALU), Load Store Unit (LSU); Instruction Cache (I-Cache); Data Cache (D-Cache)

[0055] The above-mentioned instruction fetching is responsible for obtaining instructions, which is the starting point of the entire process, fetching the instructions to be executed from memory; the instruction buffer is used to temporarily store the fetched instructions for subsequent processing; the decoding unit decodes the instructions in the instruction buffer to make them into executable operations; the scoreboard is used to record the execution status and resource usage of each instruction to avoid conflicts and ensure the correct execution order; the selection unit selects the next instruction to be executed according to information such as the scoreboard; the SIMT engine is the core of the entire SIMT part, responsible for managing and executing instructions of multiple threads to achieve parallel computing; the operand collector is used to collect the operands required for executing instructions; the register is used to temporarily store data and intermediate results and is a storage unit with fast access in the processor; the special function unit is used to execute some special mathematical operations such as trigonometric functions and logarithms; the arithmetic logic unit executes basic arithmetic and logical operations such as addition, subtraction, AND, and OR; the load / store unit is responsible for the loading and storing operations of data between memory and registers; the instruction cache is used to cache the fetched instructions to improve the speed of instruction fetching and reduce the number of accesses to memory; the data cache is used to cache frequently used data to improve the speed of data access and reduce the access latency to memory.

[0056] An algorithm calculation process based on the above architecture diagram is as follows: The SIMT front end fetches instructions from the I-Cache by the Fetch module and sends them to the Decode module for decoding, and sets the flag bits corresponding to the source and destination registers in the Score module. At the same time, the control information is stored in the I-Buffer. The Select module and the SIMT Engine module detect the control information in the I-Buffer and distribute it to the Operand Collector modules of different channels. Then, according to the type of instruction and other situations, the register file is determined and relevant register operations are performed. Finally, the SFU, ALU, LSU and other modules are called according to the different instructions.

[0057] In another embodiment of the present application, the above-mentioned multiple short register combinations are obtained based on the bit width of the computing task, specifically including:

[0058] According to the bit width of the computing task, the corresponding required register length is determined. According to the required register length and the current available status of the registers, the candidate register combinations in the preset available register set are determined. Among them, the preset available register set contains several 32-bit registers. According to each candidate register combination and the preset register cost function, the selected register combination is determined to obtain the vector register according to the selected register combination.

[0059] Generally understood, for example, for a computing task that requires a register length of 256 bits, 8 registers with a length of 32 bits are combined; for a computing task that requires a register length of 128 bits, 4 registers with a length of 32 bits are combined; for a computing task that requires a register length of 64 bits, 2 registers with a length of 32 bits are combined; and for a register length of 32 bits, 1 register with a length of 32 bits is used. The register combination is as Figure 3 shown. To effectively select the register combination, the present application comprehensively considers the following factors: the bit-width requirement of the computing task, that is, the register length required for each computing task. The performance optimization goal, considering how to combine registers to achieve the highest throughput and the lowest latency of the computing task. The current available resource status, that is, the number and usage of existing registers in the system.

[0060] The present application defines the following variables which represents the set of all available registers in the system, which represents the set of all computing tasks to be processed, which represents the computing task the required register length, which represents the register the current allocation status (0 means not allocated, 1 means allocated), which represents the task using the register the preset register cost function, which is an overall evaluation of time delay and energy consumption, where represents the cost brought by time delay, represents the cost brought by energy consumption. These two goals are usually conflicting. For example, in some cases, longer registers may be needed to reduce time delay, but using longer registers will increase energy consumption.

[0061] Furthermore, based on the above-mentioned candidate register combinations and the preset register cost function, the selected register combination is determined, specifically including:

[0062] Through the gradient descent algorithm, an optimization operation is performed on the preset register cost function. Among them, the preset register cost function at least includes the time delay cost and the energy consumption cost corresponding to the corresponding register combination. When the gradient norm corresponding to the gradient descent algorithm is less than the preset threshold, the iteration is stopped, and the corresponding candidate register combination is determined as the selected register combination.

[0063] That is to say, the present application uses the gradient descent algorithm to optimize the above-mentioned preset register cost function. The specific steps are as follows: Set the objective function as , select an initial register allocation scheme, that is, and It has been determined that this set of parameters is used as the initial value and substituted into the calculation of the objective function score , and calculate the gradient function . When the gradient is small enough, stop the iteration, and at this time, a set of local minima is found; otherwise, obtain Perform iteration: , , let , Iterate in the opposite direction of the gradient

[0064] Through the above solution, a flexibly combinable vector register is configured for each SP, and the register allocation is optimized through comprehensive factor consideration and the gradient descent algorithm to improve the utilization rate of register resources

[0065] S102. In response to the fetch instruction information, the fetch unit obtains the first cache instruction corresponding to the first warp and sends it to the decoding unit of the SIMT front end for parsing, so as to dispatch the parsed instruction to the SIMD back end

[0066] That is to say, after the decoding unit completes the decoding of the instruction, it sends the relevant information of the instruction (such as the operation code, operand address, etc.) to the dispatch unit, and the dispatch unit decides when to dispatch the instruction to the SIMD back end for execution according to the information provided by modules such as the scoreboard

[0067] Among them, the instruction cache unit of the first cache instruction includes vector instructions added by RISC-V. The vector instructions include: configuration setting instructions, vector load and store instructions, vector integer instructions, vector fixed-point instructions, and vector floating-point instructions

[0068] For example, the configuration setting instructions: vsetvli, vsetivli, vsetvl. The vector load and store instructions: vle <eew>.v, vlm.v, vlseg <nf> e <eew>.v, vlsseg <nf> e <eew>.v, vl <nf> re <eew>.v, vse <eew>.v, vsm.v, vsseg <nf> e <eew>.v, vs <nf>r.v. Among them, eew represents the bit width of the register, and multiple values such as 32, 64, 128, 256, 512, etc. can be specified. ".v" indicates a vectorized instruction. Vector integer instructions: vadd, vsub, vminu, vmin, vmaxu, vmax, vand, vor, vxor, vadc, vmadc, vsbc, vmseq, vmsne, vmsleu, vmsle, vmsgtu, vmsgt, vdivu, vdiv, vremu, vrem, vsll, vmul, vsrl, vsra, vmadd, vnmsub, vmacc, vnmsac, vwaddu, vwadd, vwsubu, vwsub, vwaddu.w, vwadd.w, vwsubu.w, vwsub.w, vwmulu, vwmul. Vector fixed-point instructions: vaaddu, vaadd, vasubu, vasub, vsaddu, vsadd, vssubu, vssub, vsmulu, vssrl, vssra. Vector floating-point instructions: vfadd, vfsub, vfmin, vfmax, vfcvt, vfwcvt, vfncvt, vfsqrt, vmfeq, vmfle, vmflt, vmfne, vmfgt, vmfge, vfdiv, vfmadd, vfnmadd, vfmsub, vfmacc, vfnmacc, vfwadd, vfwsub, vfwadd.w, vfwsub.w, vfwmul.

[0069] That is to say, when this application supports the cooperative computing combining SIMD and SIMT, new instructions can be pre-included in the instruction set, including the encoding, opcode, register usage, operands, and execution effects of the instructions. Specifically, RISC-V is used to implement the addition of vector instructions. Warp corresponds to 7 CSR registers, including: vector start position register, fixed-point saturation flag register, fixed-point rounding mode register, vector control and status register, vector length register, vector data type register, vector register length register. They are vstart, vxsat, vxrm, vscr, vl, vtype, vlenb respectively. vstart is the vector start position, vxsat is the fixed-point saturation flag, vxm is the fixed-point rounding mode, vscr is the vector control and status register, vl is the vector length, vtype is the vector data type register, and vlenb is the byte length of the vector register.

[0070] Among them, vl and vtype are the registers most frequently involved by users. vl calculates the instruction vector length through vset(i)vl(i), and vtype represents the element attributes in each instruction with several fields including vill, vma, vta, vlmul, and vsew. vill is an illegal identifier. If its value is 1, the other fields are invalid; vma is a mask bit; the vmul field represents the vector register grouping, which can be a fraction or an integer. The integer n means that n registers are grouped for calculation, and the fraction 1 / n means that 1 register is split into n registers for calculation; vsew represents the element width.

[0071] When adding vector instructions in this application, it includes modifications to both the hardware part and the software part, such as Figure 4 shown, including the hardware part: determining the hardware operation type, customizing instructions, modifying Verilog, modifying macro definitions, modifying the decoding module, modifying the hardware logic, adjusting the data and control links, and instruction simulation; the software part: customizing instructions, opcode definition, compiler and toolchain modification, Intrinsic interface, and testing.

[0072] Specifically, start: the starting point of the process. Determine the hardware operation type: First, it is necessary to clarify the operation type that the hardware needs to perform, which is the basis and foundation for subsequent work. Custom instruction: According to the determined hardware operation type, design custom instructions to meet specific computing requirements. Modify Verilog: Verilog is a hardware description language, and here it needs to be modified to implement the functions of custom instructions in hardware. Adjust the data and control links: Adjust the data transmission path and the control signal transmission path in the hardware to ensure that the custom instructions can correctly process data and control the operation of the hardware. Instruction simulation: Simulate the custom instructions to verify whether their functions and performance in hardware meet the expectations. Modify macro definitions, modify the decoding module, and modify the hardware logic: These steps are to modify and optimize some underlying definitions, modules, and logics during the hardware design process to support the implementation of custom instructions and the overall hardware functions. Software part; Custom instruction: Corresponding to the custom instructions in the hardware part, the software part also needs to define these instructions so that the software can work in coordination with the hardware. Opcode definition: Define the opcode for the custom instructions. The opcode is an important part of the computer instruction system and is used to identify the operation type and function of the instruction. Modify the compiler and toolchain: Since custom instructions are introduced, it is necessary to modify the compiler and related toolchains so that they can recognize and process these new instructions and correctly compile the programs written in high-level languages into code that can run on the hardware. Intrinsic interface: The Intrinsic interface is a programming interface used to directly call the functions of the underlying hardware in high-level languages. Here, corresponding settings and modifications are required so that the software can efficiently utilize the custom instructions of the hardware through the Intrinsic interface. Test: Test the software part to ensure that the software can correctly interact with the hardware and that the custom instructions can work properly at the software level and meet the expected performance and function requirements. End: The end point of the entire process, indicating that the development or improvement work of the hardware and software parts is completed.

[0073] Through the above solution, the development and optimization of a computing system or chip with custom instruction functions are achieved.

[0074] In addition, before adding vector instructions to RISC-V in this application, the front-end vector operation interface of the front-end instructions is defined in the vector instruction description file of RISC-V through the tablegen tool. The front-end vector operation interface is at least used for vector addition, vector multiplication, and vector loading. Store the interface header file corresponding to the front-end vector operation interface in the preset toolchain so that the encapsulated vector instructions can be called through the preset toolchain for RISC-V vector programming.

[0075] The front-end interface is generated using tablegen. The front-end instruction generation built-in interface is defined in riscv_vector.td. The interface header file will be located in lib / clang / 12.0.1 / include / riscv_vector.h of the toolchain. This file contains all the front-end built-in interfaces of RISC-V Vector. By using LLVM compilation optimization, it is possible to optimize the potentially vectorized code in the code and use the optimized instruction set for performance acceleration.

[0076] Through the above solution, the compiler and toolchain can be updated to enable them to recognize and generate code containing new instructions.

[0077] In addition, when using compiler optimization, on the one hand, the upper layer cannot be optimized in the way expected by programmers, resulting in poor flexibility and generality. On the other hand, since the compiler cannot accelerate in every case, the performance improvement brought by optimization is limited, and there are many limitations in applying relevant optimizations. Therefore, this application also provides an Intrinsic interface; the Intrinsic interface is used to utilize the hardware register resources for vectorization acceleration when writing device-side code at the upper layer.

[0078] In another embodiment of this application, to better improve the execution efficiency of vector instructions, the following embodiments are also provided:

[0079] According to the functional category and operation complexity of the vector instruction, and the preset opcode subspace comparison table, map the opcode of the corresponding vector instruction to the opcode space of the RISC-V basic instruction format. Among them, the preset opcode subspace comparison table includes the correspondence between opcodes of different functional categories and different operation complexities and different subspaces.

[0080] For example, for vector configuration instructions (such as vsetvli, etc.), the opcode is mapped to a specific high-order opcode subspace. Among them, according to different vector configuration parameters, the lower bits of the opcode further refine the opcode information, which not only ensures the uniqueness of the opcode but also realizes the logical grouping of functions. For vector arithmetic instructions (such as vadd, vsub, etc.), the opcode is mapped to another subspace, and different arithmetic operations and data types are distinguished by the lower bits of the opcode. The above technical solution improves the scalability of the instructions. When new vector instructions need to be added, only the opcode needs to be extended within the corresponding functional category subspace, without the need to readjust the entire opcode layout, reducing the hardware and software modification costs during instruction set extension and improving the maintainability of the instruction set.

[0081] S103, through the operand collection unit of the SIMD backend, obtain one or more required operands corresponding to the parsed instruction.

[0082] Among them, when the instruction is ready to be executed (for example, conditions such as the operands being ready are met), the dispatch unit notifies the operand collection unit to start collecting the operands required by the instruction. For example, after an addition instruction is dispatched, the operand collection unit obtains two operands from the register file / RAM and uses them for the addition operation.

[0083] S104, through the SIMD backend, distributes each required operand to the corresponding computing management unit to execute the corresponding computing task, and writes back the computing result corresponding to the computing task to the storage unit through the write-back unit.

[0084] After the operand collection unit collects the required operands, it sends these operands to the cross-data distribution module, and the cross-data distribution module then distributes the operands to the integer computing unit, FPU computing unit, etc. for parallel computing to execute the computing task. The write-back unit is responsible for storing the result in a specified location for subsequent instructions to use.

[0085] This application combines the SIMD instructions of chips such as the CPU and the working modes such as SIMT of the GPU, fully utilizes the vector processing ability of SIMD in single-core processing and the advantages of multi-core parallel execution of SIMT, can greatly improve the performance of application programs, adds vector instructions based on RISC-V to support the compiler to analyze and optimize potential code, and provides an Intrinsic interface method for developers to flexibly apply the hardware vector processing ability to accelerate the performance of application programs.

[0086] Figure 5 Another flowchart of a collaborative computing method based on SIMD and SIMT provided by an embodiment of this application is as Figure 5 shown, including the instruction fetch stage, decoding stage, and dispatch stage executed by the SIMT front end; the operand collection stage, execution stage, and write-back stage executed by the SIMD backend.

[0087] Taking the execution of a vector addition computing task as an example:

[0088] In the instruction fetch stage: Instruction cache: Assume that the vector addition instruction to be executed has been stored in the instruction cache. Thread block scheduler: Since this is a simple computing task, perhaps only one thread block (such as W0) is in the active state, and the corresponding program counter PC0 points to the position of the vector addition instruction in the instruction cache. Instruction fetch unit: According to the indication of the thread block scheduler, fetches the vector addition instruction from the instruction cache.

[0089] Decode Stage: Decoding Unit: Decodes the fetched vector addition instruction, determines that this is an operation involving the addition of two vectors, and information such as the addresses of the operands; Instruction Buffer: The decoded instruction is temporarily stored in the instruction buffer. Scoreboard: Records the relevant information of the instruction, for example, it needs to read the operands of two vectors and is currently in a state of waiting for operands. Dispatch Unit: Since the scoreboard indicates that the instruction is not ready to execute (lacking operands), it is not dispatched temporarily.

[0090] Operand Collection Stage: Register File / RAM: Assume that the two vectors to be added are already stored in the register file or RAM. Round-Robin Arbiter RR: When the dispatch unit of the SIMT front-end is ready to dispatch an instruction, multiple potentially concurrent operand requests (including other possible instructions) are determined by the round-robin arbiter RR to decide who accesses the register file / RAM first. Operand Collection Unit: Assume that after the round-robin arbitration, the vector addition instruction obtains the access right, and the operand collection unit collects the corresponding elements of the two vectors from the register file / RAM as operands.

[0091] Execution Stage: Cross Data Distribution: Distributes the operands of the two collected vectors to the integer computing unit (assuming the vector elements are integers). Integer Computing Unit: Performs addition operations on the corresponding elements of the two vectors, for example, adding the first element of the first vector to the first element of the second vector, and so on, to achieve parallel computing. Branch Management Unit: In this simple vector addition task, branch operations may not be required, so the branch management unit does not play a role temporarily.

[0092] Write-back Stage: Write-back Unit: Writes the result calculated by the integer computing unit, that is, the new vector after adding the two vectors, back to the register file / RAM to update the data.

[0093] Through the instruction fetching, decoding, and scheduling of the SIMT front-end and the efficient data parallel processing of the SIMD back-end above, a simple vector addition calculation task is completed. In practical applications, there may be more complex instructions and more warps working simultaneously. This architecture can make full use of the multi-thread management of SIMT and the data parallelism of SIMD to improve the computing efficiency and performance, especially suitable for tasks with a large amount of parallelism and data-level parallelism such as graphics processing, scientific computing, and deep learning.

[0094] Through the above solution, by combining SIMD and SIMT, the advantages of SIMD's single-core vectorization processing ability and SIMT's multi-core parallel execution are integrated. It can make full use of SIMT's multi-thread management and SIMD's data parallelism ability to improve computing efficiency and performance, and is especially suitable for tasks with a large amount of parallelism and data-level parallelism such as graphics processing, scientific computing, and deep learning, significantly enhancing the performance of the application program. At the same time, a vector register with a flexible combination is configured for each SP, optimizing the register allocation and improving resource utilization.

[0095] In addition, this application also adds vector instructions based on RISC-V, updates the compiler and tool chain, and provides an Intrinsic interface to facilitate developers to utilize the hardware vectorization ability and improve development flexibility.

[0096] Figure 6 The following is a schematic structural diagram of a collaborative computing system based on SIMD and SIMT provided by an embodiment of this application. As Figure 6 shown, the system is applied to a thread bundle Warp composed of multiple stream processors SP. The Warp includes a SIMT front end and a SIMD back end. Each SP corresponds to a vector register composed of multiple short registers. The collaborative computing system 600 based on SIMD and SIMT includes:

[0097] A generation module 601, configured to generate fetch instruction information when the thread bundle scheduler at the SIMT front end determines that the first thread bundle corresponding to the computing task is in an active state. Wherein, the active state is determined based on the execution-dependent resources of the computing task in the task queue state. An acquisition and sending module 602, configured to, in response to the fetch instruction information, acquire the first cache instruction corresponding to the first thread bundle through the fetch unit and send it to the decoding unit at the SIMT front end for parsing, so as to dispatch the parsed instruction to the SIMD back end. An acquisition module 603, configured to acquire one or more required operands corresponding to the parsed instruction through the operand collection unit at the SIMD back end. A distribution module 604, configured to distribute each required operand to the corresponding computing management unit through the SIMD back end to execute the corresponding computing task, and write back the computing result corresponding to the computing task to the storage unit through the write-back unit.

[0098] Figure 7 The following is a schematic structural diagram of a collaborative computing device based on SIMD and SIMT provided by an embodiment of this application. As Figure 7 shown, the device includes:

[0099] At least one processor; and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute:

[0100] When the warp scheduler at the SIMT front-end determines that the first warp corresponding to the computing task is in an active state, fetch indication information is generated. Among them, the active state is determined based on the execution-dependent resources of the computing task in the task queue state. In response to the fetch indication information, the first cache instruction corresponding to the first warp is obtained through the fetch unit and sent to the decoding unit at the SIMT front-end for parsing, so as to dispatch the parsed instruction to the SIMD back-end. Through the operand collection unit at the SIMD back-end, one or more required operands corresponding to the parsed instruction are obtained. Through the SIMD back-end, each required operand is distributed to the corresponding computing management unit to execute the corresponding computing task, and the computing result corresponding to the computing task is written back to the storage unit through the write-back unit.

[0101] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system and device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0102] The systems and devices provided in the embodiments of this application correspond one-to-one with the methods. Therefore, the systems and devices also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and devices will not be elaborated here.

[0103] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.

[0104] The above description is only for the embodiments of this application and is not intended to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.< / nf> < / eew> < / nf> < / eew> < / eew> < / nf> < / eew> < / nf> < / eew> < / nf> < / eew>

Claims

1. A collaborative computing method based on SIMD and SIMT, characterized in that: The method is applied to a thread bundle Warp composed of multiple stream processors SP, the Warp includes a SIMT front end and a SIMD back end; each SP corresponds to a vector register composed of multiple short registers; the method includes: When the thread warp scheduler of the SIMT front end determines that the first thread warp corresponding to the computing task is in an activated state, generating instruction fetch indication information; wherein the activation state is determined based on the execution dependent resource of the computing task in the task queue state; In response to the instruction fetch indication information, obtaining a first cache instruction corresponding to the first thread warp through an instruction fetch unit, and sending the first cache instruction to a decoding unit of the SIMT front end for parsing, so as to dispatch the parsed instruction to a SIMD back end; Acquire one or more required operands corresponding to the parsed instruction through the operand collection unit of the SIMD backend; Distribute each of the required operands to a corresponding computing management unit through the SIMD backend to execute the corresponding computing task, and write the computing result corresponding to the computing task back to the storage unit through the write-back unit; The multiple short register combinations are obtained based on the bit width of the computing task; specifically including: Determining the corresponding required register length according to the bit width of the computing task; According to the required register length and the current register available state, determine the register combination to be selected in the preset available register set; wherein the preset available register set includes a plurality of 32-bit registers; for a computing task that needs to use a register length of 256 bits, eight register combinations with a length of 32 bits are used; for a computing task that needs to use a register length of 128 bits, four register combinations with a length of 32 bits are used; for a computing task that needs to use a register length of 64 bits, two register combinations with a length of 32 bits are used; and for a register length of 32 bits, one register with a length of 32 bits is used; According to each of the candidate register combinations and the preset register cost function, a selected register combination is determined to obtain the vector register according to the selected register combination; wherein the following variables are defined Represents the set of all available registers in the system, represents the set of all pending computing tasks, Represents a computing task The required register length, Indicates register The current allocation status of Indicates the task Using Registers The preset register cost function is an overall evaluation of time delay and energy consumption, ,in represents the cost of time delay, Represents the cost caused by energy consumption; Wherein, determining the selected register combination according to each of the candidate register combinations and the preset register cost function specifically includes: By using a gradient descent algorithm, an optimization operation is performed on the preset register cost function; wherein the preset register cost function at least includes a time delay cost and an energy consumption cost corresponding to a corresponding register combination; When the gradient modulus corresponding to the gradient descent algorithm is less than a preset threshold, the iteration is stopped, and the corresponding candidate register combination is determined to be the selected register combination; Wherein, the instruction cache unit of the first cache instruction includes vector instructions added by RISC-V; the vector instructions include: configuration setting instructions, vector load and store instructions, vector integer instructions, vector fixed-point instructions, and vector floating-point instructions; Wherein, the Warp corresponds to 7 CSR registers, including: vector starting position register, fixed-point saturation flag register, fixed-point rounding mode register, vector control and status register, vector length register, vector data type register, vector register length register; Wherein, the method further comprises: According to the functional category and operation complexity of the vector instruction and a preset operation code subspace comparison table, the operation code of the corresponding vector instruction is mapped to the operation code space of the RISC-V basic instruction format; wherein the preset operation code subspace comparison table includes the comparison relationship between operation codes of different functional categories and different operation complexities and different subspaces.

2. The collaborative computing method based on SIMD and SIMT according to claim 1, characterized in that: The method further comprises: Before adding vector instructions using RISC-V, a front-end vector operation interface of the front-end instruction is defined in a vector instruction description file of RISC-V by using a tablegen tool; the front-end vector operation interface is at least used for vector addition, vector multiplication and vector loading; The interface header file corresponding to the front-end vector operation interface is stored in a preset tool chain so that the encapsulated vector instructions can be called through the preset tool chain to perform RISC-V vector programming.

3. The collaborative computing method based on SIMD and SIMT according to claim 1, characterized in that: An Intrinsic interface is provided; the Intrinsic interface is used to utilize hardware register resources for vectorization acceleration when writing device-side code at an upper layer.

4. A collaborative computing system based on SIMD and SIMT, characterized in that: The system is applied to a thread bundle Warp composed of multiple stream processors SP, the Warp includes a SIMT front end and a SIMD back end; each SP corresponds to a vector register composed of multiple short registers; the system includes: A generating module, configured to generate instruction fetch indication information when a thread warp scheduler of a SIMT front end determines that a first thread warp corresponding to a computing task is in an activated state; wherein the activation state is determined based on an execution dependent resource of the computing task in a task queue state; an acquisition and sending module, configured to, in response to the instruction fetch indication information, acquire the first cache instruction corresponding to the first thread warp through an instruction fetch unit, and send the first cache instruction to the decoding unit of the SIMT front end for parsing, so as to dispatch the parsed instruction to the SIMD back end; An acquisition module, configured to acquire one or more required operands corresponding to the parsed instruction through an operand collection unit of the SIMD backend; A distribution module, used to distribute each of the required operands to a corresponding calculation management unit through the SIMD backend to execute the corresponding calculation task, and write the calculation result corresponding to the calculation task back to the storage unit through a write-back unit; The plurality of short register combinations are obtained based on the bit width of the computing task; and the system further comprises: Determining the corresponding required register length according to the bit width of the computing task; According to the required register length and the current register available state, determine the register combination to be selected in the preset available register set; wherein the preset available register set includes a plurality of 32-bit registers; for a computing task that needs to use a register length of 256 bits, eight register combinations with a length of 32 bits are used; for a computing task that needs to use a register length of 128 bits, four register combinations with a length of 32 bits are used; for a computing task that needs to use a register length of 64 bits, two register combinations with a length of 32 bits are used; and for a register length of 32 bits, one register with a length of 32 bits is used; According to each of the candidate register combinations and the preset register cost function, a selected register combination is determined to obtain the vector register according to the selected register combination; wherein the following variables are defined Represents the set of all available registers in the system, represents the set of all pending computing tasks, Represents a computing task The required register length, Indicates register The current allocation status of Indicates the task Using Registers The preset register cost function is an overall evaluation of time delay and energy consumption, ,in represents the cost of time delay, Represents the cost caused by energy consumption; Wherein, the system further comprises: By using a gradient descent algorithm, an optimization operation is performed on the preset register cost function; wherein the preset register cost function at least includes a time delay cost and an energy consumption cost corresponding to a corresponding register combination; When the gradient modulus corresponding to the gradient descent algorithm is less than a preset threshold, the iteration is stopped, and the corresponding candidate register combination is determined to be the selected register combination; Wherein, the instruction cache unit of the first cache instruction includes vector instructions added by RISC-V; the vector instructions include: configuration setting instructions, vector load and store instructions, vector integer instructions, vector fixed-point instructions, and vector floating-point instructions; Wherein, the Warp corresponds to 7 CSR registers, including: vector starting position register, fixed-point saturation flag register, fixed-point rounding mode register, vector control and status register, vector length register, vector data type register, vector register length register; Wherein, the system further comprises: According to the functional category and operation complexity of the vector instruction and a preset operation code subspace comparison table, the operation code of the corresponding vector instruction is mapped to the operation code space of the RISC-V basic instruction format; wherein the preset operation code subspace comparison table includes the comparison relationship between operation codes of different functional categories and different operation complexities and different subspaces.

5. A collaborative computing device based on SIMD and SIMT, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a collaborative computing method based on SIMD and SIMT as described in any one of claims 1 to 3 above.