Implementation method of RISC-V non-single width vector instruction and processor
By splitting RISC-V non-single-width vector instructions into multiple micro-operations, the problems of large area and high power consumption are solved, and more efficient instruction execution is achieved.
Patent Information
- Application Number
- CN202411706919.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-09-19
AI Technical Summary
The splitting and implementation of RISC-V non-single-width vector instructions have problems with large area and high power consumption.
By splitting and fetching multiple micro-operations based on non-single-width vector instructions, each micro-operation corresponds to a fixed number of source operands, and issuing and executing the micro-operations.
The area and power consumption of splitting and implementing RISC-V non-single-width vector instructions are reduced.
Smart Images

Figure CN120670025A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and more specifically, to a method and processor for implementing RISC-V non-single-width vector instructions. Background Art
[0002] The fifth-generation Reduced Instruction Set Computer-five (RISC-V) has the advantages of being fully open source, simple architecture, and modular. RISC-V Vector Extension (V Extension) version 1.0 was officially launched in 2021. Compared to the traditional Single Instruction Multiple Data (SIMD) architecture, the vector architecture has variable vector length. The same set of binary code can run on processors with different vector lengths without recompilation, which provides better flexibility and portability. In addition to the above advantages, RISC-V's vector extension also introduces the concept of a vector length multiplier (LMUL), which allows the vector length of instruction operations to exceed the register length implemented in the architecture.
[0003] RISC-V V extension instructions are divided into the following three categories based on the element widths of the source and destination registers: Single-width (SEW = SEW OP SEW); Narrow (SEW = 2SEW OP SEW); and Widen (2SEW = SEW OP SEW or 2SEW = 2SEW OP SEW). The Selected Element Width (SEW) represents the size of each element in the vector register, and OP stands for Operation (OP).
[0004] During the execution of instructions in the RISC-V V extension architecture, for single-width instructions, ignoring the mask register, it is sufficient to read the values of three registers: two source operand registers and one destination register. The purpose of reading the destination register as a source is to accommodate the peculiarities of the RISC-V V extension, which may require preserving the original values of inactive elements. Similarly, the execution unit only requires three source interfaces, each of which is the width of a vector register (VLEN). The write-back channel only requires the width of one register.
[0005] In the related art, the splitting and implementation of RISC-V non-single-width vector instructions have problems of large area and high power consumption. Summary of the Invention
[0006] The embodiments of the present disclosure provide a method and processor for implementing RISC-V non-single-width vector instructions, so as to at least solve the problem of large area and high power consumption in the splitting and implementation of RISC-V non-single-width vector instructions in the related art.
[0007] According to one embodiment of the present disclosure, a method for implementing a RISC-V non-single-width vector instruction is provided, comprising: splitting and acquiring a plurality of micro-operations based on the non-single-width vector instruction, each of the micro-operations corresponding to a fixed number of source operands; and sending and executing the micro-operations.
[0008] According to another embodiment of the present disclosure, a processor of a RISC-V architecture is provided, including a controller and an execution unit, wherein the controller includes an instruction decoding unit, wherein the instruction decoding unit obtains multiple micro-operations based on non-single-width vector instruction splitting, wherein each of the micro-operations corresponds to a fixed number of source operands; and the execution unit is configured to receive and execute the micro-operations.
[0009] This disclosure provides a method for implementing RISC-V non-single-width vector instructions. This method involves splitting non-single-width vector instructions to obtain multiple micro-operations, each corresponding to a fixed number of source operands; and then sending and executing the micro-operations. This method addresses the problem of large footprint and high power consumption associated with the splitting and implementation of RISC-V non-single-width vector instructions in related technologies, thereby reducing the footprint and power consumption associated with the splitting and implementation of RISC-V non-single-width vector instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a diagram of the number of source operand registers and destination registers for Narrow type instructions;
[0011] Figure 2 This is a diagram of the number of source operand registers and destination registers for Widen type instructions (2SEW=SEW OP SEW);
[0012] Figure 3 This is a diagram of the number of source operand registers and destination registers for Widen type instructions (2SEW=2SEW OP SEW);
[0013] Figure 4 is a flowchart of a method for implementing a RISC-V non-single-width vector instruction according to an embodiment of the present disclosure;
[0014] Figure 5 1 is a diagram illustrating an example of the structure of a processor of the RISC-V architecture according to an embodiment of the present disclosure;
[0015] Figure 6This is a schematic diagram of the module design of a device for implementing RISC-V non-single-width vector instructions according to an embodiment of the present disclosure;
[0016] Figure 7 This is a diagram of the disassembly and data selection of the Widen type instruction (2SEW=SEW OP SEW);
[0017] Figure 8 This is a diagram of the disassembly and data selection of the Widen type instruction (2SEW=2SEW OP SEW);
[0018] Figure 9 This is a diagram showing the disassembly and data selection of a Narrow type instruction (SEW=2SEW OP SEW). DETAILED DESCRIPTION
[0019] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings and in conjunction with embodiments.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0021] In the related art, if LMUL of the control status register vtype in the V extension is set to 1, VLEN=4*SEW. Figure 1 This is a diagram of the number of source operand registers and destination registers for Narrow type instructions, such as Figure 1 As shown, for the Narrow instruction, it can be found that one of the sources of the instruction (vs2 in the figure) requires two registers, and the other source and destination register only require one register. If four source operand registers are required to adapt to this situation, this will increase the reading of the source operand register and the interface input of the execution unit source operand, increasing the area and power consumption.
[0022] Figure 2 This is a diagram of the number of source operand registers and destination registers for Widen type instructions (2SEW=SEW OP SEW), as shown in the following example: Figure 2 As shown, both sources require one register, and the destination requires two registers. To accommodate this situation, four source operand registers are also required. Furthermore, the result is written back using two registers, which requires an additional write-back interface, further increasing the area.
[0023] Figure 3 This is a diagram of the number of source operand registers and destination registers for Widen type instructions (2SEW=2SEW OP SEW), as shown in the figure below: Figure 3As shown, one of the sources requires two registers, the destination register requires two registers, and a total of five registers are required. Writing back still requires two registers, which consumes more resources.
[0024] The present disclosure provides a method for implementing a RISC-V non-single-width vector instruction. Figure 4 is a flowchart of a method for implementing a RISC-V non-single-width vector instruction according to an embodiment of the present disclosure, such as Figure 4 The process includes the following steps:
[0025] Step S402 : Split the non-single-width vector instruction to obtain multiple micro-operations, each micro-operation corresponding to a fixed number of source operands.
[0026] In the disclosed embodiment, before splitting a non-single-width vector instruction, it is necessary to first obtain the non-single-width vector instruction and decode it to confirm the type of the non-single-width vector instruction. In an exemplary embodiment, each micro-operation corresponds to a fixed number of source operands, including: each micro-operation corresponds to three source operands, wherein the third source operand is read from a destination register.
[0027] In the disclosed embodiment, when the vector length multiplier LMUL is 1 / 2 / 4, each disassembled micro-operation uop only needs to read three source operands, and no additional source operand read path is added.
[0028] In the disclosed embodiment, by limiting each micro-operation to correspond to three source operands, the splitting and implementation area of RISC-V non-single-width vector instructions can be effectively reduced.
[0029] In an exemplary embodiment, the non-single-width vector instruction includes at least one of the following: a narrow vector instruction; and an extended wide vector instruction.
[0030] In the embodiment of the present disclosure, during the splitting process, the non-single-width vector instruction is split according to its type to obtain multiple micro-operations.
[0031] In an exemplary embodiment, multiple micro-operations are obtained by splitting based on non-single-width vector instructions, including: when two source operand registers corresponding to the Widen vector instruction need to be expanded, the first source operand register and the second source operand register of each two adjacent micro-operations are the same, and the third source operand register of each micro-operation stores the original value of the destination register.
[0032] In the embodiment of the present disclosure, the meaning of every two adjacent micro-operation uops is that the first uop is adjacent to the second uop, the third uop is adjacent to the fourth uop, and so on.
[0033] In an exemplary embodiment, multiple micro-operations are split and obtained based on non-single-width vector instructions, including: when a source operand register corresponding to a Widen vector instruction needs to be expanded, the first source operand registers of every two adjacent micro-operations correspond to the same, and the third operand register of each micro-operation stores the original value of the destination register.
[0034] In an exemplary embodiment, multiple micro-operations are split and obtained based on non-single-width vector instructions, including: for Narrow vector instructions, the first source operand register and destination register of each two adjacent micro-operations correspond to the same, and the third operand register of each micro-operation stores the original value of the destination register.
[0035] In an exemplary embodiment, multiple micro-operations are split and acquired based on a non-single-width vector instruction, including: for a Narrow vector instruction, every two adjacent micro-operations form a read-after-write dependency.
[0036] In an exemplary embodiment, multiple micro-operations are split and acquired based on a non-single-width vector instruction, including: for a Narrow vector instruction, the destination register of the first micro-operation of every two adjacent micro-operations is the same as the third source operand register corresponding to the second micro-operation.
[0037] In an exemplary embodiment, multiple micro-operations are split and obtained based on non-single-width vector instructions, including: when the first source operand register and the second source operand register of each two adjacent micro-operations correspond to the same, or when the first source operand registers of each two adjacent micro-operations correspond to the same, the corresponding source operand in the corresponding source operand register is selected based on the counter value LMUM to perform an operation.
[0038] In the disclosed embodiment, for Wide vector instructions, the split uop selects the corresponding value in the source operand register to participate in the operation based on the counter value LMUM, that is, the upper half or the lower half of the source operand register. For Narrow vector instructions, the split uop also selects the corresponding value in the source operand register to participate in the operation based on the counter value LMUM.
[0039] Step S404: Send and execute micro-operations.
[0040] In an exemplary embodiment, sending micro-operations includes: for Widen vector instructions, randomly sending each micro-operation, or sending each micro-operation in sequence, or pairing two corresponding micro-operations using the same source operand register into groups and sending the micro-operations adjacently.
[0041] In the disclosed embodiment, for Widen vector instructions, since there is no read-after-write dependency, each disassembled uop can be randomly emitted. Here, it is best to emit two adjacent instructions in sequence, such as uop1 and uop2, because they share some registers. This can reduce register flips and reduce power consumption. For Narrow vector instructions, there is a read-after-write dependency between two adjacent uops. Then, uop2 can only be emitted after a certain number of clock cycles after uop1 is emitted. The number of cycles is determined by the instruction execution delay. Ideally, when the result of uop1 is just written back, uop2 is sent to the execution unit for execution. At this time, the write-back result of uop1 can be directly used by uop2 without having to read it from the register stack again, reducing power consumption.
[0042] In an exemplary embodiment, issuing micro-operations includes: for Narrow vector instructions, pairing two corresponding micro-operations using the same source operand register into a group, and issuing the micro-operations sequentially and adjacently.
[0043] In an exemplary embodiment, sending micro-operations sequentially and adjacently includes sending a next micro-operation after a preset number of clock cycles have elapsed after sending a current micro-operation.
[0044] In an exemplary embodiment, sending a micro-operation includes: for a Narrow vector instruction, controlling the writing back of an execution result of a current micro-operation and obtaining a source operand of a next micro-operation simultaneously.
[0045] In the disclosed embodiment, by adopting different strategies for sending micro-operations for different vector instruction types, it is possible to reduce the splitting and implementation power consumption of RISC-V non-single-width vector instructions.
[0046] In an exemplary embodiment, micro-operations are sent and executed, including: for Narrow vector instructions, the first micro-operation of every two adjacent micro-operations saves the original value of the high half of the destination register, and the execution result of the first micro-operation is written to the low half of the destination register; the second micro-operation of every two adjacent micro-operations obtains the destination register value of the first micro-operation, and writes the execution result of the second micro-operation into the high half of the corresponding destination register, and the low half of the destination register remains unchanged.
[0047] In an exemplary embodiment, executing a micro-operation includes: determining the number of execution units that execute the micro-operation corresponding to the non-single-width vector instruction based on the number of elements in the non-single-width vector instruction; or, when the number of execution units that execute the micro-operation corresponding to the non-single-width vector instruction is fixed, increasing the number of execution times of the execution unit.
[0048] In the disclosed embodiment, the above-mentioned determination of the number of execution units for executing micro-operations corresponding to the non-single-width vector instructions based on the number of elements in the non-single-width vector instructions makes the number of execution units adjustable, which can effectively reduce the number of execution units and further reduce the splitting and implementation area of RISC-V non-single-width vector instructions.
[0049] Through the disclosed embodiments, a method for implementing RISC-V non-single-width vector instructions is provided. This method involves splitting the non-single-width vector instructions to obtain multiple micro-operations, each corresponding to a fixed number of source operands; and then sending and executing the micro-operations. This method addresses the problem of large area and high power consumption associated with the splitting and implementation of RISC-V non-single-width vector instructions in related technologies, thereby reducing the area and power consumption associated with the splitting and implementation of RISC-V non-single-width vector instructions.
[0050] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present disclosure.
[0051] The present disclosure also provides a RISC-V architecture processor. Figure 5 : is a structural example diagram of a processor of the RISC-V architecture according to an embodiment of the present disclosure, such as Figure 5 As shown, the processor 50 of the RISC-V architecture includes a controller 510 and an execution unit 520. The controller 510 includes an instruction decoding unit 5101. The instruction decoding unit 5101 obtains multiple micro-operations based on non-single-width vector instruction splitting, where each micro-operation corresponds to a fixed number of source operands; the execution unit 520 is configured to receive and execute micro-operations.
[0052] In the disclosed embodiment, the RISC-V architecture processor 50 further includes conventional structures such as a data processing unit and registers, which are not described in detail in the disclosed embodiment. For example, the data processing unit can be used to process the source data and execution results of a micro-operation, and select the source operand register element to participate in the operation based on the micro-operation counter value LMUM. The registers can be used to store and read data from the source operand register and the destination register, and update the stored results based on the write-back mechanism and dependency relationships of the micro-operation.
[0053] In this embodiment, a device for implementing RISC-V non-single-width vector instructions is also provided, which is used to implement the above-mentioned embodiments and preferred implementation methods, and the details that have been described will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0054] The implementation device of the RISC-V non-single-width vector instruction provided by the embodiment of the present disclosure may include a splitting module and a sending module, wherein the splitting module is configured to split the non-single-width vector instruction to obtain multiple micro-operations, each micro-operation corresponding to a fixed number of source operands; the sending module is configured to send and execute the micro-operations.
[0055] In the embodiments of the present disclosure, the implementation device of the above-mentioned RISC-V non-single-width vector instructions can also include different modules, and the module naming and functional division can also be selected in different ways according to actual conditions, and no specific restrictions are made here.
[0056] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0057] An embodiment of the present disclosure further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when run.
[0058] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0059] An embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.
[0060] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0061] The embodiments of the present disclosure further provide a computer program product, including a computer program, which implements the steps of any of the above method embodiments when executed by a processor.
[0062] In an exemplary embodiment, the above-mentioned computer program product includes a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method described in each embodiment of the present application are implemented.
[0063] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0064] Obviously, those skilled in the art should understand that the modules or steps of the present disclosure described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present disclosure is not limited to any particular combination of hardware and software.
[0065] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, they are described below in conjunction with different embodiments.
[0066] Example 1
[0067] The disclosed embodiment provides an implementation method for RISC-V non-single-width vector instructions. Compared with single-width vector (single-width) instructions that need to be simply split into corresponding micro-operations (micro-operations, uops) through LMUL and the source operand registers of the preceding and following uops and the destination registers are not correlated, non-single-width instructions need to split the instructions into uops that only require three source operands, and the source operand registers of the preceding and following uops and the destination registers are correlated. In the disclosed embodiment, the case where LMUL is 1 / 2 / 4 is considered, and no disassembly is required for other values. Each uop disassembled in this way only needs to read three source operands, and no additional source operand read path, operand interface of the execution unit, and execution unit result write-back channel are added.
[0068] Figure 6 This is a schematic diagram of the module design of the implementation device of the RISC-V non-single-width vector instruction of the embodiment of the present disclosure, such as Figure 6As shown, it includes an instruction cache module, an instruction decoding and instruction splitting module, a renaming module, an instruction emission module, a data processing module, an instruction execution module and a register stack.
[0069] In the disclosed embodiment, the instruction cache module is used to store non-single-width instructions to be executed, so as to facilitate the acquisition of non-single-width instructions from this module. The instruction decoding and instruction splitting module is used to split the corresponding non-single-width instructions into uops that only require three source operands according to different instruction types, and then send them to the renaming module and the instruction emission module. The renaming module is used for register renaming. The instruction emission module is used to send the disassembled uop to the execution unit, and the information of the source operand register is sent to the register stack to read the data. For narrow type instructions, it is necessary to check the write-after-read dependency to ensure data dependency. The data processing module is used to process the elements of the source operand and write back the execution results according to the number of uop counters and the elements of the execution unit operation. The instruction execution module is used to perform corresponding operations according to different instruction types, and the number of execution units can be configured. The register stack is used to realize the reading, writing back and storage of source operands.
[0070] Based on the above-mentioned RISC-V non-single-width vector instruction implementation device, the embodiment of the present disclosure provides a RISC-V non-single-width vector instruction implementation method, including the following steps:
[0071] S1. Instruction decoding and splitting.
[0072] In this embodiment, the instructions obtained from the instruction cache module are decoded and split according to instruction types.
[0073] In this embodiment, in the most complex case under the RISC-V V architecture, LMUL=4, both Widen and Narrow type instructions will be split into 8 uops, but the disassembly situations are different.
[0074] For the Widen instruction (2SEW=SEW OP SEW), assuming the instruction format is op1 v24, v16, v8, its disassembly is shown in Table 1.
[0075] In this embodiment, src0, src1, and src2 represent source operand registers, where src0 is the first source operand register in the above embodiment, src1 is the second source operand register in the above embodiment, and src2 is the third source operand register in the above embodiment. dst represents the destination register, and LMUM represents the counter value.
[0076] Table 1 Example of splitting the Widen instruction (2SEW=SEW OP SEW)
[0077]
[0078]
[0079] In this embodiment, for this type of Widen instruction, the source operand registers src0 and src1 read by the two adjacent uops are the same, and the destination registers dst correspond to different ones. src2 is used to save the original value of the destination register dst.
[0080] In this embodiment, for the Widen instruction (2SEW=2SEW OP SEW), assuming the instruction format is op1 v24, v16, v8, the disassembly is shown in Table 2.
[0081] Table 2 Example of splitting the Widen instruction (2SEW=2SEW OP SEW)
[0082] uop number Opcode src0 src1 src2 dst LMUM 1 op1 v24, v16, v8 v8 v16 v24 v24 0 2 op1 v25, v17, v8 v8 v17 v25 v25 1 3 op1 v26, v18, v9 v9 v18 v26 v26 2 4 op1 v27, v19, v9 v9 v19 v27 v27 3 5 op1 v28, v20, v10 v10 v20 v28 v28 4 6 op1 v29, v21, v10 v10 v21 v29 v29 5 7 op1 v30, v22, v11 v11 v22 v30 v30 6 8 op1 v31, v23, v11 v11 v23 v31 v31 7
[0083] In this embodiment, for this type of Widen instruction, the source operand register src0 read by the two adjacent uops is the same, the other source operand register src1 corresponds to different destination registers dst, and src2 is used to save the original value of the destination register dst.
[0084] In this embodiment, for a Narrow instruction (SEW=2SEW OP SEW), assuming the instruction format is op1 v24, v16, v8, the disassembly is shown in Table 3.
[0085] Table 3 Narrow instruction (SEW=2SEW OP SEW) split example table
[0086] uop number Opcode src0 src1 src2 dst LMUM 1 op1 v24, v16, v8 v8 v16 v24 v24 0 2 op1 v24, v17, v8 v8 v17 v24 v24 1 3 op1 v26, v18, v9 v9 v18 v25 v25 2 4 op1 v27, v19, v9 v9 v19 v25 v25 3 5 op1 v28, v20, v10 v10 v20 v26 v26 4 6 op1 v29, v21, v10 v10 v21 v26 v26 5 7 op1 v30, v22, v11 v11 v22 v27 v27 6 8 op1 v31, v23, v11 v11 v23 v27 v27 7
[0087] For this type of Narrow instruction, the source operand register src0 and destination register dst read by the two adjacent uops are the same, and the other source operand register src1 corresponds to a different one, and src2 is used to save the original value of the destination register dst. At this time, unlike the Widen instruction, the destination register of the previous instruction (dst of uop1) is the source of the next instruction (src2 of uop2), so a write-after-read dependency occurs at this time to enable uop1 and uop2 to write the result back to the same register.
[0088] In the embodiment of the present disclosure, the aforementioned two adjacent uops mean that the first uop is adjacent to the second uop, the third uop is adjacent to the fourth uop, and so on.
[0089] S2. Command transmission.
[0090] In this embodiment, for the Widen instruction, since there is no read-after-write dependency, each disassembled uop can be issued randomly. Here, it is best to issue two adjacent instructions in sequence, such as uop1 and uop2, because they share some registers. This can reduce register flipping and reduce power consumption.
[0091] However, for Narrow instructions, there is a write-after-read dependency between two adjacent uops. Therefore, uop2 can only be issued after a certain number of clock cycles after uop1 is issued. The number of cycles is determined by the instruction execution delay. Ideally, when the result of uop1 is just written back, uop2 is sent to the execution unit for execution. At this time, the write-back result of uop1 can be directly used by uop2 without having to read it from the register stack again, reducing power consumption.
[0092] S3. Instruction execution.
[0093] In this embodiment, in the RISC V-V architecture, a vector register contains many elements. In order to reduce area, the number of execution units is sometimes reduced, and a uop is split into multiple times and sent to the execution units for operation, which is mostly used in floating-point execution units. For example, if there are 8 elements in the vector register and there are only 2 execution units, each execution unit can only process one element at a time, then it is necessary to split it into 4 times to process the elements in sequence. This requires more sophisticated processing in the data processing module, including sequential selection before sending the data to the execution unit and sequential saving of the results after writing them back, and then writing them back to the register file.
[0094] S4. Data processing.
[0095] In this embodiment, the data processing process is described by taking LMUL=1 and VLEN=4SEW as an example.
[0096] Figure 7 This is a diagram of the disassembly and data selection of the Widen type instruction (2SEW=SEW OP SEW), as shown in the following example: Figure 7 As shown, the uops split from this type of Widen instruction select the corresponding value in the source operand register based on the counter value LMUM to participate in the operation, that is, the upper half or the lower half of the two source operand registers. If the number of execution units is less than the number of elements, the uop can be executed multiple times, each time selecting the number of elements in the number of execution units to participate in the operation.
[0097] Figure 8 This is a diagram of the disassembly and data selection of the Widen type instruction (2SEW=2SEW OP SEW), as shown in the following example: Figure 8 As shown, the uop split out of this type of Widen instruction will also select the corresponding value in the source operand register according to the counter value LMUM to participate in the operation. At this time, only one source operand register vs2 requires this operation.
[0098] Figure 9 This is a diagram of the Narrow type instruction (SEW=2SEW OP SEW) instruction disassembly and data selection, as shown in the following figure: Figure 9 As shown, the uops split out of this type of Narrow instruction will also select the corresponding value in the source operand register to participate in the operation based on the counter value LMUM, which is vs1 in the figure. Special processing is required for writing back the result. For uop0, the original value of the high half of the destination register needs to be saved, and the result generated by its execution unit is only written back to the low half of the destination register; for uop1, the output of uop0 needs to be read, and then the low half of this output needs to be saved. The high half of the destination register is filled with the result of uop1. Under the RISCV-V architecture, the execution unit will have a mask to control whether the corresponding element is an active element (the element participating in the operation). If mask = 0, the result is the original value of the destination register; if mask = 1, the result is the result obtained by the operation of the source operand register. The original values of the high and low halves of the destination register can be saved by controlling the high and low halves of the mask to all be 0 to achieve the update of an entire register.
[0099] In summary, the disclosed embodiments provide a method for implementing RISC-V non-single-width vector instructions, which limits the bit width and number of the source operand interface of the execution unit and the interface for reading source operands from the register stack, thereby reducing the area. For non-single-width instructions, namely, the Widen and Narrow instructions, when LMUL ≥ 1, the instructions are disassembled to implement the function of the instruction, and the power consumption is effectively reduced through processing during instruction issuance. By reducing the number of execution units and related operations in data processing, the area is further reduced while ensuring that the instruction function remains unchanged.
[0100] The implementation method of the RISC-V non-single-width vector instructions provided by the embodiments of the present disclosure can be applied to a general-purpose central processing unit that supports vector instructions, or a digital signal processor or a vector processor or a graphics processor.
[0101] The foregoing description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations of the present disclosure are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present disclosure shall be included within the scope of protection of the present disclosure.
Claims
1. A method for implementing RISC-V non-single-width vector instructions, characterized in that: include: Splitting and acquiring a plurality of micro-operations based on a non-single-width vector instruction, each of the micro-operations corresponding to a fixed number of source operands; The micro-operation is sent and executed.
2. The method according to claim 1, characterized in that Each of the micro-operations corresponds to a fixed number of source operands, including: Each of the micro-operations corresponds to three source operands, wherein a third source operand among the source operands is read from a destination register.
3. The method according to claim 1, characterized in that in, The non-single-width vector instructions include at least one of the following: Narrow vector instructions; expand wide vector instructions.
4. The method according to claim 3, characterized in that The method of splitting and acquiring multiple micro-operations based on a non-single-width vector instruction includes: When the two source operand registers corresponding to the Widen vector instruction need to be expanded, the first source operand register and the second source operand register of each two adjacent micro-operations correspond to the same value, and the third source operand register of each micro-operation stores the original value of the destination register.
5. The method according to claim 3, characterized in that The method of splitting and acquiring multiple micro-operations based on a non-single-width vector instruction includes: When a source operand register corresponding to the Widen vector instruction needs to be expanded, the first source operand registers of every two adjacent micro-operations correspond to the same one, and the third operand register of each micro-operation stores the original value of the destination register.
6. The method according to claim 3, characterized in that The method of splitting and acquiring multiple micro-operations based on a non-single-width vector instruction includes: For the Narrow vector instruction, the first source operand register and the destination register of each two adjacent micro-operations correspond to the same value, and the third operand register of each micro-operation stores the original value of the destination register.
7. The method according to claim 3, characterized in that The method of splitting and acquiring multiple micro-operations based on a non-single-width vector instruction includes: For the Narrow vector instruction, every two adjacent micro-operations form a read-after-write dependency.
8. The method according to claim 7, characterized in that The method of splitting and acquiring multiple micro-operations based on a non-single-width vector instruction includes: For the Narrow vector instruction, the destination register of the first micro-operation in every two adjacent micro-operations is the same as the third source operand register corresponding to the second micro-operation.
9. The method according to claim 3, characterized in that Sending the micro-operation includes: For the Widen vector instruction, each micro-operation is sent randomly, or each micro-operation is sent sequentially, or two corresponding micro-operations using the same source operand register are paired into a group and sent adjacently.
10. The method according to claim 3, characterized in that Sending the micro-operation includes: For the Narrow vector instruction, two corresponding micro-operations using the same source operand register are paired into a group, and the micro-operations are sent adjacently in sequence.
11. The method according to claim 10, characterized in that The sending of the micro-operations sequentially and adjacently includes: After the current micro-operation is sent and a preset number of clock cycles have passed, the next micro-operation is sent.
12. The method according to claim 3, characterized in that Sending the micro-operation includes: For the Narrow vector instruction, the writing back of the execution result of the current micro-operation is controlled and the source operand of the next micro-operation is obtained simultaneously.
13. The method according to claim 3, characterized in that Sending and executing the micro-operation includes: For the Narrow vector instruction, the first micro-operation of every two adjacent micro-operations saves the original value of the high half of the destination register, and the execution result of the first micro-operation is written to the low half of the destination register; The second micro-operation of each two adjacent micro-operations obtains the destination register value of the first micro-operation, and writes the execution result of the second micro-operation into the high half of the corresponding destination register, while the low half of the destination register remains unchanged.
14. The method according to claim 1, wherein Executing the micro-operation includes: Determining the number of execution units that execute the micro-operations corresponding to the non-single-width vector instruction according to the number of elements in the non-single-width vector instruction; Alternatively, when the number of execution units of the micro-operation corresponding to the non-single-width vector instruction is fixed, the execution times of the execution unit are increased.
15. The method according to claim 3, characterized in that The method of splitting and acquiring multiple micro-operations based on a non-single-width vector instruction includes: When the first source operand register and the second source operand register of each two adjacent micro-operations correspond to the same, or when the first source operand registers of each two adjacent micro-operations correspond to the same, the corresponding source operand in the corresponding source operand register is selected based on the counter value LMUM for calculation operation.
16. A RISC-V architecture processor, comprising a controller and an execution unit, wherein the controller comprises an instruction decoding unit, characterized in that: The instruction decoding unit is configured to split and obtain a plurality of micro-operations based on a non-single-width vector instruction, wherein each of the micro-operations corresponds to a fixed number of source operands; and the execution unit is configured to receive and execute the micro-operations.
Citation Information
Patent Citations
Execution method of CASP instruction, microprocessor and computer equipment
CN110515656A
Instruction transmitting unit, instruction executing unit, related device and method
CN114428638A
Pipeline decoding micro-architecture design method for RISC-V vector instruction
CN118012504A
Instruction transmitting unit, instruction execution unit, and related apparatus and method
US20220147351A1