An optimization method for improving the parallelism and performance of hardware in executing RISC-V vector instructions
By using dynamic physical register ready delay queue and transmission conditions in the RISC-V processor, the pipeline blocking problem caused by vector instructions relying on old values is solved, and the parallelism and performance improvement of vector instructions is achieved.
Patent Information
- Application Number
- CN202510609397.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When RISC-V processors execute vector instructions, relying on the old value of the destination register causes the instruction to wait until the old value is fully ready before it can be transmitted, resulting in pipeline blockage, limited parallelism and reduced performance.
Dynamic physical register ready delay queue and transmission condition are used to dynamically calculate, allowing instructions to transmit in advance when the old value is not fully ready, and obtain the old value later in the operation through the data pre-transfer mechanism.
Improves the parallelism and performance of RISC-V vector instructions, reduces instruction transmission delay, and improves pipeline throughput.
Smart Images

Figure CN120123006B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optimization methods, and particularly relates to an optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions. Background Art
[0002] For traditional out-of-order processors, the target register of an instruction is written with a new value. Therefore, when renaming, only the available physical register number needs to be obtained from the register free list, and after the instruction gets the result, a new value is written into this physical register. However, the V Extension of the RISC-V processor brings new challenges. In the V Extension, some vector instructions require the processor to implement the function of vector masks. When the bit vm enabling the vector mask function in the instruction is valid, the arithmetic unit needs to obtain the value of the vector logical register v0 as the mask. Each valid bit of the mask represents that the corresponding element of the target vector operand should use the new value as the result. And for the elements of the target vector operand corresponding to the non-valid bits of the mask, it is necessary to decide whether to retain the old value in the logical register or use the all-1 value as the result according to the vma configuration in the control register vtype. Similarly, there is a control register in the V Extension that can configure the vector length. Among them, the elements of the target vector operand within the vector length should use the new value as the result, and for the elements exceeding the vector length, it is necessary to decide whether to retain the old value in the logical register or use the all-1 value as the result according to the vta configuration in the control register vtype. This means that in some configurations, the destination operand also depends on the old value in the destination vector logical register. However, during the execution stage of the instruction, the old value of the destination vector register does not need to participate in the operation. It only needs to read back the old value before the last cycle of the operation, select the final result according to the mask in the last cycle of the operation, and write it back to the physical register in the next cycle. Therefore, the destination register does not need to be fetched before entering the arithmetic unit like an operand.
[0003] Disadvantages of the prior art:
[0004] 1. When a vector instruction depends on the old value of the destination register (controlled by the mask vm, vector length vlen, and vtype), the instruction must wait until the old value is completely ready before it can be issued, resulting in pipeline stalls and limiting parallelism and performance;
[0005] 2. Instruction issue latency, instructions that depend on old values must wait until the previous instructions are completely written back before they can be issued;
[0006] 3. Limited parallelism. Traditional data forwarding can only issue two cycles after the data is ready, unable to fully utilize the data stream during the operation stage. The present invention proposes an optimization method to improve the parallelism and performance of hardware executing RISC-V vector instructions. It allows instructions to be issued in advance and obtain the correct old value after meeting certain conditions, instead of waiting for the old value of the destination vector register to be ready. Thus, the purpose of improving the parallelism and performance of hardware executing RISC-V vector instructions is achieved. The present invention reduces the instruction issue latency through a dynamic physical register ready delay queue + dynamic calculation of issue conditions. Summary of the Invention
[0007] The purpose of the present invention is to reduce the instruction waiting time, improve the pipeline throughput rate, and be applicable to the existing out-of-order processor architecture, adapting to the RISC-V vector extension scenario, thereby overcoming the problems in the above background technology.
[0008] Based on the above technical ideas, the technical solution adopted by the present invention is as follows:
[0009] An optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions, comprising the following steps:
[0010] Step S1: Fetching and decoding stage steps, which include a fetching link and a decoding link. The fetching link reads instructions from memory in program order and stores them in the instruction queue. The decoding link includes parsing the instructions to identify whether they are vector mask instructions (through vm) or vector length control instructions (through vta). If the instruction requires the old value of the destination register (controlled by the mask or length), mark the instruction with an old value dependency flag.
[0011] Step S2: Register renaming stage steps, which include an operand mapping link and a destination operand mapping link. The operand mapping link includes obtaining the currently mapped physical register numbers for the operands of the instruction (such as vs1, vs2) from the register renaming table. The destination operand mapping link includes allocating a new physical register (such as P_new) from the free list to store the result of the instruction.
[0012] Step S3: Issue control stage steps, which include a dependency check link and an issue decision link. The dependency check link includes operand checks to confirm that the physical registers of all operands are ready (traditional logic), and old value dependency checks to query the current ready delay value of the old value physical register P_old. The issue decision link includes allowing the instruction to be issued and carrying an old value acquisition method flag (divided into three categories). If the old value delay = 2, forward data from the last cycle of the operation stage (EX3). If the old value delay = 1, forward data from the write-back stage (WB). If the old value delay = 0, directly read from the physical register.
[0013] S4 Execution Phase Steps, which include an old value acquisition section and an operation and result merging section. The old value acquisition section includes selecting a data source according to the old value acquisition method identifier carried in the emission phase, obtaining the old value from the forwarding network inside the execution unit in the EX2 phase, obtaining the old value from the write-back phase forwarding network in the EX3 phase, and directly reading the old value from the register file; the operation and result merging section includes the execution unit calculating a new result according to the instruction function (such as vector addition). At the last cycle of the operation (EX3), according to the values of vm and vector register v0, the old and new values are selected. According to vlen and vta, it is determined whether to use the old value or all-ones value for the elements exceeding the length, generating the final result, and preparing to write back to the new physical register P_new;
[0014] S5 Write-back and Commit Phase Steps, which include a write-back result section and an in-order commit section. The write-back result section includes writing the final result into the new physical register P_new and setting its ready latency counter to 0 (indicating that the data is available). The in-order commit section includes committing the instruction results in program order, updating the register renaming table, and releasing the old physical register P_old to the free list;
[0015] S6 Physical Register Ready Latency Maintenance Steps, which include a continuous operation section. The continuous operation section includes automatically decrementing the ready latency counter of all physical registers by 1 (until it reaches zero). For example: for an instruction with a total latency of 4 cycles, its counter changes as 4 → 3 → 2 → 1 → 0.
[0016] For further limitation of the above technical solution, in the S2 register renaming phase steps, this step further includes an old value physical register binding section and an initialization of the ready latency counter section. The old value physical register binding section includes obtaining the old physical register (such as P_old) currently mapped to the destination register from the renaming table and recording it in the instruction information. The initialization of the ready latency counter section includes initializing the ready latency counter of the new physical register P_new to the total execution cycles of this instruction (for example, 3 cycles in the EX phase + 1 cycle in the WB phase → total latency of 4 cycles).
[0017] For further limitation of the above technical solution, in the S1 instruction fetching and decoding phase steps, the following operation details are included in the instruction fetching section:
[0018] The program counter (PC) sequentially accesses the instruction cache (I-Cache) and prefetches instructions in batches according to the basic block (Basic Block);
[0019] The fetched instructions are stored in the instruction queue (Instruction Buffer), and the queue depth is designed according to the number of pipeline stages of the processor (such as 16 - 32 entries).
[0020] For further limitation of the above technical solution, in the instruction fetch and decoding stage step S1, the following operation details are included in the decoding link of this step:
[0021] Analyze the instruction field, identify the RISC-V vector instruction format (OP-V type), extract the control bits of vm (mask enable bit), vta (tail element reservation configuration), and vma (mask element reservation configuration), and determine the logical register numbers of the operands (vs1, vs2, vd).
[0022] Dependency flag setting: If vm = 1 or the vta / vma configuration requires retaining the old value, set the old value dependency flag (Old Value Dependency Flag, OVDF) of the instruction. The decoded micro-operation (μOp) includes the operation type, operand mapping, and OVDF flag, and is sent to the renaming stage.
[0023] Record the dependency type (mask dependency, length dependency, or both).
[0024] For further limitation of the above technical solution, in the register renaming stage step S2, the following operation details are included in the operand physical register mapping link of this step:
[0025] Query the renaming table to obtain the current physical register numbers (such as P1, P2) of the operands (vs1, vs2). If the source register is not ready (the ready latency counter of the corresponding physical register > 0), mark this operand as NotReady. The following operation details are included in the destination operand mapping link:
[0026] Allocate an idle physical register (such as P_new) from the free list as the new mapping of the destination register vd, update the renaming table, and map the logical register vd to P_new.
[0027] For further limitation of the above technical solution, in the register renaming stage step S2, the following operation details are included in the old value physical register binding link of this step:
[0028] Obtain the old mapped physical register (P_old) of vd from the renaming table (that is, the physical register before renaming, such as P_old), and append the P_old number to the micro-operation of the instruction for use in subsequent stages.
[0029] In the register renaming stage step S2, the following operation details are included in the initialization of the ready latency counter link of this step:
[0030] Determine the total number of latency cycles according to the instruction type (such as vector addition, masking operation) (e.g., 3 beats in the EX stage + 1 beat in the WB stage → total latency of 4 beats), initialize the counter of P_new to the total latency value (e.g., 4), and the counter automatically decrements with each clock cycle.
[0031] For further limitation of the above technical solution, in the step of the S3 emission control stage, the following operation details are included in the operand ready check link in this step:
[0032] Check whether the physical registers of all operands (P1, P2) are ready (counter = 0). If any operand is not ready, the instruction is temporarily stored in the emission queue and retried in subsequent cycles.
[0033] For further limitation of the above technical solution, in the step of the S3 emission control stage, the following operation details are included in the old value dependency check link in this step:
[0034] Query the current counter value (such as L_old) of the old value physical register P_old, and calculate the allowable emission condition: allowable emission condition = L_old ≤ (total latency of this instruction - 2).
[0035] For further limitation of the above technical solution, in the step of the S4 execution stage, the following operation details are included in the old value acquisition link in this step:
[0036] In the EX2 stage, obtain the old value from the middle pipeline stage of the execution unit (such as the result of the EX1 stage). The forwarding network directly passes the intermediate result of P_old to the EX2 stage of the current instruction. In the EX3 stage, obtain the old value from the forwarding bus in the write-back stage (WB).
[0037] For further limitation of the above technical solution, in the step of the S4 execution stage, the following operation details are included in the operation and result merging link in this step:
[0038] Calculate the new value, the execution unit completes the vector operation (such as vs1 + vs2), generates a temporary result, and performs mask control (if vm = 1);
[0039] Read the mask bits of the vector register v0, select to retain the old value or use the new value for each element, control the vector length (vlen and vta), determine the valid element range according to vlen, and for elements outside the range, retain the old value or set them to all 1 according to the vta configuration. At the last beat of the EX3 stage, write the final result after mask and length control into the temporary register to prepare for write-back.
[0040] For further limitation of the above technical solution, in the step of the S5 write-back and commit stage, the following operation details are included in the result write-back to the physical register link in this step:
[0041] Write the final result to the destination physical register P_new, set the ready latency counter of P_new to 0, and mark the data as ready;
[0042] The S5 write-back and commit phase steps, in which the sequential commit and resource release steps include the following operation details:
[0043] Commit the instructions in program order, update the architectural state, release the old physical register, release P_old back to the free list (provided that subsequent instructions no longer depend on its old value), update the rename table, and ensure that the mapping of the logical register vd points to P_new.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] 1. Without considering other data dependencies, the vector instruction in this solution that depends on the previous instruction as the old value of the destination operand of this instruction can be issued 2 clock cycles earlier than the previous solution; A dynamic latency calculation and early issue mechanism is proposed, allowing instructions to be issued early when the old value is not fully ready, and obtaining the old value through the data forward mechanism in the later stage of the operation, reducing the waiting time. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is the timing diagram of the optimization method of the present invention;
[0048] Figure 2 It is the improved timing diagram of the optimization method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The following combines the attached Figure 1-2 to further elaborate on the present invention.
[0050] Embodiment 1: This embodiment provides an optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions, as Figure 1-2 shown, including the following steps:
[0051] Step of the instruction fetching and decoding stage S1, which includes an instruction fetching link and a decoding link. The instruction fetching link reads instructions from the memory in program order and stores them in the instruction queue. The decoding link includes parsing the instructions to identify whether they are vector mask instructions (through vm) or vector length control instructions (through vta). If an instruction requires the old value of the destination register (controlled by the mask or length), mark that the instruction needs to carry the old value dependency identifier.
[0052] Step of the register renaming stage S2, which includes an operand mapping link and a destination operand mapping link. The operand mapping link includes obtaining the currently mapped physical register numbers for the operands of the instruction (such as vs1, vs2) from the register renaming table. The destination operand mapping link includes allocating a new physical register (such as P_new) from the free list for storing the result of the instruction.
[0053] Step of the issue control stage S3, which includes a dependency check link and an issue decision link. The dependency check link includes operand checks to confirm that the physical registers of all operands are ready (traditional logic), and old value dependency checks to query the current ready latency value of the old value physical register P_old. The issue decision link includes allowing the instruction to be issued and carrying the old value acquisition method identifier (divided into three categories). If the old value latency = 2, forward the data from the last cycle of the execution stage (EX3). If the old value latency = 1, forward the data from the write-back stage (WB). If the old value latency = 0, directly read from the physical register.
[0054] Step of the execution stage S4, which includes an old value acquisition link and an operation and result merging link. The old value acquisition link includes selecting the data source according to the old value acquisition method identifier carried in the issue stage, obtaining the old value from the internal forward network of the execution unit in the EX2 stage, obtaining the old value from the write-back stage forward network in the EX3 stage, or directly reading the old value from the register file. The operation and result merging link includes the execution unit calculating the new result according to the instruction function (such as vector addition). At the last cycle of the operation (EX3), select the new and old values according to vm and the value of the vector register v0, and decide whether to use the old value or all-1 values for the elements exceeding the length according to vlen and vta, generate the final result, and prepare to write back to the new physical register P_new.
[0055] Step of the write-back and commit stage S5, which includes a write-back result link and an in-order commit link. The write-back result link includes writing the final result into the new physical register P_new and setting its ready latency counter to 0 (indicating that the data is available). The in-order commit link includes committing the instruction results in program order, updating the register renaming table, and releasing the old physical register P_old to the free list.
[0056] S6 Physical Register Ready Latency Maintenance Step, which includes a continuous operation phase. The continuous operation phase includes automatically decrementing the ready latency counter of all physical registers by 1 (until it reaches zero). For example, for an instruction with a total latency of 4 beats, its counter changes as 4 → 3 → 2 → 1 → 0.
[0057] The S2 Register Renaming Phase Step, which also includes an old value physical register binding phase and an initializing ready latency counter phase. The old value physical register binding phase includes obtaining the old physical register (such as P_old) currently mapped to the destination register from the renaming table and recording it in the instruction information. The initializing ready latency counter phase includes initializing the ready latency counter of the new physical register P_new to the total execution cycles of the instruction (for example, 3 beats in the EX phase + 1 beat in the WB phase → total latency of 4 beats).
[0058] The S1 Instruction Fetch and Decode Phase Step. In the instruction fetch link of this step, the following operation details are included:
[0059] The program counter (PC) sequentially accesses the instruction cache (I-Cache) and prefetches instructions in batches according to the basic block (Basic Block).
[0060] The fetched instructions are stored in the instruction queue (Instruction Buffer), and the queue depth is designed according to the number of pipeline stages of the processor (such as 16 - 32 entries).
[0061] The S1 Instruction Fetch and Decode Phase Step. In the decode link of this step, the following operation details are included:
[0062] Parse the instruction fields, identify the RISC-V vector instruction format (OP-V type), extract the control bits of vm (mask enable bit), vta (tail element reservation configuration), and vma (mask element reservation configuration), and determine the logical register numbers of the operands (vs1, vs2, vd).
[0063] Dependency flag setting. If vm = 1 or the vta / vma configuration requires retaining the old value, set the old value dependency flag (Old Value Dependency Flag, OVDF) of the instruction. The decoded micro-operation (μOp) includes the operation type, operand mapping, and OVDF flag, and is sent to the renaming phase.
[0064] Record the dependency type (mask dependency, length dependency, or both).
[0065] The S1 Instruction Fetch and Decode Phase Step. In the decode link of this step, the following operation details are included:
[0066] Parse the instruction field, identify the RISC-V vector instruction format (OP-V type), extract the control bits of vm (mask enable bit), vta (tail element reservation configuration), and vma (mask element reservation configuration), and determine the logical register numbers of the operands (vs1, vs2, vd).
[0067] Dependency identification setting: If vm = 1 or the vta / vma configuration requires retaining the old value, set the old value dependency flag (OVDF) of the instruction. The decoded micro-operation (μOp) includes the operation type, operand mapping, and OVDF flag, and is sent to the rename stage.
[0068] Record the dependency type (mask dependency, length dependency, or both).
[0069] The steps of the S2 register rename stage. In the operand physical register mapping link of this step, the following operation details are included:
[0070] Query the rename table to obtain the current physical register numbers (such as P1, P2) of the operands (vs1, vs2). If the source register is not ready (the ready latency counter of the corresponding physical register > 0), mark the operand as NotReady. The following operation details are included in the destination operand mapping link:
[0071] Allocate an idle physical register (such as P_new) from the free list as the new mapping of the destination register vd, update the rename table, and map the logical register vd to P_new.
[0072] The steps of the S2 register rename stage. In the old value physical register binding link of this step, the following operation details are included:
[0073] Obtain the old mapped physical register (P_old) of vd from the rename table (that is, the physical register before renaming, such as P_old), and append the P_old number to the micro-operation of the instruction for use in subsequent stages.
[0074] The steps of the S2 register rename stage. In the initialization of the ready latency counter link of this step, the following operation details are included:
[0075] Determine the total number of latency cycles according to the instruction type (such as vector addition, mask operation) (for example, 3 beats in the EX stage + 1 beat in the WB stage → total latency of 4 beats), initialize the counter of P_new to the total latency value (such as 4), and the counter automatically decrements with each clock cycle.
[0076] The steps of the S3 issue control stage. In the operand ready check link of this step, the following operation details are included:
[0077] Check whether the physical registers of all operands (P1, P2) are ready (counter = 0). If any operand is not ready, the instruction is temporarily stored in the issue queue and retried in subsequent cycles;
[0078] In the S3 issue control stage step, the following operation details are included in the old value dependency check link:
[0079] Query the current counter value of the old value physical register P_old (such as L_old), and calculate the allowable issue condition: allowable issue condition = L_old ≤ (total delay of this instruction - 2).
[0080] In the S4 execution stage step, the following operation details are included in the old value acquisition link:
[0081] In the EX2 stage, obtain the old value from the middle pipeline stage of the execution unit (such as the result of the EX1 stage). The forwarding network directly passes the intermediate result of P_old to the EX2 stage of the current instruction. In the EX3 stage, obtain the old value from the forwarding bus of the write-back stage (WB).
[0082] In the S4 execution stage step, the following operation details are included in the operation and result merging link:
[0083] Calculate the new value. The execution unit completes the vector operation (such as vs1 + vs2), generates a temporary result, and mask control (if vm = 1);
[0084] Read the mask bits of the vector register v0, select to retain the old value or use the new value for each element, vector length control (vlen and vta), determine the valid element range according to vlen, and for elements outside the range, retain the old value or set to all 1 according to the vta configuration. At the last cycle of the EX3 stage, write the final result after mask and length control into the temporary register to prepare for write-back.
[0085] In the S5 write-back and commit stage step, the following operation details are included in the result write-back to the physical register link:
[0086] Write the final result into the destination physical register P_new, set the ready delay counter of P_new to 0, and mark the data as ready;
[0087] In the S5 write-back and commit stage step, the following operation details are included in the sequential commit and resource release link:
[0088] Submit instructions in program order, update the Architectural State, release the old physical registers, release P_old back to the free list (provided that subsequent instructions no longer depend on its old value), update the rename table, and ensure that the mapping of the logical register vd points to P_new.
[0089] Embodiment 2: This embodiment provides an optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions. As Figure 1-2 shown, it further includes the following steps:
[0090] 1. The instruction fetch stage is the same as that of a traditional out-of-order processor.
[0091] 2. In the decoding stage, according to the vm field of the instruction, the vector length, and the value of the control register vtype, determine whether the result of this instruction needs to use the old value of the destination register. If so, carry a flag for reading the old value.
[0092] 3. In the register renaming stage, in addition to finding the physical register mapping for the operands in the instruction and mapping the logical register of the target register of the instruction to an idle physical register, it is also necessary to check whether the instruction carries a flag for reading the old value of the target register. If it carries this flag, it is also necessary to use the register renaming table to find the old physical register mapping of the logical register of the target register.
[0093] 4. In the instruction scheduling and instruction issue stages, if the instruction carries a flag, when checking whether the operands are ready, it is also necessary to check whether the value of the old physical register of the target register is ready. Only when all these dependent data are ready can the instruction be issued.
[0094] 5. In the instruction execution stage, if the instruction carries a flag, it is necessary to fetch both the operands and the old value of the destination operand, then participate in the operation, and determine whether each element of the destination operand is obtained from the old value or from the result obtained by calculating the operands according to the mask and the vector length. Finally, write the result back to the newly mapped physical register of the destination operand.
[0095] The instruction completion and instruction submission stages are the same as those of a traditional out-of-order processor.
[0096] Embodiment 3: This embodiment provides an optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions. As Figure 1-2 shown, it further includes the instruction execution process of an out-of-order processor:
[0097] 1. Instruction Fetch
[0098] The processor fetches instructions from memory and fetches them in the order of the program counter (PC). The fetched instructions are stored in the instruction queue for further processing.
[0099] 2. Instruction Decode
[0100] The fetched instructions are decoded to determine the type of instruction, the operands, and the execution units required. If the required operands are ready, the instruction will enter the execution stage; otherwise, it will wait for the operands to be ready.
[0101] 3. Register Renaming
[0102] To avoid register conflicts between instructions (such as read-after-write or write-write conflicts), the processor performs register renaming. This means that logical registers are mapped to physical registers to eliminate potential data dependency problems. There is a set of physical registers in the processor's architecture. The processor uses a register renaming table (Renaming Table) to find the physical register mapping for the operands in the instruction, and uses a register free list (Freelist) to map the logical register of the destination register of the instruction to a free physical register. In this way, even if different instructions modify the same logical register, they can still use different physical registers, thus avoiding data dependency conflicts.
[0103] 4. Instruction Dispatch
[0104] After the instruction is decoded and the necessary operands are obtained, it is dispatched to the appropriate execution unit. Due to the characteristics of an out-of-order processor, the instruction decides when to execute based on the available execution units and resources, rather than in program order.
[0105] 5. Instruction Issue
[0106] After instruction dispatch, the instruction is cached in an instruction queue until its operands are ready and there are available execution units to execute it.
[0107] 6. Out-of-Order Execution
[0108] In an out-of-order processor, if the pre-dependencies of an instruction are satisfied, it can be executed earlier even if its position in the program is relatively late. The processor looks for free execution units and executes the instruction in one of them. The execution order of these instructions is determined dynamically, rather than in a fixed order.
[0109] 7. Instruction Completion
[0110] After the execution unit completes the instruction operation, the result is written back to a register or memory. An out-of-order processor may save the result of an instruction in a temporary buffer first instead of writing it back to the register immediately to ensure the correct execution order of the instructions.
[0111] 8. Instruction Commit
[0112] Finally, the results of the instructions are committed in the original order of the program to ensure the correct semantics of the program. If an exception or branch error occurs, the processor will withdraw the out-of-order executed instructions (by rolling back) and restore to the correct state.
[0113] Embodiment 4: This embodiment provides an optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions. As Figure 1-2 shown, it further includes the following instruction execution stages:
[0114] In the execution stage of the instruction, the old value of the destination register does not need to participate in the operation. It only needs to read back the old value before the last cycle of the operation, select the final result according to the mask in the last cycle of the operation, and write it back to the physical register in the next cycle. Therefore, the destination register does not need to be fetched before entering the arithmetic unit like the operands. In the instruction scheduling and issuing stages, when querying whether the old value of the destination operand is ready, information on whether the instruction can be issued in advance can be obtained by recording the latency of data readiness.
[0115] As Figure 1 shown, assume that the old values of the destination operands that OP2, OP3, and OP4 need to fetch all come from the result of the destination operand of OP1 (usually this is not the case. If the logical encodings of the destination registers of OP2, OP3, and OP4 are the same, then OP3 depends on OP2 and OP4 depends on OP3. This figure can be considered as the situation where OP2, OP3, and OP4 are the same instruction that depends on the result of OP1 arriving at different times), and other dependencies are all satisfied. At the second-to-last cycle (EX2 in the above figure) when OP1's operation ends, the message that the destination vector register of this instruction is ready is broadcast to the instruction queue in the instruction scheduling. Then the next instruction can be issued as early as T3 (i.e., the case of OP2), and it will obtain the data from the last cycle (EX3) of the previous instruction's operation; if it is issued at T4, it will obtain the data from the write-back cycle of the previous instruction; if it is issued after T5, it can read the data from the physical register.
[0116] Embodiment 4: This embodiment provides an optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions. AsFigure 1-2 As shown, it further includes an improved method for Embodiment 3, comprising the following steps:
[0117] In the execution stage of the instruction, the old value of the destination register does not need to participate in the operation. It only needs to read back the old value before the last cycle of the operation, and select the final result according to the mask in the last cycle of the operation, and write it back to the physical register in the next cycle. Therefore, the destination register does not need to be fetched before entering the operation unit like the operand. In the instruction scheduling and issuing stages, when querying whether the old value of the destination operand is ready, the information on whether the instruction can be issued in advance can be obtained by recording the latency of data readiness.
[0118] As Figure 2 shown, assume that the old values of the destination operands required by OP2, OP3, and OP4 in the figure all come from the destination operand result of OP1 (usually this situation does not occur. If the logical encodings of the destination registers of OP2, OP3, and OP4 are the same, then OP3 depends on OP2, and OP4 depends on OP3. In this figure, it can be considered that OP2, OP3, and OP4 are the same instruction that depends on the result of OP1 arriving at different times), and other dependencies are all satisfied. A queue of physical register data readiness latency will be maintained within the instruction scheduling module. In the cycle after OP1 is issued, the readiness latency corresponding to the physical number of the destination register of OP1 will be set to the latency of this instruction (including the latency of execution (EX) and writeback (WB)). Thereafter, this readiness latency will be decremented by one each cycle until it reaches 0. The latency of all instructions in the figure is 4 (3 cycles for calculation and 1 cycle for writeback).
[0119] When subsequent instructions enter the instruction scheduling / issuing module, querying whether the old value of the target operand is ready can be obtained by calculating the readiness latency of the target register in this cycle and the latency of the instruction itself. Subtracting 2 from the latency of this instruction (because at least one cycle of operation and one cycle of writeback are included, so the latency of the instruction must be greater than 2), the number of cycles after which the old value of the target operand needs to be read if this instruction is issued in this cycle can be obtained. Subtracting the latency of the instruction to read the old value from the readiness latency of the target register (if the result is negative, the result is considered 0), the result of whether the instruction can be issued and where to fetch the data can be obtained.
[0120] When the calculated value is less than 2, it can be considered that the old value of the target operand is ready. If other dependencies are also satisfied, this instruction can be issued. When the instruction is issued, the value calculated in this cycle will be carried to the subsequent pipeline, and the old value will be read two cycles before this instruction reaches the write-back stage (for the instruction latency of 4 in the figure, so the old value will be read in the EX2 cycle). When the value is 2, data is forwarded from the last cycle of the operation; when the value is 1, data is forwarded from the write-back; when the value is 0, data is fetched from the register file.
[0121] Compared with the previous solution, this instruction can be issued 2 clock cycles earlier than its latency. The above content is a further detailed description of the present invention in combination with specific preferred implementation schemes, which is convenient for those skilled in the art of this technology to understand and apply the present invention. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions.
Claims
1. An optimization method for improving the parallelism and performance of hardware in executing RISC-V vector instructions, characterized in that, It includes the following steps: The steps of the fetch and decode stage, which include a fetch link and a decode link. In the fetch link, instructions are read from the memory in program order and stored in the instruction queue. The decode link includes parsing the instructions to identify whether they are vector mask instructions or vector length control instructions. If the instruction requires the old value of the destination register, mark that the instruction needs to carry an old value dependency flag. The steps of the register renaming stage, which include an operand mapping link and a destination operand mapping link. The operand mapping link includes obtaining the currently mapped physical register numbers for the operands of the instruction from the register renaming table. The destination operand mapping link includes allocating a new physical register from the free list to store the result of the instruction. The steps of the issue control stage, which include a dependency check link and an issue decision link. The dependency check link includes operand checks to confirm that the physical registers of all operands are ready, and old value dependency checks to query the current ready latency value of the old value physical register P_old. The issue decision link includes allowing the instruction to be issued and carrying an old value acquisition method flag. If the old value latency = 2, forward the data from the second last cycle of the execution stage. If the old value latency = 1, forward the data from the write-back stage. If the old value latency = 0, directly read from the physical register. The steps of the execution stage, which include an old value acquisition link and an operation and result merging link. The old value acquisition link includes selecting the data source according to the old value acquisition method flag carried in the issue stage, obtaining the old value from the internal forwarding network of the execution unit at the second last cycle EX2 of the operation, obtaining the old value from the write-back forwarding network at the last cycle EX3 of the operation, or directly reading the old value from the physical register in the register file. The operation and result merging link includes the execution unit calculating the new result according to the instruction function. At the last cycle of the operation, select the old and new values according to the mask enable bit vm and the value of the vector register v0, and decide whether to use the old value or all 1 values for the elements exceeding the length according to the vector length vlen and the tail element retention configuration vta, and generate the final result and prepare to write it back to the new physical register P_new. The steps of the write-back and commit stage, which include a write-back result link and an in-order commit link. The write-back result link includes writing the final result into the new physical register P_new and setting its ready latency counter to 0. The in-order commit link includes committing the instruction results in program order, updating the register renaming table, and releasing the old physical register P_old to the free list. The steps of maintaining the ready latency of physical registers, which include a continuous operation link. The continuous operation link includes automatically decrementing the ready latency counter of all physical registers by 1.
2. An optimization method for improving the parallelism and performance of hardware in executing RISC-V vector instructions according to claim 1, characterized in that, The steps of the S2 register renaming stage also include an old value physical register binding link and an initializing ready latency counter link. The old value physical register binding link includes obtaining the old physical register currently mapped to the destination register from the renaming table and recording it in the instruction information. The initializing ready latency counter link includes initializing the ready latency counter of the new physical register P_new to the total execution cycles of the instruction.
3. An optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions according to claim 2, characterized in that, The steps of the S1 instruction fetching and decoding stage. In the instruction fetching part of this step, the following operation details are included: The program counter sequentially accesses the instruction cache and prefetches instructions in batches according to the basic block. The fetched instructions are stored in the instruction queue, and the queue depth is designed according to the number of pipeline stages of the processor.
4. An optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions according to claim 3, characterized in that The steps of the S1 instruction fetching and decoding stage. In the instruction decoding part of this step, the following operation details are included: Parse the instruction fields, identify the RISC-V vector instruction format, extract the mask enable bit vm, the tail element reservation configuration vta, the mask element reservation configuration vma control bit, and determine the logical register numbers of the operands. Dependency flag setting. If vm = 1 or the vta / vma configuration requires retaining the old value, set the old value dependency flag of the instruction. The decoded micro-operations include the operation type, operand mapping, and OVDF flag, and are sent to the renaming stage. Record the dependency type.
5. An optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions according to claim 4, characterized in that, The steps of the S2 register renaming stage. In the physical register mapping of operands part of this step, the following operation details are included: Query the renaming table to obtain the current physical register number of the operand. If the source register is not ready, mark this operand as not ready. In the destination operand mapping part of this step, the following operation details are included: Allocate an idle physical register from the free list as the new mapping of the logical register vd, update the renaming table, and map the logical register vd to P_new.
6. An optimization method for improving the parallelism and performance of hardware in executing RISC-V vector instructions according to claim 5, characterized in that, The steps of the S2 register renaming stage. In the old value physical register binding part of this step, the following operation details are included: Obtain the old mapped physical register of vd from the renaming table, and append the P_old number to the micro-operations of the instruction for use in subsequent stages. The steps of the S2 register renaming stage. In the initialization of the ready latency counter part of this step, the following operation details are included: Determine the total number of latency cycles according to the instruction type, initialize the counter of P_new to the total latency value, and the counter automatically decrements with each clock cycle.
7. An optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions according to claim 6, characterized in that, The steps of the S3 issue control stage. In the operand ready check part of this step, the following operation details are included: Check whether the physical registers of all operands are ready. If any operand is not ready, the instruction is temporarily stored in the issue queue and retried in subsequent cycles. The steps of the S3 issue control stage. In the old value dependency check part of this step, the following operation details are included: Query the current counter value of the old value physical register P_old, and calculate the allowable issue condition: Allowable issue condition = L_old ≤ the total latency of this instruction - 2, where L_old is the current counter value of the old value physical register P_old.
8. An optimization method for improving the parallelism and performance of hardware in executing RISC-V vector instructions according to claim 7, characterized in that, The steps of the S4 execution stage. In the old value acquisition part of this step, the following operation details are included: In the EX2 stage, obtain the old value from the middle pipeline stage of the execution unit. The forwarding network directly passes the intermediate result of P_old to the EX2 stage of the current instruction. In the EX3 stage, obtain the old value from the forwarding bus in the write-back stage.
9. An optimization method for improving the parallelism and performance of hardware in executing RISC-V vector instructions according to claim 7, characterized in that, The steps of the S4 execution stage. In the operation and result merging part of this step, the following operation details are included: Calculate the new value, the execution unit completes the vector operation, generates a temporary result, and mask control. Read the mask bits of vector register v0, select to retain the old value or use the new value element by element, control the vector length, determine the valid element range according to vlen, and for elements outside the range, retain the old value or set them to all 1 according to the vta configuration. At the last cycle of the EX3 stage, write the final result after mask and length control to a temporary register, preparing for write-back.
10. An optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions according to claim 7, characterized in that, The S5 write-back and commit stage steps. In the step of writing the result back to the physical register, the following operation details are included: Write the final result to the new physical register P_new, set the ready latency counter of P_new to 0, and mark the data as ready; The S5 write-back and commit stage steps. In the step of sequential commit and resource release, the following operation details are included: Commit the instructions in program order, update the architecture state, release the old physical register, release P_old back to the free list, update the rename table, and ensure that the mapping of the logical register vd points to P_new.
Citation Information
Patent Citations
Six-stage pipeline processor based on RISC-V instruction set
CN114721724A
Four-emission RISC-V processor micro-architecture and working method thereof
CN115454504A