Optimization method for improving parallelism and performance of RISC-V vector instruction executed by hardware

By dynamically calculating the transmission conditions of instructions and using data pre-transfer mechanism, the pipeline blocking and performance limitations caused by the dependence of old values ​​by RISC-V processors when executing vector instructions is solved, achieving higher parallelism and performance.

CN120123006AActive Publication Date: 2025-06-10BEIJING YIHUA CLOUD NETWORK TECH CO LTD

Patent Information

Application Number
CN202510609397.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

When executing vector instructions, the RISC-V processor needs to rely on the old value of the destination register, so the instructions must wait until the old value is fully ready before they can be transmitted, resulting in pipeline blocking, limiting parallelism and performance.

Method used

Dynamic calculation of dynamic physical register ready delay queue and transmit conditions allows instructions to be transmitted in advance when the old value is not fully ready, and obtain the old value later in the operation through the data pre-transfer mechanism, reducing the latency time.

Benefits of technology

It realizes reducing instruction transmission delay, improving the parallelism and performance of hardware execution of RISC-V vector instructions, and overcoming the problems of pipeline blocking and parallelism limitation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123006A_ABST
    Figure CN120123006A_ABST
Patent Text Reader

Abstract

The invention provides an optimization method for improving parallelism and performance of hardware executing RISC-V vector instructions, which comprises the following steps: S1, an instruction fetching and decoding stage step: the step comprises an instruction fetching link and a decoding link, and the instruction fetching link reads instructions from a memory according to a program sequence and stores the instructions into an instruction queue; the decoding link comprises the steps of analyzing an instruction, identifying whether the instruction is a vector mask instruction or a vector length control instruction, and if the instruction needs an old value of a destination register, marking that the instruction needs to carry an old value dependency identifier, belonging to the technical field of optimization methods. According to the method, the instruction is transmitted without waiting for the old value of the target vector register to be ready, the instruction can be transmitted in advance after a certain condition is met, and the correct old value is obtained, so that the aim of improving the parallelism and the performance of hardware for executing the RISC-V vector instruction is fulfilled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimization methods, and particularly to an optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions. Background Art

[0002] For traditional out-of-order processors, the target register of an instruction will be written with a new value. Therefore, when renaming, only the available physical register number needs to be obtained from the register free list, and after the instruction gets the result, a new value is written into this physical register. However, the V Extension of the RISC-V processor brings new challenges. In the V Extension, some vector instructions require the processor to implement the function of vector masks. When the bit vm enabling the vector mask function in the instruction is valid, the arithmetic unit needs to obtain the value of the vector logical register v0 as the mask. Each valid bit of the mask represents that the corresponding element of the target vector operand is to use the new value as the result. And for the elements of the target vector operand corresponding to the non-valid bits of the mask, it is necessary to decide whether to retain the old value in the logical register or use the all-1 value as the result according to the vma configuration in the control register vtype. Similarly, there is a control register in the V Extension that can configure the vector length. Among them, for the elements of the target vector operand within the vector length, the new value is used as the result, and for the elements exceeding the vector length, it is necessary to decide whether to retain the old value in the logical register or use the all-1 value as the result according to the vta configuration in the control register vtype. This means that in some configurations, the destination operand also depends on the old value in the destination vector logical register. However, in the execution stage of the instruction, the old value of the destination vector register does not need to participate in the operation. It only needs to read back the old value before the last cycle of the operation, select the final result according to the mask in the last cycle of the operation, and write it back to the physical register in the next cycle. Therefore, the destination register does not need to be fetched before entering the arithmetic unit like the operands.

[0003] Disadvantages of the prior art: 1. When a vector instruction needs to depend on the old value of the destination register (controlled by the mask vm, vector length vlen, and vtype), the instruction must wait until the old value is completely ready before it can be issued, resulting in pipeline stalls and limiting parallelism and performance; 2. Instruction issue latency. Instructions that depend on old values must wait until the previous instructions are completely written back before they can be issued; 3. Limited parallelism. Traditional data forwarding can only issue instructions two cycles after the data is ready, unable to fully utilize the data stream during the operation stage. The present invention proposes an optimization method to improve the parallelism and performance of hardware executing RISC-V vector instructions. It allows instructions to be issued in advance and obtain the correct old value after meeting certain conditions, instead of waiting for the old value of the destination vector register to be ready. Thus, the purpose of improving the parallelism and performance of hardware executing RISC-V vector instructions is achieved. The present invention realizes the effect of reducing instruction issue latency through a dynamic physical register ready delay queue + dynamic calculation of issue conditions. Summary of the Invention

[0004] The object of the present invention is to reduce the instruction waiting time, improve the pipeline throughput rate, and be applicable to the existing out-of-order processor architecture, adapting to the RISC-V vector extension scenario, thereby overcoming the problems in the above background technology.

[0005] Based on the above technical ideas, the technical solution adopted by the present invention is as follows: An optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions, comprising the following steps: Step S1: Fetching and decoding stage. This step includes a fetching link and a decoding link. In the fetching link, instructions are read from memory in program order and stored in the instruction queue. The decoding link includes parsing the instructions to identify whether they are vector mask instructions (through vm) or vector length control instructions (through vta). If the instruction requires the old value of the destination register (controlled by the mask or length), mark the instruction with an old value dependency flag. Step S2: Register renaming stage. This step includes an operand mapping link and a destination operand mapping link. The operand mapping link includes obtaining the currently mapped physical register numbers for the operands of the instruction (such as vs1, vs2) from the register renaming table. The destination operand mapping link includes allocating a new physical register (such as P_new) from the free list to store the result of the instruction. Step S3: Issue control stage. This step includes a dependency check link and an issue decision link. The dependency check link includes operand checks to confirm that the physical registers of all operands are ready (traditional logic), and old value dependency checks to query the current ready delay value of the old value physical register P_old. The issue decision link includes allowing the instruction to be issued and carrying an old value acquisition method flag (divided into three categories). If the old value delay = 2, forward data from the last cycle of the operation stage (EX3). If the old value delay = 1, forward data from the write-back stage (WB). If the old value delay = 0, directly read from the physical register. S4 Execution Phase Steps, which include an old value acquisition section and an operation and result merging section. The old value acquisition section includes selecting a data source according to the old value acquisition method identifier carried in the emission phase, obtaining the old value from the forwarding network inside the execution unit in the EX2 phase, obtaining the old value from the write-back phase forwarding network in the EX3 phase, and directly reading the old value from the register file; the operation and result merging section includes the execution unit calculating a new result according to the instruction function (such as vector addition). At the last cycle of the operation (EX3), according to the values of vm and vector register v0, the old and new values are selected. According to vlen and vta, it is determined whether to use the old value or all-ones value for the elements exceeding the length, generating the final result and preparing to write back to the new physical register P_new; S5 Write-back and Commit Phase Steps, which include a write-back result section and an in-order commit section. The write-back result section includes writing the final result into the new physical register P_new and setting its ready latency counter to 0 (indicating that the data is available). The in-order commit section includes committing the instruction results in program order, updating the register renaming table, and releasing the old physical register P_old to the free list; S6 Physical Register Ready Latency Maintenance Steps, which include a continuous operation section. The continuous operation section includes automatically decrementing the ready latency counter of all physical registers by 1 (until it reaches zero). For example, for an instruction with a total latency of 4 cycles, its counter changes as 4 → 3 → 2 → 1 → 0.

[0006] For further limitation of the above technical solution, in the S2 register renaming phase steps, this step further includes an old value physical register binding section and an initializing ready latency counter section. The old value physical register binding section includes obtaining the old physical register currently mapped to the destination register (such as P_old) from the renaming table and recording it in the instruction information. The initializing ready latency counter section includes initializing the ready latency counter of the new physical register P_new to the total execution cycles of this instruction (for example, 3 cycles in the EX phase + 1 cycle in the WB phase → total latency of 4 cycles).

[0007] For further limitation of the above technical solution, in the S1 instruction fetch and decode phase steps, the following operation details are included in the instruction fetch section: The program counter (PC) sequentially accesses the instruction cache (I-Cache) and prefetches instructions in batches according to the basic block; The fetched instructions are stored in the instruction queue (Instruction Buffer), and the queue depth is designed according to the number of pipeline stages of the processor (such as 16 - 32 entries).

[0008] For further limitation of the above technical solution, in the S1 instruction fetch and decode phase steps, the following operation details are included in the decode section: Parse the instruction field, identify the RISC-V vector instruction format (OP-V type), extract the control bits of vm (mask enable bit), vta (tail element reservation configuration), and vma (mask element reservation configuration), and determine the logical register numbers of the operands (vs1, vs2, vd). Dependency identification setting: If vm = 1 or the vta / vma configuration requires retaining the old value, set the old value dependency flag (OVDF) of the instruction. The decoded micro-operation (μOp) includes the operation type, operand mapping, and OVDF flag, and is sent to the rename stage. Record the dependency type (mask dependency, length dependency, or both).

[0009] Further limitation of the above technical solution, in the step of the S2 register rename stage, the following operation details are included in the physical register mapping link of the operand: Query the rename table to obtain the current physical register numbers (such as P1, P2) of the operands (vs1, vs2). If the source register is not ready (the ready latency counter of the corresponding physical register > 0), mark the operand as NotReady. The following operation details are included in the destination operand mapping link: Allocate an idle physical register (such as P_new) from the free list as the new mapping of the destination register vd, update the rename table, and map the logical register vd to P_new.

[0010] Further limitation of the above technical solution, in the step of the S2 register rename stage, the following operation details are included in the old value physical register binding link: Obtain the old mapped physical register (P_old) of vd from the rename table (i.e., the physical register before renaming, such as P_old), and append the P_old number to the micro-operation of the instruction for use in subsequent stages. In the step of the S2 register rename stage, the following operation details are included in the initialization of the ready latency counter link: Determine the total number of latency cycles according to the instruction type (such as vector addition, mask operation) (for example, 3 beats in the EX stage + 1 beat in the WB stage → total latency of 4 beats), initialize the counter of P_new to the total latency value (such as 4), and the counter automatically decrements with each clock cycle.

[0011] Further limitation of the above technical solution, in the step of the S3 issue control stage, the following operation details are included in the operand ready check link: Check whether the physical registers of all operands (P1, P2) are ready (counter = 0). If any operand is not ready, the instruction is temporarily stored in the issue queue and retried in subsequent cycles; The steps in the S3 issue control stage, and the following operation details are included in the old value dependency check link in this step: Query the current counter value (such as L_old) of the old value physical register P_old, and calculate the allowable issue condition: allowable issue condition = L_old ≤ (total delay of this instruction - 2).

[0012] Further limitation of the above technical solution, the steps in the S4 execution stage, and the following operation details are included in the old value acquisition link in this step: In the EX2 stage, obtain the old value from the middle pipeline stage of the execution unit (such as the result of the EX1 stage). The forwarding network directly passes the intermediate result of P_old to the EX2 stage of the current instruction. In the EX3 stage, obtain the old value from the forwarding bus in the write-back stage (WB).

[0013] The steps in the S4 execution stage, and the following operation details are included in the operation and result merging link in this step: Calculate the new value, the execution unit completes the vector operation (such as vs1 + vs2), generates a temporary result, and mask control (if vm = 1); Read the mask bits of the vector register v0, select to retain the old value or use the new value for each element, vector length control (vlen and vta), determine the effective element range according to vlen, and for elements outside the range, retain the old value or set to all 1 according to the vta configuration. At the last cycle of the EX3 stage, write the final result after mask and length control into the temporary register to prepare for write-back.

[0014] The steps in the S5 write-back and commit stage, and the following operation details are included in the result write-back to the physical register link in this step: Write the final result into the destination physical register P_new, set the ready delay counter of P_new to 0, and mark the data as ready; The steps in the S5 write-back and commit stage, and the following operation details are included in the sequential commit and resource release link in this step: Commit the instructions in program order, update the architectural state, release the old physical register, release P_old back to the free list (provided that subsequent instructions no longer depend on its old value), update the rename table, and ensure that the mapping of the logical register vd points to P_new.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. Without considering other data dependencies, the vector instruction in this solution that depends on the previous instruction as the old value of the destination operand of this instruction can be issued 2 clock cycles earlier than the previous solution by reducing the latency of this instruction; a dynamic latency calculation and early issue mechanism is proposed, allowing instructions to be issued early when the old value is not fully ready, and obtaining the old value through the data forward mechanism in the later stage of the operation to reduce the waiting time. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 is the timing diagram of the optimization method of the present invention; Figure 2 is the improved timing diagram of the optimization method of the present invention. Detailed Description of the Embodiments

[0018] The following combines the attached Figure 1-2 to further elaborate on the present invention.

[0019] Embodiment 1: This embodiment provides an optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions. As Figure 1-2 shown, it includes the following steps: Step S1 of the instruction fetching and decoding stage, which includes an instruction fetching link and a decoding link. The instruction fetching link reads instructions from the memory in program order and stores them in the instruction queue; the decoding link includes parsing the instructions to identify whether they are vector mask instructions (through vm) or vector length control instructions (through vta). If the instruction requires the old value of the destination register (controlled by the mask or length), mark that this instruction needs to carry the old value dependency identifier. Step S2 of the register renaming stage, which includes an operand mapping link and a destination operand mapping link. The operand mapping link includes obtaining the currently mapped physical register numbers for the operands of the instruction (such as vs1, vs2) from the register renaming table; the destination operand mapping link includes allocating a new physical register (such as P_new) from the Free List to store the result of the instruction. Steps of the S3 emission control stage, which include a dependency check link and an emission decision link. The dependency check link includes an operand check to confirm that the physical registers of all operands are ready (traditional logic), and an old value dependency check to query the current ready latency value of the old value physical register P_old. The emission decision link includes allowing the instruction to be emitted and carrying the old value acquisition method identifier (divided into three categories). If the old value latency = 2, forward the data from the last cycle of the operation stage (EX3). If the old value latency = 1, forward the data from the write-back stage (WB). If the old value latency = 0, directly read from the physical register. Steps of the S4 execution stage, which include an old value acquisition link and an operation and result merging link. The old value acquisition link includes selecting the data source according to the old value acquisition method identifier carried in the emission stage, obtaining the old value from the forward network inside the execution unit at the EX2 stage, obtaining the old value from the forward network in the write-back stage at the EX3 stage, and directly reading the old value from the physical register in the register file. The operation and result merging link includes the execution unit calculating the new result according to the instruction function (such as vector addition). At the last cycle of the operation (EX3), select the old and new values according to the values of vm and the vector register v0, and according to vlen and vta, decide whether to use the old value or all 1 values for the elements exceeding the length, generate the final result, and prepare to write back the new physical register P_new. Steps of the S5 write-back and commit stage, which include a write-back result link and an in-order commit link. The write-back result link includes writing the final result into the new physical register P_new and setting its ready latency counter to 0 (indicating that the data is available). The in-order commit link includes committing the instruction results in program order, updating the register renaming table, and releasing the old physical register P_old to the free list. Steps of the S6 physical register ready latency maintenance, which include a continuous operation link. The continuous operation link includes automatically decrementing the ready latency counter of all physical registers by 1 (until it reaches zero). For example, for an instruction with a total latency of 4 cycles, its counter changes as 4 → 3 → 2 → 1 → 0.

[0020] The steps of the S2 register renaming stage also include an old value physical register binding link and an initializing ready latency counter link. The old value physical register binding link includes obtaining the old physical register (such as P_old) currently mapped to the destination register from the renaming table and recording it in the instruction information. The initializing ready latency counter link includes initializing the ready latency counter of the new physical register P_new to the total execution cycles of the instruction (for example, 3 cycles in the EX stage + 1 cycle in the WB stage → total latency of 4 cycles).

[0021] In the steps of the S1 fetch and decode stage, the following operation details are included in the fetch link: The program counter (PC) accesses the instruction cache (I-Cache) sequentially and prefetches instructions in batches according to basic blocks; The fetched instructions are stored in the instruction buffer, and the queue depth is designed according to the number of pipeline stages of the processor (such as 16 - 32 entries).

[0022] In the instruction fetching and decoding stage step S1, the following operation details are included in the decoding link: Parse the instruction fields, identify the RISC-V vector instruction format (OP-V type), extract the control bits of vm (mask enable bit), vta (tail element reservation configuration), and vma (mask element reservation configuration), and determine the logical register numbers of the operands (vs1, vs2, vd); Dependency flag setting: If vm = 1 or the vta / vma configuration requires retaining the old value, set the old value dependency flag (Old Value Dependency Flag, OVDF) of the instruction. The decoded micro-operation (μOp) includes the operation type, operand mapping, and OVDF flag, and is sent to the renaming stage; Record the dependency type (mask dependency, length dependency, or both).

[0023] In the instruction fetching and decoding stage step S1, the following operation details are included in the decoding link: Parse the instruction fields, identify the RISC-V vector instruction format (OP-V type), extract the control bits of vm (mask enable bit), vta (tail element reservation configuration), and vma (mask element reservation configuration), and determine the logical register numbers of the operands (vs1, vs2, vd); Dependency flag setting: If vm = 1 or the vta / vma configuration requires retaining the old value, set the old value dependency flag (Old Value Dependency Flag, OVDF) of the instruction. The decoded micro-operation (μOp) includes the operation type, operand mapping, and OVDF flag, and is sent to the renaming stage; Record the dependency type (mask dependency, length dependency, or both).

[0024] In the register renaming stage step S2, the following operation details are included in the physical register mapping link of the operands: Query the renaming table to obtain the current physical register numbers (such as P1, P2) of the operands (vs1, vs2). If the source register is not ready (the ready latency counter of the corresponding physical register > 0), mark the operand as NotReady; The following operation details are included in the destination operand mapping link: Allocate an idle physical register (such as P_new) from the idle list as the new mapping of the destination register vd, update the rename table, and map the logical register vd to P_new.

[0025] In the S2 register renaming stage step, the following operation details are included in the old value physical register binding link: Obtain the old mapped physical register (P_old) of vd from the rename table (i.e., the physical register before renaming, such as P_old), and append the number of P_old to the micro-operations of the instruction for use in subsequent stages; In the S2 register renaming stage step, the following operation details are included in the initialization of the ready latency counter link: Determine the total latency cycles according to the instruction type (such as vector addition, mask operation) (e.g., 3 beats in the EX stage + 1 beat in the WB stage → total latency of 4 beats), initialize the counter of P_new to the total latency value (e.g., 4), and the counter automatically decrements with each clock cycle.

[0026] In the S3 issue control stage step, the following operation details are included in the operand ready check link: Check whether the physical registers of all operands (P1, P2) are ready (counter = 0). If any operand is not ready, the instruction is temporarily stored in the issue queue and retried in subsequent cycles; In the S3 issue control stage step, the following operation details are included in the old value dependency check link: Query the current counter value of the old value physical register P_old (such as L_old), and calculate the allowable issue condition: allowable issue condition = L_old ≤ (total latency of this instruction - 2).

[0027] In the S4 execution stage step, the following operation details are included in the old value acquisition link: In the EX2 stage, obtain the old value from the middle pipeline stage of the execution unit (such as the result of the EX1 stage). The forwarding network directly passes the intermediate result of P_old to the EX2 stage of the current instruction. In the EX3 stage, obtain the old value from the forwarding bus in the write-back stage (WB).

[0028] In the S4 execution stage step, the following operation details are included in the operation and result merging link: Calculate the new value, the execution unit completes the vector operation (such as vs1 + vs2), generates a temporary result, and mask control (if vm = 1); Read the mask bits of vector register v0, select to retain the old value or use the new value element by element, vector length control (vlen and vta), determine the valid element range according to vlen, and for elements outside the range, retain the old value or set to all 1 according to the vta configuration. At the last cycle of the EX3 stage, write the final result after mask and length control to a temporary register, preparing for write-back.

[0029] The steps in the S5 write-back and commit stage. In the step of writing the result back to the physical register, the following operation details are included: Write the final result to the destination physical register P_new, set the ready latency counter of P_new to 0, and mark the data as ready; The steps in the S5 write-back and commit stage. In the step of sequential commit and resource release, the following operation details are included: Commit the instructions in program order, update the architectural state, release the old physical register, release P_old back to the free list (provided that subsequent instructions no longer depend on its old value), update the rename table, and ensure that the mapping of the logical register vd points to P_new.

[0030] Embodiment 2: This embodiment provides an optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions. As Figure 1-2 shown, the following steps are further included: 1. The instruction fetch stage is the same as that of a traditional out-of-order processor.

[0031] 2. In the decoding stage, according to the values of the vm field, vector length, and control register vtype of the instruction, determine whether the result of this instruction needs to use the old value of the destination register. If so, carry a flag for reading the old value.

[0032] 3. In the register renaming stage, in addition to finding the physical register mapping for the operands in the instruction and mapping the logical register of the target register of the instruction to an idle physical register, it is also necessary to check whether the instruction carries a flag for reading the old value of the target register. If it carries this flag, it is also necessary to use the register renaming table to find the old physical register mapping of the logical register of the target register.

[0033] 4. In the instruction scheduling and instruction issue stage, if the instruction carries a flag, when checking whether the operands are ready, it is also necessary to check whether the value of the old physical register of the target register is ready. Only when all these dependent data are ready can the instruction be issued.

[0034] 5. During the instruction execution stage, if the instruction carries an identifier, both the operand and the old value of the destination operand need to be fetched, and then participate in the operation. According to the mask and vector length, it is determined whether each element of the destination operand is obtained from the old value or from the result obtained by calculating the operand. Finally, the result is written back to the newly mapped physical register of the destination operand.

[0035] The instruction completion and instruction commit stages are the same as those of a traditional out-of-order processor.

[0036] Embodiment 3: This embodiment provides an optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions. As Figure 1-2 shown, it also includes the out-of-order processor instruction execution flow: 1. Instruction Fetch The processor fetches instructions from memory in the order of the program counter (PC). The fetched instructions are stored in the instruction queue for further processing.

[0037] 2. Instruction Decode The fetched instructions are decoded to determine the type of the instruction, the operands, and the execution units required. If the required operands are ready, the instruction will enter the execution stage; otherwise, it will wait for the operands to be ready.

[0038] 3. Register Renaming To avoid register conflicts between instructions (such as read-after-write or write-write conflicts), the processor performs register renaming. This means that logical registers are mapped to physical registers to eliminate potential data dependency problems. There is a set of physical registers in the processor's architecture. The processor uses a register renaming table (Renaming Table) to find the physical register mapping for the operands in the instruction, and uses a register free list (Freelist) to map the logical register of the destination register of the instruction to an idle physical register. In this way, even if different instructions modify the same logical register, they can still use different physical registers, thus avoiding data dependency conflicts.

[0039] 4. Instruction Dispatch After the instruction is decoded and the necessary operands are obtained, it is dispatched to the appropriate execution unit. Due to the characteristics of the out-of-order processor, the instruction decides when to execute according to the available execution units and resources, rather than in program order.

[0040] 5. Instruction Issue After instruction scheduling, it will be cached in an instruction queue until its operands are ready and there are available execution units to execute it.

[0041] 6. Out-of-Order Execution In an out-of-order processor, if the dependencies of an instruction are satisfied, it can be executed earlier even if it is located later in the program. The processor will look for idle execution units and execute the instruction in one of them. The execution order of these instructions is determined dynamically rather than in a fixed order.

[0042] 7. Instruction Completion After the execution unit completes the instruction operation, the result will be written back to the register or memory. An out-of-order processor may save the result of the instruction in a temporary buffer first instead of writing it back to the register immediately to ensure the correct execution order of the instructions.

[0043] 8. Instruction Commit Finally, the results of the instructions will be committed in the original order of the program to ensure the correct semantics of the program. If an exception or branch error occurs, the processor will withdraw the out-of-order executed instructions (by rolling back) and restore to the correct state.

[0044] Embodiment 4: This embodiment provides an optimization method for improving the parallelism and performance of hardware executing RISC-V vector instructions, as Figure 1-2 shown, and further includes the following instruction execution stages: In the execution stage of the instruction, the old value of the destination register does not need to participate in the operation. It only needs to read back the old value before the last cycle of the operation, select the final result according to the mask in the last cycle of the operation, and write it back to the physical register in the next cycle. Therefore, the destination register does not need to be fetched before entering the arithmetic unit like the operands. When querying whether the old value of the destination operand is ready in the instruction scheduling and issuing stages, information on whether the instruction can be issued earlier can be obtained by recording the latency of data readiness.

[0045] As Figure 1As shown in the figure, assume that the old values of the destination operands to be fetched by OP2, OP3, and OP4 in the figure all come from the result of the destination operand of OP1 (usually this is not the case. If the logical encodings of the destination registers of OP2, OP3, and OP4 are the same, then OP3 depends on OP2, and OP4 depends on OP3. In this figure, it can be considered that OP2, OP3, and OP4 are the cases where the same instruction depending on the result of OP1 arrives at different times), and all other dependencies are satisfied. In the second-to-last cycle when OP1's operation ends (EX2 in the above figure), the message that the destination vector register of this instruction is ready will be broadcast to the instruction queue within the instruction scheduler. Then the next instruction can be issued as early as T3 (i.e., the case of OP2), and it will obtain the data from the last cycle (EX3) of the previous instruction's operation; if it is issued at T4, it will obtain the data from the write-back cycle of the previous instruction; if it is issued after T5, it can read the data from the physical register.

[0046] Embodiment 4: This embodiment provides an optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions. As Figure 1-2 shown, it also includes an improvement method for Embodiment 3, including the following steps: In the execution stage of the instruction, the old value of the destination register does not need to participate in the operation. It only needs to read back the old value before the last cycle of the operation ends, select the final result according to the mask in the last cycle of the operation, and write it back to the physical register in the next cycle. Therefore, the destination register does not need to be fetched before entering the operation unit like an operand. In the instruction scheduling and issue stages, when querying whether the old value of the destination operand is ready, information on whether the instruction can be issued early can be obtained by recording the latency of data readiness.

[0047] As Figure 2 shown, assume that the old values of the destination operands to be fetched by OP2, OP3, and OP4 in the figure all come from the result of the destination operand of OP1 (usually this is not the case. If the logical encodings of the destination registers of OP2, OP3, and OP4 are the same, then OP3 depends on OP2, and OP4 depends on OP3. In this figure, it can be considered that OP2, OP3, and OP4 are the cases where the same instruction depending on the result of OP1 arrives at different times), and all other dependencies are satisfied. A queue of physical register data readiness latency will be maintained within the instruction scheduling module. In the cycle after OP1 is issued, the readiness latency corresponding to the physical number of OP1's destination register will be set to the latency of this instruction (including the latency of execution (EX) and write-back (WB)). Thereafter, this readiness latency will be decremented by one each cycle until it reaches 0. The latency of all instructions in the figure is 4 (3 cycles for calculation and 1 cycle for write-back).

[0048] When subsequent instructions enter the instruction scheduling / issue module, it is queried whether the old value of the target operand is ready, which can be obtained by calculating the ready latency of the target register in this cycle and the latency of the instruction itself. Subtracting 2 from the instruction latency (because at least one cycle for the operation and one cycle for write-back are included, so the instruction latency must be greater than 2) can obtain how many cycles later the old value of the target operand needs to be read if the instruction in this cycle is issued. Subtracting the latency for the instruction to read the old value from the ready latency of the target register (if the result is negative, the result is considered 0) can obtain the result of whether the instruction can be issued and where to fetch the data.

[0049] When the calculated value is less than 2, it can be considered that the old value of the target operand is ready. If other dependencies are also satisfied, this instruction can be issued. When the instruction is issued, the value calculated in this cycle will be carried to the subsequent pipeline, and the old value will be read two cycles before this instruction reaches the write-back stage (for the instruction latency of 4 in the figure, so the old value will be read in the EX2 cycle). When the value is 2, data is forwarded from the last cycle of the operation; when the value is 1, data is forwarded from the write-back; when the value is 0, data is fetched from the register file.

[0050] The instruction can be issued 2 clock cycles earlier than the previous solution. The above content is a further detailed description of the present invention in combination with specific preferred implementation solutions, facilitating those skilled in the art of this technology to understand and apply the present invention. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions.

Claims

1. An optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions, characterized in that: The following steps are involved: S1 instruction fetch and decode phase step, which includes instruction fetch and decoding. The instruction fetch phase reads instructions from the memory in program order and stores them in the instruction queue. The decoding phase includes parsing instructions to identify whether they are vector mask instructions or vector length control instructions. If the instruction requires the old value of the destination register, the instruction is marked to carry the old value dependency flag. S2 register renaming phase step, which includes an operand mapping link and a destination operand mapping link. The operand mapping link includes obtaining the currently mapped physical register number from the register renaming table for the operand of the instruction; the destination operand mapping link includes allocating a new physical register from the free list for storing the result of the instruction; S3 is the emission control phase step, which includes a dependency check link and an emission decision link. The dependency check link includes an operand check to confirm that the physical registers of all operands are ready, an old value dependency check to query the current ready delay value of the old value physical register P_old; the emission decision link includes allowing the instruction to be emitted and carrying the old value acquisition method identifier. If the old value delay = 2, data is forwarded from the last beat of the operation phase; if the old value delay = 1, data is forwarded from the write-back phase; if the old value delay = 0, it is directly read from the physical register; S4 execution phase step, which includes the old value acquisition link and the operation and result merging link. The old value acquisition link includes selecting the data source according to the old value acquisition method identifier carried in the emission phase, acquiring the old value from the forward network inside the execution unit in the EX2 phase, acquiring the old value from the forward network in the write-back phase in the EX3 phase, and directly reading the old value from the physical register in the register file; the operation and result merging link includes the execution unit calculating the new result according to the instruction function, selecting the new and old values ​​according to the values ​​of vm and vector register v0 at the last beat of the operation, and determining whether the elements exceeding the length use the old value or the all-1 value according to vlen and vta, generating the final result, and preparing to write back to the new physical register P_new; S5 write back and commit phase step, which includes a write back result phase and a sequential commit phase. The write back result phase includes writing the final result into a new physical register P_new and setting its ready delay counter to 0. The sequential commit phase includes committing the instruction result in program order, updating the register renaming table, and releasing the old physical register P_old to the free list. S6 is a physical register ready delay maintenance step, which includes a continuous operation link, and the continuous operation link includes automatically decrementing the ready delay counters of all physical registers by 1.

2. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 1, characterized in that: The S2 register renaming phase step also includes an old value physical register binding link and an initialization ready delay counter link. The old value physical register binding link includes obtaining the old physical register currently mapped to the destination register from the renaming table and recording it in the instruction information. The initialization ready delay counter link includes initializing the ready delay counter of the new physical register P_new to the total number of execution cycles of the instruction.

3. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 2, characterized in that: The S1 instruction fetch and decoding stage step includes the following operation details in the instruction fetch link: The program counter accesses the instruction cache sequentially and prefetches instructions in batches by basic block; The fetched instructions are stored in the instruction queue, and the queue depth is designed according to the number of processor pipeline stages.

4. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 3, characterized in that: The S1 instruction fetch and decoding stage step includes the following operation details in the decoding link: Parse the instruction field, identify the RISC-V vector instruction format, extract the vm, vta, and vma control bits, and determine the logical register number of the operand; Dependency flag setting: if vm=1 or the vta / vma configuration requires retaining the old value, the old value dependency flag of the instruction is set. The decoded micro-operation includes the operation type, operand mapping, and OVDF flag, and is sent to the renaming stage. Records dependency types.

5. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 4, characterized in that: The S2 register renaming phase step includes the following operation details in the operand physical register mapping link: Query the renaming table to obtain the current physical register number of the operand. If the source register is not ready, mark the operand as not ready. The destination operand mapping link includes the following operation details: Allocate a free physical register from the free list as the new mapping of the destination register vd, update the rename table, and map the logical register vd to P_new.

6. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 5, characterized in that: The S2 register renaming phase step includes the following operation details in the old value physical register binding link: Get the old mapped physical register of vd from the rename table and append the P_old number to the micro-op of the instruction for use in subsequent stages; The S2 register renaming phase step includes the following operation details in the initialization of the ready delay counter: The total delay cycle number is determined according to the instruction type, and the counter of P_new is initialized to the total delay value, and the counter is automatically decremented with each clock cycle.

7. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 6, characterized in that: The S3 emission control phase step includes the following operation details in the operand readiness check link: Check whether the physical registers of all operands are ready. If any operand is not ready, the instruction is temporarily stored in the issue queue and will be retried in the subsequent cycle; The S3 emission control phase step includes the following operation details in the old value dependency check link: Query the current counter value of the old value physical register P_old and calculate the condition for enabling issuance: the condition for enabling issuance = L_old ≤ the total delay of this instruction - 2.

8. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 7, characterized in that: The S4 execution phase step includes the following operation details in the old value acquisition link: In the EX2 stage, the old value is obtained from the middle pipeline stage of the execution unit, and the forward network passes the intermediate result of P_old directly to the EX2 stage of the current instruction. In the EX3 stage, the old value is obtained from the forward bus of the write-back stage.

9. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 7, characterized in that: The S4 execution phase step includes the following operation details in the calculation and result merging process: Calculate new values, execute units to complete vector operations, generate temporary results, and mask control; Read the mask bit of vector register v0, and choose to keep the old value or use the new value element by element. The vector length is controlled. The valid element range is determined according to vlen. Elements out of the range retain the old value or are set to all 1 according to the vta configuration. In the last beat of the EX3 stage, the final result after mask and length control is written to the temporary register and prepared to be written back.

10. The optimization method for improving the parallelism and performance of hardware execution of RISC-V vector instructions according to claim 7, characterized in that: The step of writing back and committing in S5 includes the following operation details in the step of writing the result back to the physical register: Write the final result into the destination physical register P_new, set the ready delay counter of P_new to 0, and mark the data as ready; The S5 write-back and commit phase includes the following operation details in the sequential commit and resource release process: Submit instructions in program order, update the architectural state, release the old physical registers, release P_old back to the free list, and update the rename table to ensure that the mapping of logical register vd points to P_new.

Citation Information

Patent Citations

  • Six-stage pipeline processor based on RISC-V instruction set

    CN114721724A

  • Four-emission RISC-V processor micro-architecture and working method thereof

    CN115454504A

  • Instruction out-of-order execution method

    CN115599445A

  • Fine-grained lock-step fault-tolerant superscale out-of-order processor design method and system

    CN117667477A

  • Low-power-consumption single-emission out-of-order execution RISC-V processor and instruction processing method

    CN119718430A

Cited By

  • Vector configuration instruction implementation method and system and storage medium

    CN120335869A

  • Implementation method, system and storage medium of vector configuration instruction

    CN120335869B

  • Vector microoperation splitting method and system based on RISC-V

    CN121233174A

  • Dynamic management method and system for RISC-V vector register pressure

    CN121255288A

  • Realization method and device of mask vector writing instruction based on RISC-V

    CN121326406A