RISC-V Vector Store Hardware Resource Release Method and Device
By introducing vector storage decommissioning table (VSRB) to assist in reordering cache, optimize the hardware resource release of RISC-V vector store instructions, the problem of difficult hardware resource release in vector store instructions in high-performance CPU cores is solved, and more efficient hardware resource utilization and performance improvement is achieved.
Patent Information
- Application Number
- CN202411158735.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-08-22
AI Technical Summary
In high-performance, out-of-order execution of RISC-V CPU cores, there are difficulties in releasing hardware resources of RISC-V vector store instructions, resulting in improper allocation of store buffer (SB) items and retire processes, resulting in waste of hardware resources and degradation of performance.
By introducing a vector storage decommissioning table (VSRB) to assist in reordering caches, the vector storage decommissioning items are used as granularity to perform hardware resources retire and release. The specific method includes decoding and generating multiple STA and STD micro-instruction pairs when the vector store instruction is executed, and allocating a hardware resource item to each pair of micro-instructions; optimizing the return and release of the SB item through the preset number of consecutive maskless store operation items stored in the VSRB item.
This method effectively reduces the waste of hardware resources, improves the utilization efficiency of store buffers, avoids the decline in CPU performance caused by the shortage of SB items, and does not need to significantly increase the hardware resources required for SB items.
Smart Images

Figure CN119127485B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technologies, and particularly to a method and apparatus for releasing RISC-V vector store hardware resources. Background Art
[0002] The store instruction in an instruction set is used to write data into memory, and its implementation in a high-performance, out-of-order execution CPU core is different from other instructions: the execution of other non-store instructions can be out-of-order (that is, when there are older instructions that have not been executed) and in a speculative state (this instruction may be on an incorrect branch path), and the result is written into a physical register; if speculation is incorrect, the content in the physical register file can be restored to the correct state and then execution can start from the correct branch path. However, once a store instruction writes data into memory / cache, it is very difficult / impossible to restore. Therefore, a store instruction will only write its data into memory / cache after it is retired, that is, after entering the architectural state. There may be multiple clock cycles between the execution of the store instruction and its writing of data into memory. During this period, the memory address and the written data of the store are generally stored in a hardware resource, the store buffer (SB). To ensure program correctness, store memory writes must be strictly in program order. Therefore, the store buffer is generally a First In First Out (FIFO) circular buffer. Since the allocation and release of store buffer entries are both sequential, the allocation must be in the front-end sequential stage of the CPU core pipeline, such as the dispatch pipeline stage.
[0003] A store operation consists of two independent sub-operations: 1) STA: Calculate the memory write address; 2) STD: Read the data to be written to memory from a register. In a high-performance, out-of-order execution CPU core, these two operations are usually decomposed into two independent micro-instructions for execution. Since the level 1 data cache (L1D) is shared by the SB and other pipelines (such as load and second-level cache read / write, etc.), competition is required to obtain it. Therefore, there is often an interval of several clock cycles between when the SB item is allowed to write to the L1D (i.e., retire the SB item) and when it actually writes to the L1D and releases the SB item. The SB write to the L1D operation is responsible for the internal logic circuit of the SB. Most RISC-V vector instructions support masked computing, that is, the instruction only writes to the masked (i.e., the corresponding mask bit is 1) fields in the destination register, and the values of the mask-off (i.e., the corresponding mask bit is 0) fields do not change. For these instructions, the vm bit in the opcode identifies whether masked computing is performed: vm = 1: No masked computing, and the calculation results of all fields are written to the register. vm = 0: Masked computing, and the mask value comes from the architectural register v0. At this time, each bit starting from the low position in v0 identifies a field. When v0[i] is 1, the calculation result can be written to the corresponding field, otherwise it is not written.
[0004] In the case of having a mask (vm = 0), the mask values of a pair of STA and STD micro-instructions in a vector store instruction are obtained in the execution stage. If mask off (v0[i] = 0), then both of these micro-instructions are converted to NOPs, and their data is not written to memory. Since the allocation of the SB item for the store operation is performed at an earlier Dispatch pipeline stage, and the value in the Mask register is not known at this time, it is necessary to allocate an SB item for each store operation, even though the mask-off store operation does not use the allocated SB item.
[0005] In a high-performance, out-of-order execution CPU core, although instruction execution is out-of-order, instruction retirement is in-order, which can ensure the correctness of program execution. Instruction in-order retirement is achieved through a reorder buffer (ROB). The ROB is a first-in-first-out (FIFO) circular buffer. An instruction is assigned a ROB entry at the end of the valid ROB entries (ROB tail) at the dispatch pipeline stage of in-order execution, and retirement always starts from the beginning position (ROB head) of the valid ROB entries. Inside a RISC-V CPU, an instruction is often decoded into one or more micro-ops or macro-ops. Correspondingly, each entry in the ROB can choose to save an instruction, a micro-op, or a macro-op. To improve the efficiency of the ROB, most ROB architectures adopt the method of having one entry corresponding to one instruction.
[0006] A simple solution to extend the existing scalar ROB and SB to support RISC-V vector store instructions. The allocation, retirement, and release of SB are still carried out at the instruction granularity. However, when this solution is extended to RISC-V vector instructions, the following problems will be faced:
[0007] SB entry allocation problem: In a CPU architecture that only supports scalar instructions, the SB usually doesn't require many entries, and generally 16 - 48 entries are enough. However, as mentioned before, when VLEN = 128b, a single vector store instruction may require up to 128 SB entries (and more SB entries are needed when VLEN is larger), otherwise the instruction cannot be dispatched. The solution of increasing the number of SB entries to meet the maximum demand of a single vector instruction will cause the SB to be too large, resulting in waste of hardware resources and increased access latency to the SB, ultimately leading to a decrease in the overall performance of the CPU and an increase in area and power consumption.
[0008] SB entry retirement and release problem: Since a single vector store instruction occupies multiple SB entries, if these SB entries can only be released when the instruction retires, then the instruction will occupy multiple SB entries for too long. Considering that a single vector store instruction requires multiple SB entries, it will cause the front-end pipeline of the CPU to stall due to a shortage of SB entries, thus resulting in a decrease in CPU performance.
[0009] The masked store operation still occupies SB entries. Coupled with the fact that the SB entries are only released when the store instruction retires, it further causes waste of SB resources.
[0010] Regarding the above problems, no effective solutions have been proposed yet. Summary of the Invention
[0011] The embodiments of this specification provide a method and apparatus for releasing RISC-V vector store hardware resources. The optimized SB allocation / retire scheme better supports RISC-V vector store instructions, and specifically optimizes the early retirement of SB entries for mask off store operations. It achieves the same performance as the simple extension scheme with fewer hardware resources and fewer SB entries.
[0012] The embodiments of this specification provide a method for releasing RISC-V vector store hardware resources, including:
[0013] In response to a vector store instruction, decode according to the number of store operations to generate multiple pairs of vector STA and STD microinstructions, and allocate a hardware resource entry for each pair of vector STA and STD microinstructions; assist the reorder buffer through a preset vector store retirement table to retire the hardware resource entries in units of vector store retirement entries in the vector store retirement table, so as to release the hardware resources corresponding to the hardware resource entries; wherein, through the number of consecutive unmasked store operation items stored in the vector store retirement entry, when retiring the hardware resource entry corresponding to the vector store retirement entry, retire all the hardware resource entries of the unmasked store operations of the vector store instruction after the vector store retirement entry.
[0014] In one embodiment, decoding according to the number of store operations of the vector store instruction to generate multiple pairs of STA and STD microinstructions includes:
[0015] Add an IS_LAST field to the vector STA microinstruction generated by the vector store instruction, and define the last bit information of the vector STA microinstruction in the vector store instruction through the value of the IS_LAST field.
[0016] In one embodiment, the method further includes:
[0017] Set a vector store retirement entry identification field, a vector STA microinstruction quantity field, and a vector STD microinstruction quantity field in the reorder buffer entry; wherein, the vector store retirement entry identification field is used to configure different values according to the vector store instruction or non-vector store instruction; the vector STA microinstruction quantity field and the vector STD microinstruction quantity field are respectively used to record the quantity of vector STA microinstructions and vector STD microinstructions generated by the vector store instruction, and adjust the recorded quantity of vector STA microinstructions and vector STD microinstructions according to the execution conditions of the vector STA microinstructions and vector STD microinstructions.
[0018] In one embodiment, the method further includes:
[0019] When the source physical register value of the vector STA micro-instruction exists and there is an idle vector store retirement entry, the vector store retirement table sequentially allocates an idle vector store retirement entry for the vector STA micro-instruction; releases the corresponding vector store retirement entry when retiring a vector STA micro-instruction of a non-masked store operation; releases the corresponding vector store retirement entry when the vector STA micro-instruction of a masked store operation is completed; wherein, the vector store retirement entry includes a vector store retirement index field, a vector store instruction index field corresponding to the reorder buffer, a store operation index field in the vector store instruction, an index field of the hardware resource item corresponding to the store operation, a store operation exception record identification field, and a pre-masked store operation quantity field before the store operation.
[0020] In one embodiment, the vector store retirement table sequentially allocating an idle vector store retirement entry for the vector STA micro-instruction includes: when creating or writing a vector store retirement entry, the number of non-masked store operation items between the corresponding valid store operation and the adjacent previous valid store operation is written into the pre-masked store operation quantity field before the store operation of the vector store retirement entry; when releasing or reading a vector store retirement entry, and when it is determined that there is valid data in the adjacent next item through the pre-masked store operation quantity field before the store operation, reads the pre-masked store operation quantity field of the adjacent next item, so as to retire all the hardware resource items of the non-masked store operations of the vector store instruction after the vector store retirement entry when retiring the hardware resource item corresponding to the vector store retirement entry.
[0021] In one embodiment, the vector store retirement table sequentially allocating an idle vector store retirement entry for the vector STA micro-instruction includes: when the vector store retirement entry flags an exception, providing the corresponding field in the store operation index field in the vector store instruction of the vector store retirement entry and the exception type in the store operation exception record identification field to the reorder buffer for exception handling.
[0022] In one embodiment, a preset vector store retirement table is used to assist the reorder buffer in retiring the hardware resource items in units of vector store retirement items in the vector store retirement table, including: sequentially releasing the hardware resource items starting from the item pointed to by the preset vector store pointer head every clock cycle; sequentially allocating the hardware resource items starting from the item pointed to by the preset vector store pointer tail every clock cycle; wherein, the vector store pointer head points to the longest-lived valid item in the vector store retirement table, and the vector store pointer tail points to the first invalid item in the vector store retirement table.
[0023] In one embodiment, a preset vector store retirement table is used to assist the reorder buffer in retiring the hardware resource items in units of vector store retirement items in the vector store retirement table, including: when all the vector store retirement items corresponding to the vector STA micro-instructions of the vector store instruction are released, the vector store retirement table provides a notification signal to the reorder buffer; when all the vector STD micro-instructions of the vector store instruction are executed, the reorder buffer releases the reorder buffer item corresponding to the vector store instruction according to the notification signal.
[0024] An embodiment of this specification also provides a RISC-V vector store hardware resource release device, which includes: an allocation module and a release module; the allocation module is configured to, in response to a vector store instruction, decode and generate multiple pairs of vector STA and STD micro-instructions according to the number of store operations, and allocate a hardware resource item to each pair of vector STA and STD micro-instructions; the release module is configured to assist the reorder buffer in retiring the hardware resource items in units of vector store retirement items in the vector store retirement table to release the hardware resources corresponding to the hardware resource items; wherein, the number of consecutive unmasked store operation items stored in the vector store retirement item is used to retire all the hardware resource items of the unmasked store operations of the vector store instruction after the vector store retirement item when retiring the hardware resource item corresponding to the vector store retirement item.
[0025] An embodiment of this specification also provides a central processing device, which includes a central processing core with a reorder buffer based on instructions, micro-instructions, and macro-instructions, and the central processing device applies the above-mentioned RISC-V vector store hardware resource release method.
[0026] For a RISC-V CPU core implementing the RISC-V vector extension, the RISC-V vector store hardware resource release method and apparatus provided by the embodiments of this specification provide an optimized solution for extending the ROB and SB to support RISC-V vector store instructions. Compared with the traditional solution, this application does not require a substantial increase in the hardware resources required for the SB items, and achieves the purpose of releasing the SB items as soon as possible by retiring the SB items as soon as possible. At the same time, this application can also release the SB items corresponding to the mask off store operation as soon as possible, thereby further improving the utilization efficiency of the SB.
[0027] Referring to the following description and drawings, specific embodiments of the present invention are disclosed in detail, indicating the ways in which the principles of the present invention can be employed. It should be understood that the embodiments of the present invention are not limited in scope thereby. Features described and / or illustrated for one embodiment can be used in the same or similar way in one or more other embodiments, combined with the features in other embodiments, or replace the features in other embodiments.
[0028] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, components, steps or assemblies, but does not exclude the presence or addition of one or more other features, components, steps or assemblies. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and do not limit this specification. In the drawings:
[0030] Figure 1 is a schematic flowchart of the RISC-V vector store hardware resource release method provided by an embodiment of this application;
[0031] Figure 2 is a schematic diagram of the application effect of the RISC-V vector store hardware resource release method provided by an embodiment of this application;
[0032] Figure 3 is a schematic diagram of the application effect of the RISC-V vector store hardware resource release method provided by another embodiment of this application;
[0033] Figure 4 is a diagram of multiple pipeline stages of a CPU architecture in the prior art;
[0034] Figure 5 is a schematic flowchart of the execution stage provided by another embodiment of this application;
[0035] Figure 6 is a schematic flowchart of the release stage provided by another embodiment of this application;
[0036] Figure 7 This is a schematic structural diagram of the RISC-V vector store hardware resource release device provided by an embodiment of the present application. Detailed implementation manners
[0037] The principles and spirit of this specification will be described below with reference to several exemplary implementation manners. It should be understood that these implementation manners are given only to enable those skilled in the art to better understand and then implement this specification, and do not limit the scope of this specification in any way. On the contrary, these implementation manners are provided to make the disclosure of this specification more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0038] Those skilled in the art know that the implementation manners of this specification can be realized as a system, device, equipment, method, or computer program product. Therefore, the disclosure of this specification can be specifically realized in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0039] Combined with the description of the accompanying drawings and the specific implementation manners of the present invention, the details of the present invention can be understood more clearly. However, the specific implementation manners of the present invention described herein are only for the purpose of explaining the present invention and cannot be understood in any way as a limitation of the present invention. Under the teaching of the present invention, those skilled in the art can conceive any possible variations based on the present invention, and these should all be regarded as belonging to the scope of the present invention. It should be noted that when an element is referred to as being "disposed on" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "mounted", "connected", and "coupled" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations. The terms "vertical", "horizontal", "upper", "lower", "left", "right", and similar expressions used herein are only for the purpose of illustration and do not represent the only implementation manner.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this specification belongs. The terms used herein in this specification are only for the purpose of describing specific implementation manners and are not intended to limit this specification. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0041] An embodiment of this specification provides a method for releasing RISC-V vector store hardware resources. Figure 1 The flowchart of the RISC-V vector store hardware resource release method in an embodiment of this specification is shown. Although this specification provides method operation steps or device structures as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or non-creative labor. In steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of this specification and shown in the drawings. When the method or module structure is applied to an actual device or terminal product, it can be executed sequentially or in parallel according to the method or module structure connection shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, or even a distributed processing environment).
[0042] Please refer to Figure 1 As shown, the RISC-V vector store hardware resource release method provided by the embodiments of this specification may specifically include:
[0043] S101: In response to a vector store instruction, decode and generate multiple pairs of vector STA and STD microinstructions according to the number of store operations, and allocate a hardware resource item to each pair of vector STA and STD microinstructions;
[0044] S102: Assist the reorder buffer through a preset vector store retirement table to retire the hardware resource items in units of vector store retirement items in the vector store retirement table, so as to release the hardware resources corresponding to the hardware resource items;
[0045] Among them, through the number of consecutive unmasked store operation items stored in the vector store retirement item, when retiring the hardware resource item corresponding to the vector store retirement item, retire all the hardware resource items of the unmasked store operations of the vector store instruction after the vector store retirement item.
[0046] In this embodiment, the RISC-V vector store instruction is divided into multiple STA / STD micro-instruction pairs according to the number of store operations, and each pair of STA / STD (whether masked off or not) micro-instructions corresponds to an SB item. Thus, the problem that the vector store instruction applies for too many SB items at a time and blocks the CPU pipeline is avoided. The increased hardware table, the vector store retire buffer (VSRB), is used to assist the ROB to retire SB items in units of VSRB items. The number of consecutive masked-off store operation items is saved in the subsequent VSRB items. By optimizing the writing of the same item NOP_CNT in the VSRB but reading the NOP_CNT of the next item, when retiring the SB item corresponding to a VSRB item, all the SB items corresponding to the subsequent masked-off stores are also retired, thereby further improving the utilization efficiency of the SB items. Thus, the same performance as the simple expansion scheme is achieved with fewer hardware resources and fewer SB items.
[0047] It should be noted that in many CPU cores, the ROB does not store instructions, but micro-operations (microuop) or macro-operations (macro op), which are specific conversion formats of instructions in a CPU core: an instruction may be converted into one or more micro-operations / macro-operations, or two instructions may be merged into one micro-operation / macro-operation, etc. The ROB based on instructions / micro-operations / macro-operations is designed for scalar instructions and faces the same problems for vector instructions. Therefore, this application is also applicable to CPU cores using a ROB based on micro-operations / macro-operations. A store operation includes two independent sub-operations: 1) STA: calculate the write memory address; 2) STD: read the data to be written to memory from a register. In a high-performance, out-of-order execution CPU core, these two operations are usually decomposed into two independent micro-instructions for execution; the CPU core of this application can use this method and is also applicable to other methods, such as the same store micro-instruction executing in two pipelines for address generation and register reading; further, many high-performance, superscalar CPU cores can rename, execute, and retire multiple instructions in one clock cycle; this application is also applicable to superscalar CPU cores; this application does not make further limitations here.
[0048] In an embodiment of this application, generating multiple STA and STD micro-instruction pairs by decoding the vector store instruction according to the number of store operations includes: adding an IS_LAST field to the vector STA micro-instruction generated from the vector store instruction, and defining the last bit information of the vector STA micro-instruction in the vector store instruction through the value of the IS_LAST field.
[0049] In this embodiment, a one-bit IS_LAST field is mainly added to the vector STA micro-instruction (uop) item. If this field is 1, it indicates that the STA uop is the last STA uop of the corresponding vector store instruction. This action is mainly applied to the Decode stage of the pipeline. A vector store instruction will be decoded into at least a pair of STA / STD micro-instructions, and each pair of micro-instructions corresponds to a store operation and an SB item; the number of SB items required for a vector store instruction depends on parameters such as VLEN, the current LMUL, and SEW. For a segment store instruction, the NFIELD parameter also needs to be considered, and the IS_LAST bit is set for the generated vector STA uop: the IS_LAST of the last uop is set to 1, and the IS_LAST of the remaining uops is set to 0.
[0050] In an embodiment of the present application, the method further includes: setting a vector store retirement item identification field, a vector STA micro-instruction quantity field, and a vector STD micro-instruction quantity field in the reorder buffer item; wherein, the vector store retirement item identification field is used to configure different values according to the vector store instruction or the non-vector store instruction; the vector STA micro-instruction quantity field and the vector STD micro-instruction quantity field are respectively used to record the quantity of vector STA micro-instructions and vector STD micro-instructions generated by the vector store instruction, and adjust the corresponding recorded quantity of vector STA micro-instructions and vector STD micro-instructions according to the execution conditions of the vector STA micro-instructions and vector STD micro-instructions.
[0051] In the above embodiments, the vector store retirement item identification field is the 1-bit use_VSRB field. If the ROB item is a vector store instruction, this bit is set to 1 when the ROB item is allocated; otherwise, this bit is set to 0. A vector store instruction has at least one VSRB item. The vector STA micro-instruction quantity field is the STA_CNTDOWN field. It records the number of STA uops generated by this vector store instruction. This field is set when the ROB item is allocated. After each vector STA uop is executed, regardless of whether it is mask off or not, this field is decremented by one. When the value of this field is 0, all STA uops of this instruction have been executed. This field is multi-bit and can accommodate the maximum number of vector STA uops implemented by the CPU core. The vector STD micro-instruction quantity field is the STD_CNTDOWN field. It records the number of STD uops generated by this vector store instruction. This field is set when the ROB item is allocated. After each vector STD uop is executed, regardless of whether it is mask off or not, this field is decremented by one. When the value of this field is 0, all STD uops of this instruction have been executed; this field is multi-bit and can accommodate the maximum number of vector STA uops implemented by the CPU core.
[0052] In an embodiment of the present application, the method further includes: when the source physical register value of the vector STA micro-instruction exists and there are idle vector store retirement items, the vector store retirement table sequentially allocates an idle vector store retirement item for the vector STA micro-instruction; releasing the corresponding vector store retirement item when retiring a vector STA micro-instruction of a non-masked store operation; releasing the corresponding vector store retirement item when finishing executing a vector STA micro-instruction of a masked store operation; wherein, the vector store retirement item includes a vector store retirement index field, a vector store instruction index field corresponding to the reorder buffer, a store operation index field in the vector store instruction, an index field of the hardware resource item corresponding to the store operation, a store operation exception record identification field, and a pre-store masked store operation quantity field.
[0053] In the above embodiments, VSRB is the same as ROB. VSRB is also a FIFO circular buffer. Each valid (that is, not mask off) vector STA micro-instruction occupies one VSRB item. The Issue pipeline stage maintains the information of the number of currently available VSRB items and allocates one VSRB item for each vector STA micro-instruction. A non-mask off vector STA micro-instruction releases its VSRB item when retiring; while a mask off vector STA micro-instruction releases its VSRB item when execution is completed.
[0054] The VSRB entry contains the following fields:
[0055] Vector Store Retirement Index field IDX: The VSRB index, used for sequential allocation and release of VSRB entries;
[0056] Vector store instruction index field IID corresponding to the reorder buffer: The instruction index, corresponding to the IID field in the ROB, used to identify the ROB entry / instruction to which this VSRB entry belongs;
[0057] Store operation index field SOID in the vector store instruction: (store operation ID) The index of this store operation in a vector store instruction;
[0058] Index field SBID of the hardware resource item corresponding to the store operation: (store buffer ID) The index of the SB item corresponding to this store operation;
[0059] Store operation exception record identification field Exception: The initial value is -1. If an exception occurs during address translation for this store operation, the exception type (value >= 0) is written into this field;
[0060] Mask off store operation quantity field NOP_CNT before the store operation: The number of mask off store operations before this store operation; These mask off store operations may come from different vector store instructions.
[0061] In an embodiment of the present application, assisting the reorder buffer to retire the hardware resource items in granularity of the vector store retirement items in the preset vector store retirement table includes: sequentially releasing the hardware resource items starting from the item pointed to by the preset vector store pointer head every clock cycle; sequentially allocating the hardware resource items starting from the item pointed to by the preset vector store pointer tail every clock cycle; wherein, the vector store pointer head points to the longest valid item in the vector store retirement table, and the vector store pointer tail points to the first invalid item in the vector store retirement table.
[0062] In the above embodiment, the present application releases and allocates SB items through two newly added pointers. Among them, the vector store pointer head VSRB head: points to the oldest valid item in the VSRB, and attempts to sequentially release SB items starting from the item pointed to by the VSRB head every clock cycle. The vector store pointer tail VSRB tail: points to the first invalid item in the VSRB, and sequentially allocates new SB items starting from the item pointed to by the VSRB tail every clock cycle.
[0063] Specifically, reference can be made toFigure 2 As shown Figure 2 is ROB+VSRB with vector instruction support added. Two iterations of a loop code similar to memcpy are shown in the ROB. In the first iteration, the instructions have IID = 0 to 6, and in the second iteration, the instructions have IID = 7 to 13. The code in the example performs a unit-stride load (IID = 1 / 8) to read data from address (a1), and a strided store (IID = 4 / 11) to selectively write the data to address (a3 + i * a4) according to the mask value. The mask values (v0.t and v0'.t) used in the two iterations are also shown in the figure.
[0064] Assume VLEN = 128 and the instructions with IID = 0 / 7 set SEW = 32 and LMUL = 1. Therefore, the instructions at the 5th and 12th (IID = 4 and 11) positions in the ROB are vector store instructions, each having 4 (= 128 / 32) store operations. So, 4 SB entries are allocated for each instruction, and both STA_CNTDOWN and STD_CNTDOWN in the SB entry are set to 4. However, as Figure 2 shown by the mask values, since each instruction has only two valid (i.e., not masked off) vector store operations, only two VSRB entries are allocated for each store instruction, and the NOP_CNT in each entry records the number of masked-off store operations before this store operation.
[0065] In an embodiment of the present application, the vector store retirement table sequentially allocates an idle vector store retirement entry for the vector STA micro-instruction, including: when creating or writing a vector store retirement entry, the number of unmasked store operation items between the corresponding valid store operation and the adjacent previous valid store operation is written into the store operation pre-masked store operation quantity field of the vector store retirement entry; when releasing or reading a vector store retirement entry, and it is determined that there is valid data in the adjacent next entry through the store operation pre-masked store operation quantity field, read the store operation pre-masked store operation quantity field of the adjacent next entry, so as to retire all the hardware resource items of the unmasked store operations of the vector store instruction after the vector store retirement entry when retiring the hardware resource item corresponding to the vector store retirement entry.
[0066] Please refer to Figure 3As shown, when a vector instruction becomes the oldest unretired instruction in the ROB (i.e., the ROB head points to its ROB entry), if the bit 1 of its use_VSRB field is set, it has one or more corresponding VSRB entries. At this time, the VSRB head pointer should point to the first VSRB entry of this vector instruction. Starting from this VSRB head entry, all consecutive VSRB entries with Exception equal to -1 of this instruction can be retired, and the SB entries corresponding to the NOP_CNT fields in these entries and their next entries will also be retired. After all VSRB entries of this instruction are retired, the VSRB head will move to the first VSRB entry of the next vector instruction.
[0067] If the oldest instruction in the ROB is a vector store instruction, but its STA / STD microinstructions have not all been executed, this instruction will wait for all these STA / STD microinstructions to be executed (i.e., the values of STA_CNTDOWN and STD_CNTDOWN of this ROB entry are both 0) before it can retire. It should be noted that the allocation and retirement of VSRB entries are only related to STA microinstructions and not related to STD microinstructions. However, when a retired SB entry writes to memory, its corresponding data field must be valid, that is, the corresponding STD microinstruction has been executed.
[0068] An object of the present application is to be able to retire the SB entries corresponding to the mask off store operations as soon as possible. For this purpose, in the above embodiments, the allocation of VSRB entries is carried out sequentially. When creating / writing a VSRB entry, the number of mask off store operation entries between the corresponding valid store operation and the previous valid store operation is written into the NOP_CNT field of the newly allocated VSRB entry. However, when retiring / reading a VSRB entry, if its next entry has valid data, the NOP_CNT of its next entry is read to retire / release more SB entries as soon as possible, as Figure 3 shown. A vector store instruction allocates at least one VSRB entry. If all store operations of this instruction are mask off, its last store operation is allocated one VSRB entry, and its NOP_CNT is also set to the number of mask off store operations before this entry.
[0069] In an embodiment of the present application, a preset vector storage retirement table is used to assist in reordering the cache to retire the hardware resource items in units of vector storage retirement items in the vector storage retirement table, including: when all vector storage retirement items corresponding to the vector STA microinstructions of the vector store instruction are released, the vector storage retirement table provides a notification signal to the reordering cache; when all vector STD microinstructions of the vector store instruction are executed, the reordering cache releases the reordering cache item corresponding to the vector store instruction according to the notification signal. In another embodiment, the vector storage retirement table sequentially allocates an idle vector storage retirement item for the vector STA microinstruction, including: when the vector storage retirement item flags an exception, the corresponding fields in the store operation index field and the exception type in the store operation exception record identification field in the vector store instruction of the vector storage retirement item are provided to the reordering cache for exception handling.
[0070] In the above embodiment, when all VSRB items of a vector store instruction are released, the VSRB notifies the ROB; if all STD microinstructions of this instruction are also executed, the ROB item corresponding to this instruction is released. If a certain VSRB item flags an exception, the corresponding fields (in the SOID field) and the exception type (in the Exception field) of this VSRB item will be passed to the ROB for exception handling.
[0071] Based on the above embodiment, for the RISC-V vector store instruction, the RISC-V vector store hardware resource release method provided by the present application optimizes the store buffer item allocation and retire scheme, allocates and retires in units of effective SB items, and can release the SB items occupied by the mask off store operation as soon as possible, thereby improving the SB utilization efficiency and achieving the CPU performance optimization with fewer SB items.
[0072] To more clearly illustrate the RISC-V vector store hardware resource release method provided by the present application, the following takes multiple pipeline stages as an example to further illustrate the application of the RISC-V vector store hardware resource release method to the above embodiment. Those skilled in the relevant art should know that this example is only an application method to help understand the RISC-V vector store hardware resource release method provided by the present application and does not make any limitations on it.
[0073] Please refer to Figure 4 As shown, the existing CPU architecture generally includes multiple pipeline stages. The RISC-V vector store hardware resource release method provided by the present application for this pipeline stage includes the following improvements:
[0074] Instruction Decoding: A vector store instruction is decoded into at least a pair of STA / STD micro-instructions, with each pair corresponding to a store operation and an SB entry. The number of SB entries required for a vector store instruction depends on parameters such as VLEN, the current LMUL, and SEW, and for segment store instructions, the NFIELD parameter also needs to be considered. And set the IS_LAST bit for the generated vector STA uop: for the last uop, set IS_LAST to 1, and for the rest of the uops, set IS_LAST to 0.
[0075] Dispatch: If a micro-instruction is a vector STA micro-instruction and there are free entries in the SB, allocate the required SB entries for this micro-instruction. If this micro-instruction is the first micro-instruction of its corresponding instruction and there are free entries in the ROB, allocate a ROB entry for this instruction. If both of the above conditions are met, this micro-instruction enters the next pipeline stage. Otherwise, this micro-instruction congests the Dispatch and all front-end pipeline stages and tries to obtain SB and ROB entries again in the next clock cycle. A STD does not need to allocate SB entries; it only needs to use the SB entries allocated by its corresponding STA micro-instruction.
[0076] Issue: In this stage, find the oldest ready uop to execute and save the current number of free VSRBs. For a vector STA uop, it is only ready and allocated a VSRB entry when the values of its source physical registers already exist and there are free VSRB entries.
[0077] Execution: Vector STA micro-instructions must be executed in program order. There is a register ST_NOP_CNT in the STA execution unit to save the number of mask off store operations that have not been written to the VSRB. A vector store instruction has at least one VSRB entry. If all its store operations are mask off, its last STA micro-instruction will use its VSRB entry. Other mask off vector STA micro-instructions will release their VSTA entries (i.e., update the number of free VSRB entries in the Issue stage).
[0078] For details, please refer to Figure 5 As shown, the main process of this execution stage is as follows:
[0079] 1. First, determine whether a vector STA uop is mask off. If so, proceed to the next step; otherwise, jump to step 3.
[0080] 2. If the IS_LAST of this STA uop is 1 and the use_VSRB field of the corresponding ROB entry is 0, that is, all uops of this vector store instruction are masked off, then jump to step 6; otherwise, increment ST_NOP_CNT by 1, release the VSRB entry, and end.
[0081] 3. Write the address to the SB entry allocated for this uop in the Dispatch stage.
[0082] 4. If the VSRB is empty (VSRB head == VSRB tail), then execute the next step; otherwise, jump to step 6.
[0083] 5. If ST_NOP_CNT is not equal to 0, notify the SB to retire consecutive ST_NOP_CNT entries, and then set the ST_NOP_CNT entry to 0.
[0084] 6. Write the value of ST_NOP_CNT to the NOP_CNT field of the VSRB entry, then set the ST_NOP_CNT entry to 0. And set the use_VSRB field of the corresponding ROB entry to 1.
[0085] 7. End the execution.
[0086] Retire: For the ROB head entry, follow the following process to release this ROB entry, release the corresponding VSRB entry, and mark multiple possible SB entries as retire.
[0087] Specifically, refer to Figure 6 As shown, the main process of this retirement stage is as follows:
[0088] 1. First, check the Use_VSRB value in the ROB head entry. If it is 1, that is, this entry has at least one VSRB entry, jump to step 4; otherwise, execute the next step.
[0089] 2. Check if the instruction without a VSRB entry is abnormal. If it is abnormal, trigger the exception handling and end; otherwise, execute the next step.
[0090] 3. Determine whether the instruction has been executed. If it has been executed, jump to step 11; otherwise, this instruction has not been executed and cannot be retired, and end.
[0091] 4. Read the VSRB head entry.
[0092] 5. If this entry is valid (SBID is not -1) and has the same IID as the ROB entry, execute the next step; otherwise, jump to step 11.
[0093] 6. If this Exception is not zero, it indicates that the corresponding vector uop has an execution exception, and proceed to the next step; otherwise, jump to step 8.
[0094] 7. Send the exception index to the ROB. After receiving it, the ROB saves the exception index in the vstart CSR, and then triggers exception handling, and ends.
[0095] 8. Release the VSRB head item, set the SBID of this item to -1, and notify the Issue stage of the release of the VSRB item; then, move the VSRB head pointer to the next item ((VSRB_head + 1) % VSRB_size).
[0096] 9. If the new VSRB head item is valid (SBID is not equal to -1), then set CNT = 1 + the NOP_CNT value of this item; otherwise, set CNT to 1.
[0097] 10. Release the consecutive CNT items in the Store buffer, and then jump to step 4.
[0098] 11. Determine whether both the STA and STD microinstructions of the ROB head item have been executed (STA_CNTDOWN is 0 and STD_CNTDOWN is 0). If so, proceed to the next step; otherwise, this instruction has not been completely executed and cannot be Retired.
[0099] 12. At this time, the instruction corresponding to the ROB head item has been completely executed and there is no exception, and it can be retired. Then release all other hardware resources corresponding to this instruction.
[0100] 13. Release this ROB head item, that is, move the ROB head pointer to the next position ((ROB_head + 1) % ROB_size), and end.
[0101] Please refer to Figure 7As shown in the figure, the embodiments of this specification also provide a RISC-V vector store hardware resource release device, which includes an allocation module 701 and a release module 702. The allocation module 701 is used to respond to a vector store instruction, decode and generate multiple vector STA and STD micro-instruction pairs according to the number of store operations, and allocate a hardware resource item to each vector STA and STD micro-instruction pair. The release module 702 is used to assist the reorder buffer through a preset vector store retirement table to retire the hardware resource items in units of vector store retirement items in the vector store retirement table, so as to release the hardware resources corresponding to the hardware resource items. Among them, through the number of consecutive unmasked store operation items stored in the vector store retirement item, when retiring the hardware resource item corresponding to the vector store retirement item, all the hardware resource items of the unmasked store operations after the vector store retirement item are retired. Since the principle of this device to solve problems is similar to that of the RISC-V vector store hardware resource release method, the implementation of this device can refer to the implementation of the RISC-V vector store hardware resource release method, and the repeated parts will not be elaborated.
[0102] The embodiments of this specification also provide a central processing device, which includes a central processing core with a reorder buffer based on instructions, micro-instructions and macro-instructions, and the central processing device applies the above-mentioned RISC-V vector store hardware resource release method.
[0103] In the above embodiments, in many CPU cores, the ROB does not store instructions, but stores micro-instructions (microuop) or macro-instructions (macro op), which are all specific conversion formats of instructions in a CPU core: one instruction may be converted into one or more micro-instructions / macro-instructions, or two instructions may be combined into one micro-instruction / macro-instruction, etc. The ROB based on instructions / micro-instructions / macro-instructions is designed for scalar instructions and faces the same problems for vector instructions. Therefore, this application is also applicable to CPU cores using a ROB based on micro-instructions / macro-instructions. Similarly, for high-performance, superscalar CPU cores that can rename, execute, and retire multiple instructions in one clock cycle, this application is also applicable. It should be noted that in a high-performance, out-of-order execution CPU core, the two operations STA and STD included in a store operation are usually decomposed into two independent micro-instructions for execution. This application is also applicable to other methods besides this method, such as the same store micro-instruction executing in two pipelines for address generation and register reading.
[0104] For a RISC-V CPU core implementing the RISC-V vector extension, the RISC-V vector store hardware resource release method and apparatus provided by the embodiments of this specification provide an optimized solution for extending the ROB and SB to support RISC-V vector store instructions. Compared with the traditional solution, this application does not require a substantial increase in the hardware resources required for SB items, and achieves the purpose of releasing SB items as soon as possible by retiring SB items as soon as possible. At the same time, this application can also release the SB items corresponding to the mask off store operation as soon as possible, thereby further improving the utilization efficiency of SB.
[0105] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of this specification can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules respectively, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of this specification are not limited to any specific combination of hardware and software.
[0106] It should be understood that the above description is for illustrative purposes and not for limitation. By reading the above description, many embodiments and many applications other than the provided examples will be obvious to those skilled in the art. Therefore, the scope of this specification should not be determined with reference to the above description, but should be determined with reference to the full scope of the foregoing claims and the equivalents of these claims.
[0107] The above is only the preferred embodiment of this specification and is not used to limit this specification. For those skilled in the art, the embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.
Claims
1. A RISC-V vector store hardware resource release method, characterized in that: The method comprises: In response to the vector store instruction, a plurality of vector STA and STD microinstruction pairs are generated by decoding according to the number of store operations, and a hardware resource item is allocated to each vector STA and STD microinstruction pair; Retiring the hardware resource item by using a preset vector storage retirement table to assist the reorder cache with a vector storage retirement item in the vector storage retirement table as a granularity, so as to release the hardware resource corresponding to the hardware resource item; Among them, the number of continuous unmasked store operation items stored in the vector storage retirement item is used to retire the hardware resource items of all unmasked store operations of the vector store instructions after the vector storage retirement item when retiring the hardware resource item corresponding to the vector storage retirement item.
2. The RISC-V vector store hardware resource release method according to claim 1, characterized in that: Decode and generate multiple STA and STD microinstruction pairs according to the number of store operations, including: An IS_LAST field is added to the vector STA microinstruction generated by the vector store instruction, and the last bit information of the vector STA microinstruction in the vector store instruction is defined by the value of the IS_LAST field.
3. The RISC-V vector store hardware resource release method according to claim 1, characterized in that: The method further comprises: Set the vector store retirement entry identification field, the vector STA microinstruction number field and the vector STD microinstruction number field in the reorder cache entry; Among them, the vector storage retirement item identification field is used to configure different numerical values according to the vector store instruction or the non-vector store instruction; the vector STA microinstruction number field and the vector STD microinstruction number field are respectively used to record the number of vector STA microinstructions and the number of vector STD microinstructions generated by the vector store instruction, and adjust the corresponding recorded number of vector STA microinstructions and the number of vector STD microinstructions according to the execution status of the vector STA microinstructions and the vector STD microinstructions.
4. The RISC-V vector store hardware resource release method according to claim 1, characterized in that: The method further comprises: When the source physical register value of the vector STA microinstruction exists and there is an idle vector storage retirement item, the vector storage retirement table sequentially allocates an idle vector storage retirement item to the vector STA microinstruction; When retiring a vector STA microinstruction with no mask store operation, releasing the corresponding vector store retirement item; When the vector STA microinstruction executing the mask store operation is completed, the corresponding vector store retirement item is released; The vector storage retirement item includes a vector storage retirement index field, a vector store instruction index field corresponding to the reorder cache, a store operation index field in the vector store instruction, an index field of a hardware resource item corresponding to the store operation, a store operation exception record identification field, and a mask store operation number field before the store operation.
5. The RISC-V vector store hardware resource release method according to claim 4, characterized in that: The vector storage retirement table sequentially allocates an idle vector storage retirement entry for the vector STA microinstruction, including: When creating or writing a vector store retirement item, the number of unmasked store operations between the corresponding valid store operation and the previous valid store operation is written into the number of masked store operations before the store operation field of the vector store retirement item; When a vector storage retirement item is released or read, and it is determined through the store operation before masking the store operation number field that the adjacent next item has valid data, the store operation before masking the store operation number field of the adjacent next item is read, so that when retiring the hardware resource item corresponding to the vector storage retirement item, all the hardware resource items of the unmasked store operations of the vector store instruction after the vector storage retirement item are retired.
6. The RISC-V vector store hardware resource release method according to claim 4, characterized in that: The vector storage retirement table sequentially allocates an idle vector storage retirement entry for the vector STA microinstruction, including: When the vector store retirement item flag is abnormal, the corresponding field in the store operation index field in the vector store instruction of the vector store retirement item and the abnormal type in the store operation abnormal record identification field are provided to the reorder cache for abnormal processing.
7. The RISC-V vector store hardware resource release method according to claim 1, characterized in that: Assisting the reorder cache by presetting a vector storage retirement table to retire the hardware resource item at the granularity of the vector storage retirement item in the vector storage retirement table includes: Each clock cycle starts to release hardware resource items sequentially by presetting the items pointed to by the vector storage pointer head; Each clock cycle starts with sequentially allocating hardware resource items by presetting the item pointed to by the tail of the vector storage pointer; The vector storage pointer head points to the longest valid entry in the vector storage retirement table, and the vector storage pointer tail points to the first invalid entry in the vector storage retirement table.
8. The RISC-V vector store hardware resource release method according to claim 1, characterized in that: Assisting the reorder cache by presetting a vector storage retirement table to retire the hardware resource item at the granularity of the vector storage retirement item in the vector storage retirement table includes: When all vector store retirement items corresponding to the vector STA microinstructions of the vector store instruction are released, the vector store retirement table provides a notification signal to the reorder cache; When all vector STD microinstructions of the vector store instruction are executed, the reorder cache releases the reorder cache entry corresponding to the vector store instruction according to the notification signal.
9. A RISC-V vector store hardware resource release device, characterized in that: The device comprises: an allocation module and a release module; The allocation module is used for generating a plurality of vector STA and STD microinstruction pairs according to the number of store operations in response to the vector store instruction, and allocating a hardware resource item to each vector STA and STD microinstruction pair; The release module is used to assist the reorder cache to retire the hardware resource item with the vector storage retirement item in the vector storage retirement table as the granularity through the preset vector storage retirement table, so as to release the hardware resource corresponding to the hardware resource item; Among them, the number of continuous unmasked store operation items stored in the vector storage retirement item is used to retire the hardware resource items of all unmasked store operations of the vector store instructions after the vector storage retirement item when retiring the hardware resource item corresponding to the vector storage retirement item.
10. A central processing unit, comprising a central processing core based on a reorder cache of instructions, microinstructions and macroinstructions, characterized in that: The central processing unit applies the RISC-V vector store hardware resource release method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Hardware resource sharing method for multiprocessor system
CN114327920A
Physical register file distribution device for RISC-V vector and floating point register
CN115437691A