Register renaming method and system of vector processor and storage medium

Through the register renaming system of the vector processor, vector instructions and scalar instructions are renamed simultaneously, which solves the problem of insufficient transmission bandwidth caused by blocking in the existing technology and improves the performance and frequency requirements of the processor.

CN120631451AActive Publication Date: 2025-09-12RIVAI TECH (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511141829.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-12
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing vector processors have a blocking problem during the renaming process, resulting in insufficient transmission bandwidth and unable to meet high performance and frequency requirements.

Method used

The instruction fetch unit, decoding unit, renaming unit, reorder buffer unit and vector history stack unit are used to realize the simultaneous renaming of vector instructions and scalar instructions. The renaming information is recorded at the microinstruction granularity through the vector history stack unit to optimize the table entry usage of the reorder buffer unit.

Benefits of technology

It effectively reduces the blocking during the renaming process, improves the processor's transmission bandwidth and performance, ensures that vector instructions are renamed within one clock cycle, and avoids instruction reception problems caused by insufficient reordering buffer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631451A_ABST
    Figure CN120631451A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of processors, and particularly relates to a register renaming method and system of a vector processor and a storage medium. An instruction fetching unit of the register renaming system of the vector processor is used for storing an instruction from a memory and sending the instruction to a decoding unit; the decoding unit is used for decoding the received instruction; the renaming unit is used for renaming destination registers corresponding to the decoded scalar instruction and vector instruction as physical registers; the reordering buffer unit is used for storing a vector instruction, a scalar instruction and renaming information of the scalar instruction; the vector history heap unit is used for recording renaming information of vector instructions with micro instruction granularity. Compared with the prior art, the method has the advantages that the vector instruction and the scalar instruction can be renamed at the same time, blocking is effectively reduced, the transmitting bandwidth is optimized, and then the performance of a processor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applicable to the field of processor technology, and in particular relates to a register renaming method, system and storage medium for a vector processor. Background Art

[0002] With the advancement of reduced instruction set computing (RISC) technology, modern processors are increasingly adopting superscalar architectures, which execute multiple instructions within a single clock cycle to meet high computational demands and complex application scenarios. Register renaming is a key technology in superscalar processors. It typically eliminates false data dependencies by renaming logical registers to physical registers, thereby improving instruction-level parallelism (ILP). To improve data-level parallelism (DLP), modern processors have further evolved vector architectures, allowing a single instruction to process multiple copies of data. To improve the execution efficiency of vector instructions, processors also need to perform register renaming for vector instructions. Unlike scalar instructions, a vector instruction writes its execution result to multiple destination registers, necessitating the simultaneous renaming of multiple registers. Therefore, processors require sufficient renaming bandwidth to support both superscalar and vector instructions.

[0003] There are many renaming schemes for vector instructions in existing vector processors, such as Figure 1 As shown, Figure 1 This is a schematic diagram of the renaming process of vector instructions in related technologies. During the decoding stage, vector instructions will be identified from multiple instruction decoding channels and handed over to the vector decoding unit for processing one by one. The vector decoding unit will split a vector instruction into several micro-instructions (micro-ops) based on the granularity of vector registers, and then the renaming unit will rename these micro-instructions according to its processing bandwidth. The reorder buffer is a hardware structure in an out-of-order processor that ensures that instructions are submitted in order. It will save the renaming information of each instruction being executed until the instruction is finally submitted. In a superscalar-vector processor, the reorder buffer will allocate a table entry for each scalar instruction or vector micro-instruction to store its information.

[0004] Because instructions must be renamed sequentially, the aforementioned superscalar-vector register renaming method presents a blocking issue. First, due to the limited bandwidth of the reorder buffer, a vector instruction cannot be renamed in the same clock cycle as several other scalar instructions. A vector instruction cannot begin renaming until all preceding scalar instructions have completed renaming. Second, the vector decode unit may require multiple cycles to split and rename a vector instruction, during which time subsequent scalar instructions cannot begin renaming. Third, the microinstructions split into vector instructions can occupy too much space in the reorder buffer, and insufficient remaining space in the reorder buffer can prevent subsequent instructions from being received. In scenarios where mixed scalar and vector instructions are issued, this register renaming method can result in insufficient transmit bandwidth. This insufficient transmit bandwidth cannot be compensated in the subsequent execution phase, ultimately resulting in performance loss.

[0005] Therefore, there is an urgent need for a new register renaming method, system and storage medium for a vector processor to solve the above technical problems. Summary of the Invention

[0006] The present invention provides a register renaming method, system and storage medium for a vector processor, aiming to enable register renaming of scalar instructions and vector instructions simultaneously, so that the processor meets higher performance and frequency requirements.

[0007] In a first aspect, the present invention provides a register renaming system for a vector processor, comprising an instruction fetch unit, a decoding unit, a renaming unit, a reordering buffer unit, and a vector history stack unit; The instruction fetch unit is used to read instructions from the memory and send the instructions to the decoding unit; The decoding unit is used to decode the received instruction and send the scalar instruction and vector instruction obtained by decoding to the renaming unit and the reorder buffer unit; The renaming unit is configured to rename the destination registers corresponding to the decoded scalar instruction and the vector instruction into physical registers, and send the renaming information corresponding to the scalar instruction to the reorder buffer unit and send the renaming information corresponding to the vector instruction to the vector history stack unit; The reorder buffer unit is used to store the vector instructions, the scalar instructions and renaming information of the scalar instructions; The vector history stack unit is used to record the renaming information of the vector instructions at a microinstruction granularity.

[0008] Preferably, a single vector instruction occupies only a single entry in the reorder buffer unit.

[0009] Preferably, when the renamed vector instruction cannot be submitted, the vector history stack unit is further configured to modify the physical register mapped by the renamed vector instruction back to the physical register mapped before the renaming according to the renaming information.

[0010] In a second aspect, the present invention further provides a register renaming method for a vector processor. The register renaming method is based on the register renaming system for the vector processor as described in the above embodiment, and the register renaming method comprises the following steps: S1, read the instruction from the memory through the instruction fetch unit, and send the instruction to the decoding unit; S2. Decoding the received instruction by the decoding unit, and sending the decoded scalar instruction and vector instruction to the renaming unit; S3, renaming the destination registers corresponding to the decoded scalar instruction and the vector instruction into physical registers through the renaming unit, sending the renaming information corresponding to the scalar instruction to the reorder buffer unit, and sending the renaming information corresponding to the vector instruction to the vector history stack unit and the reorder buffer unit; S4. When the scalar instruction is successfully submitted, the entry corresponding to the scalar instruction in the reorder buffer unit is deleted; when the vector instruction is successfully submitted, the entry corresponding to the vector instruction in the reorder buffer unit is deleted, and the entry corresponding to the vector instruction in the vector history stack unit is deleted. Preferably, a single vector instruction occupies only a single entry in the reorder buffer unit.

[0011] Preferably, when the vector instruction cannot be submitted, the physical register mapped by the renamed vector instruction is modified back to the physical register mapped before the renaming according to the renaming information.

[0012] In a third aspect, the present invention further provides a computer device comprising: a memory, a processor, and a register renaming program for a vector processor stored in the memory and executable on the processor, wherein the processor implements the steps of the register renaming method for a vector processor as described in any one of the above embodiments when executing the register renaming program for the vector processor.

[0013] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a register renaming program for a vector processor is stored. When the register renaming program for the vector processor is executed by a processor, the steps in the register renaming method for the vector processor as described in any one of the above embodiments are implemented.

[0014] Compared with the prior art, the present invention enables vector instructions and scalar instructions to be renamed simultaneously through an instruction fetch unit, a decoding unit, a renaming unit, a reorder buffer unit, and a vector history stack unit, effectively reducing congestion, optimizing the transmission bandwidth, and thus improving processor performance. Since the renaming information of the vector instruction is recorded at the microinstruction granularity by the vector history stack unit, a vector instruction can be completely renamed within one clock cycle without consuming multiple clock cycles to split each vector instruction into microinstructions during the renaming period. At the same time, a single vector instruction only occupies a single table entry in the reorder buffer unit, thereby avoiding the problem that the vector instruction is split into several microinstructions that occupy too much reorder buffer, and the insufficient remaining space in the reorder buffer will cause subsequent instructions to be unable to be received. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The present invention will be described in detail below with reference to the accompanying drawings. The above and other aspects of the present invention will become clearer and easier to understand through the detailed description made with reference to the following drawings. In the accompanying drawings: Figure 1 It is a schematic diagram of the renaming process of vector instructions in the related art; Figure 2 1 is a schematic structural diagram of a register renaming system for a vector processor provided by an embodiment of the present invention; Figure 3 1 is a flow chart of a rename vector instruction of a register renaming system of a vector processor provided by an embodiment of the present invention; Figure 4 This is a flowchart of a register renaming method for a vector processor provided by an embodiment of the present invention; Figure 5 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] Example 1 The embodiment of the present invention provides a register renaming system 100 for a vector processor, please refer to Figure 2 , Figure 2 1 is a schematic structural diagram of a register renaming system 100 for a vector processor provided by an embodiment of the present invention, which includes an instruction fetch unit 1, a decoding unit 2, a renaming unit 3, a reordering buffer unit 4, and a vector history stack unit 5; The instruction fetch unit 1 is used to read instructions from a memory and send the instructions to the decoding unit 2; The decoding unit 2 is used to decode the received instructions and send the decoded scalar instructions and vector instructions to the renaming unit 3 and the reordering buffer unit 4; specifically, when decoding the vector instructions, the decoding unit 2 does not split the vector instructions into microinstructions.

[0018] The renaming unit 3 is used to rename the destination registers corresponding to the decoded scalar instructions and vector instructions into physical registers, and send the renaming information corresponding to the scalar instructions to the reorder buffer unit 4 and send the renaming information corresponding to the vector instructions to the vector history stack unit 5. Specifically, the renaming unit 3 can perform renaming operations on one vector instruction and multiple scalar instructions at the same time. The reorder buffer unit 4 is used to store the vector instructions, the scalar instructions and the renaming information of the scalar instructions; The vector history stack unit 5 is used to record the renaming information of the vector instruction at the microinstruction granularity. Specifically, since the renaming information storage units of vector instructions and scalar instructions are different, the renaming processing of scalar instructions and vector instructions is decoupled to a certain extent, avoiding the blocking problem that may be caused by the existing solution and improving the overall performance. Microinstruction granularity means that the vector instruction is split into several microinstructions with the vector register as the granularity and stored in the vector history stack unit 5. For example, the destination register of the vector instruction vadd includes two consecutive registers starting with vector register No. 4 (v4), such as v4 and v5, then the renaming information corresponding to v4 and v5 will be stored in two table entries of the vector history stack unit 5 respectively.

[0019] In the embodiment of the present invention, a single vector instruction occupies only a single entry in the reorder buffer unit 4. A vector instruction updates multiple entries in the vector register renaming table and the vector history stack unit 5, but only occupies one entry in the reorder buffer unit 4. This allows the reorder buffer unit 4 to simultaneously receive several other scalar instructions, thereby improving the overall instruction issuance bandwidth.

[0020] In an embodiment of the present invention, when the renamed vector instruction cannot be submitted, the vector history stack unit 5 is further configured to modify the physical register mapped by the renamed vector instruction back to the physical register mapped before renaming according to the renaming information.

[0021] For details, please refer to Figure 3 , Figure 3This is a flow chart illustrating the renaming of vector instructions by the register renaming system 100 of the vector processor provided by an embodiment of the present invention. When a vector instruction is renamed by the register renaming system 100 of the vector processor, the process is completed within a single clock cycle. Taking a vector addition (vadd) instruction as an example, it is assumed here that the calculation result is written to two destination registers. It should be noted that other numbers of destination registers are also possible, and this is for illustrative purposes only. This vector instruction writes the calculation result to two consecutive registers, namely, v4 and v5, starting with vector register 4 (v4). The processor's free register list allocates a physical register for each destination register of the vector instruction, which is used to store the actual calculation result of the vector instruction during subsequent execution. Simultaneously, the mappings of v4 and v5 in the register renaming table are modified by the renaming unit 3 to the newly allocated physical registers, namely, physical registers P8 and P11.

[0022] The vector history stack unit 5 records the renaming information of the vector instructions (ie, the modification history of the vector register renaming table) for physical register release and state backtracking. Figure 3 As shown, the vector history stack unit 5 records the modified mapping of the destination registers v4 and v5 before the vector instruction is renamed, and records the physical registers previously mapped to them, namely physical registers P6 and P7. After the vector instruction is successfully executed and submitted, P6 and P7 will be released and returned to the free register list. If an abnormal condition (such as a branch misprediction) subsequently causes the vector instruction to be unable to be submitted, the renaming information recorded in the vector history stack unit 5 is restored to the register renaming table, returning them to their original state before being modified by the vector instruction (i.e., P6 and P7).

[0023] Compared with the prior art, the present invention enables vector instructions and scalar instructions to be renamed simultaneously through an instruction fetch unit, a decoding unit, a renaming unit, a reorder buffer unit, and a vector history stack unit, effectively reducing congestion, optimizing the transmission bandwidth, and thus improving processor performance. Since the renaming information of the vector instruction is recorded at the microinstruction granularity by the vector history stack unit, there is no need to split and rename a vector instruction for multiple cycles. At the same time, a single vector instruction only occupies a single table entry in the reorder buffer unit, thereby avoiding the problem that the vector instruction is split into several microinstructions that occupy too much reorder buffer, and the insufficient remaining space in the reorder buffer will cause subsequent instructions to be unable to be received.

[0024] Example 2 Please refer to Figure 4The present invention further provides a register renaming method for a vector processor. The register renaming method is based on the register renaming system 100 for a vector processor as described in the above embodiment. The register renaming method includes the following steps: S1, reading an instruction from a memory through the instruction fetch unit 1, and sending the instruction to the decoding unit 2; S2, decoding the received scalar instruction and the received vector instruction through the decoding unit 2, and sending the decoded scalar instruction and vector instruction to the renaming unit 3; S3, renaming the destination registers corresponding to the decoded scalar instruction and the vector instruction into physical registers through the renaming unit 3, and sending the renaming information corresponding to the scalar instruction to the reorder buffer unit 4, and sending the renaming information corresponding to the vector instruction to the vector history stack unit 5 and the reorder buffer unit 4; S4. When the scalar instruction is successfully submitted, the table entry corresponding to the scalar instruction in the reorder buffer unit 4 is deleted; when the vector instruction is successfully submitted, the table entry corresponding to the vector instruction in the reorder buffer unit 4 and the table entry corresponding to the vector instruction in the vector history stack unit 5 are deleted.

[0025] In the embodiment of the present invention, a single vector instruction occupies only a single entry in the reorder buffer unit 4. A vector instruction updates multiple entries in the vector register renaming table and the vector history stack unit 5, but only occupies one entry in the reorder buffer unit 4. This allows the reorder buffer unit 4 to simultaneously receive several other scalar instructions, thereby improving the overall instruction issuance bandwidth.

[0026] In an embodiment of the present invention, when the renamed vector instruction cannot be submitted, the vector history stack unit 5 is further configured to modify the physical register mapped by the renamed vector instruction back to the physical register mapped before renaming according to the renaming information.

[0027] The register renaming method for a vector processor is based on the modules in the register renaming system 100 for a vector processor in the above embodiment and can achieve the same technical effects. Please refer to the description in the above embodiment and will not be repeated here.

[0028] Example 3 The embodiment of the present invention also provides a computer device, please refer to Figure 5 , Figure 52 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. The computer device 200 includes: a memory 202, a processor 201, and a register renaming program of a vector processor stored in the memory 202 and executable on the processor 201.

[0029] The processor 201 calls the register renaming program of the vector processor stored in the memory 202 and executes the steps of the register renaming method of the vector processor provided by the embodiment of the present invention. Figure 4 , specifically including the following steps: S1, reading an instruction from a memory through the instruction fetch unit 1, and sending the instruction to the decoding unit 2; S2, decoding the received instruction through the decoding unit 2, and sending the scalar instruction and vector instruction obtained by decoding to the renaming unit 3; S3, renaming the destination registers corresponding to the decoded scalar instruction and the vector instruction into physical registers through the renaming unit 3, and sending the renaming information corresponding to the scalar instruction to the reorder buffer unit 4, and sending the renaming information corresponding to the vector instruction to the vector history stack unit 5 and the reorder buffer unit 4; S4. When the scalar instruction is successfully submitted, the table entry corresponding to the scalar instruction in the reorder buffer unit 4 is deleted; when the vector instruction is successfully submitted, the table entry corresponding to the vector instruction in the reorder buffer unit 4 and the table entry corresponding to the vector instruction in the vector history stack unit 5 are deleted.

[0030] In the embodiment of the present invention, a single vector instruction occupies only a single entry in the reorder buffer unit 4. A vector instruction updates multiple entries in the vector register renaming table and the vector history stack unit 5, but only occupies one entry in the reorder buffer unit 4. This allows the reorder buffer unit 4 to simultaneously receive several other scalar instructions, thereby improving the overall instruction issuance bandwidth.

[0031] In an embodiment of the present invention, when the renamed vector instruction cannot be submitted, the vector history stack unit 5 is further configured to modify the physical register mapped by the renamed vector instruction back to the physical register mapped before renaming according to the renaming information.

[0032] The computer device 200 provided in the embodiment of the present invention can implement the steps in the register renaming method of the vector processor in the above embodiment and can achieve the same technical effects. Please refer to the description in the above embodiment and will not be repeated here.

[0033] Example 4 An embodiment of the present invention further provides a computer-readable storage medium, on which a register renaming program for a vector processor is stored. When the register renaming program for a vector processor is executed by a processor, the register renaming program for the vector processor implements the various processes and steps in the register renaming method for a vector processor provided in an embodiment of the present invention, and can achieve the same technical effects. To avoid repetition, the details are not described here.

[0034] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0035] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0036] Through the above description of the embodiments, those skilled in the art will clearly understand that the methods of the above embodiments can be implemented using software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is the more preferred implementation method. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, or optical disk) and includes a number of instructions for enabling a terminal (such as a mobile phone, computer, server, air conditioner, or network device) to execute the methods described in the various embodiments of the present invention.

[0037] The embodiments of the present invention are described above in conjunction with the accompanying drawings. What is disclosed is only a preferred embodiment of the present invention. However, the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms and equivalent changes without departing from the scope of protection of the purpose of the present invention and the claims, which are all within the protection of the present invention.

Claims

1. A register renaming system for a vector processor, characterized in that: Includes instruction fetch unit, decoding unit, renaming unit, reordering buffer unit and vector history stack unit; The instruction fetch unit is used to read instructions from the memory and send the instructions to the decoding unit; The decoding unit is used to decode the received instruction and send the scalar instruction and vector instruction obtained by decoding to the renaming unit and the reorder buffer unit; The renaming unit is configured to rename the destination registers corresponding to the decoded scalar instruction and the vector instruction into physical registers, and send the renaming information corresponding to the scalar instruction to the reorder buffer unit and send the renaming information corresponding to the vector instruction to the vector history stack unit; The reorder buffer unit is used to store the vector instructions, the scalar instructions and renaming information of the scalar instructions; The vector history stack unit is used to record the renaming information of the vector instructions at a microinstruction granularity.

2. The register renaming system for a vector processor according to claim 1, wherein: A single vector instruction only occupies a single entry in the reorder buffer unit.

3. The register renaming system for a vector processor according to claim 1, wherein: When the renamed vector instruction cannot be submitted, the vector history stack unit is further configured to modify the physical register mapped by the renamed vector instruction back to the physical register mapped before the renaming according to the renaming information.

4. A register renaming method for a vector processor, characterized in that: The register renaming method is based on the register renaming system for the vector processor according to claim 1, and the register renaming method comprises the following steps: S1, read the instruction from the memory through the instruction fetch unit, and send the instruction to the decoding unit; S2. Decoding the received instruction by the decoding unit, and sending the decoded scalar instruction and vector instruction to the renaming unit; S3, renaming the destination registers corresponding to the decoded scalar instruction and the vector instruction into physical registers through the renaming unit, sending the renaming information corresponding to the scalar instruction to the reorder buffer unit, and sending the renaming information corresponding to the vector instruction to the vector history stack unit and the reorder buffer unit; S4. When the scalar instruction is successfully submitted, the table entry corresponding to the scalar instruction in the reorder buffer unit is deleted; when the vector instruction is successfully submitted, the table entry corresponding to the vector instruction in the reorder buffer unit and the table entry corresponding to the vector instruction in the vector history stack unit are deleted.

5. The register renaming method for a vector processor according to claim 4, wherein: A single vector instruction only occupies a single entry in the reorder buffer unit.

6. The register renaming method for a vector processor according to claim 4, wherein: When the vector instruction cannot be submitted, the physical register mapped by the renamed vector instruction is modified back to the physical register mapped before the renaming according to the renaming information.

7. A computer device, characterized in that: include: A memory, a processor, and a register renaming program for a vector processor stored in the memory and executable on the processor, wherein the processor implements the steps of the register renaming method for a vector processor according to any one of claims 4 to 6 when executing the register renaming program for the vector processor.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a register renaming program for a vector processor, and when the register renaming program for a vector processor is executed by a processor, the steps of the register renaming method for a vector processor according to any one of claims 4 to 6 are implemented.

Citation Information

Patent Citations

  • Implementation method of floating point physical register file

    CN112181494A

  • Register renaming method, device, system, equipment and medium

    CN119003004A

  • Asynchronous recording out-of-order processor register renaming check point and rollback method

    CN120123005A

  • Vector configuration instruction implementation method and system and storage medium

    CN120335869A

  • Register renaming for power conservation

    US20230305852A1