An Instruction Dispatching Method and Device Based on Instruction Dependencies

By dispatching the instructions with dependencies to the same reservation station according to the dependencies between instructions, and dispatching the instructions with dependencies to different execution units, the blocking problem during parallel execution of instructions is solved and the performance of the processor is improved.

CN114996017BActive Publication Date: 2025-07-29GUANGDONG STARFIVE TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210671746.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-07-29
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

In the prior art, the dependencies between instructions cannot be effectively processed, resulting in blocking when the instructions are executed in parallel, affecting processor performance.

Method used

According to the dependency relationship between instructions, instructions with dependencies are dispatched to the same reserved station, and instructions without dependencies are dispatched to different execution units. Through the multi-port queue structure of the reserved station and the early wake-up control signal of the execution unit, the parallel transmission and execution of instructions are achieved.

Benefits of technology

Improves the parallel execution efficiency of instructions, avoids blocking of dependency-free instructions, and improves processor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114996017B_ABST
    Figure CN114996017B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of microprocessors, and specifically relates to an instruction dispatch method and device based on instruction dependencies, including the following steps: S1 After instruction renaming, determine the dependency relationship between instructions according to the decoding information of the decoder; S2 Enter the instruction dispatch stage, and dispatch instructions with dependency relationships to the same reservation station; S3 After the source operands of the instructions are prepared, dispatch instructions without dependency relationships to different execution units for execution; S4 After the instructions are executed, broadcast the destination registers of the instructions and the instruction execution structure. The present invention can, according to the dependency relationship between instructions, dispatch instructions with dependency relationships to the same reservation station, so as not to block the parallel dispatch and execution of instructions without dependency relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microprocessors, and particularly relates to an instruction dispatch method and device based on instruction correlation. Background Art

[0002] In just a few decades, the development of microprocessors has made great progress. The performance of processors has been continuously improved in terms of hardware architecture, process, and the combination of software and hardware. The hardware architecture has evolved from single-issue scalar to multi-issue superscalar; from the initial 3-stage pipeline to dozens of stages; from sequential instruction execution to out-of-order instruction execution; from no cache to multi-level cache storage structures; from physical single-core to physical multi-core (CMP, Chip Multi-Processors) and logical single-core to logical multi-core (SMT, Simultaneous Multi-Threading); and even in the cluster systems for supercomputing, the instruction-level parallelism and thread-level parallel execution of processors have been developed to the extreme. The instruction-level parallel bandwidth requirement of single-core microprocessors is getting higher and higher, and the logical complexity of chip implementation increases exponentially. Summary of the Invention

[0003] Aiming at the deficiencies of the prior art, the present invention discloses an instruction dispatch method and device based on instruction correlation, which is used to dispatch instructions with dependency relationships to the same reservation station according to the dependency relationships between instructions, so as not to block the parallel issue and execution of instructions without dependency relationships.

[0004] The present invention is realized through the following technical solutions:

[0005] In a first aspect, the present invention provides an instruction dispatch method based on instruction correlation, including the following steps:

[0006] S1 After instruction renaming, judge the dependency relationships between instructions according to the decoding information of the decoder;

[0007] S2 Enter the instruction dispatch stage, and dispatch instructions with dependency relationships to the same reservation station;

[0008] S3 After the source operands of the instructions are ready, dispatch instructions without dependency relationships to different execution units for execution;

[0009] S4 After the instructions are executed, broadcast the destination registers of the instructions and the instruction execution structure.

[0010] Furthermore, in the method, after the false correlations between instructions are renamed, the instructions can be executed simultaneously on the same type of execution components.

[0011] Further, in the method, the reservation station is a multi-port queue structure. After an instruction enters the reservation station, the instruction monitors the Forward path of the execution unit in real time.

[0012] Further, in the method, after the source operands of the instruction in the reservation station are awakened to the ready state, the instruction is issued in advance. In the EX0 stage, the instruction can be Forwarded to the data returned by the execution unit.

[0013] Further, in the method, the execution unit generates an early wake-up control signal for waking up the instructions in the reservation station in advance according to the execution cycle of the instruction.

[0014] Further, in the method, when an instruction is dispatched to the execution unit, it first judges whether there is free space in the corresponding reservation station RV of the execution unit. If there is no free space in the reservation station RV, the current instruction cannot be dispatched to this reservation station.

[0015] Further, in the method, when an instruction is dispatched to the execution unit, it also judges whether there is free space in the reservation stations RV corresponding to other integer execution units at the same time. If there is no free space in the RVs corresponding to all integer execution units, the current instruction cannot be dispatched to the reservation station until there is free space in the reservation station, and then it can be dispatched to the reservation station.

[0016] Further, in the method, when an instruction is dispatched to the execution unit, if there is free space in the reservation stations of multiple integer execution units, the integer instruction can be dispatched to the reservation station RV of any integer execution unit.

[0017] In a second aspect, the present invention provides an instruction dispatch device based on instruction dependencies, on which instructions are stored. When the instructions are executed, the instruction dispatch method based on instruction dependencies described in the first aspect is implemented.

[0018] The beneficial effects of the present invention are as follows:

[0019] The present invention can dispatch instructions with dependency relationships to the same reservation station according to the dependency relationships between instructions, so as not to block the parallel issuance and execution of instructions without dependency relationships. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0021] Figure 1 It is the multi-core CPU architecture diagram of the embodiment of the present invention;

[0022] Figure 2 It is the CPU pipeline architecture diagram of the embodiment of the present invention;

[0023] Figure 3 It is the instruction dispatch diagram of the embodiment of the present invention;

[0024] Figure 4 It is the forward result diagram of the execution unit of the embodiment of the present invention;

[0025] Figure 5 It is the instruction dependency relationship diagram of the embodiment of the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Embodiment 1

[0028] This embodiment provides an instruction dispatch method based on instruction correlation, including the following steps:

[0029] S1 After instruction renaming, judge the dependency relationship between instructions according to the decoding information of the decoder;

[0030] S2 Enter the instruction dispatch stage, and dispatch the instructions with dependency relationship to the same reservation station;

[0031] S3 After the source operands of the instruction are ready, dispatch the instructions without dependency relationship to different execution units for execution;

[0032] S4 After the instruction execution is completed, broadcast the destination register of the instruction and the instruction execution structure.

[0033] This embodiment proposes to dispatch the instructions with dependency relationship to the same reservation station according to the correlation of instructions, so as not to block the parallel execution of the instructions without dependency relationship.

[0034] In this embodiment, when there are multiple execution components of the same type in the execution unit, the instructions with dependency relationship will not be dispatched in the same clock cycle, but the instructions without dependency relationship can be dispatched to different execution units in the same cycle.

[0035] When dispatching instructions without dependency relationships to the same reservation station and dispatching instructions with dependency relationships to different execution units, it is very likely that only one instruction can be dispatched in the current cycle.

[0036] When dispatching instructions with dependency relationships to the same reservation station and dispatching instructions without dependency relationships to different execution units, it is very likely that multiple instructions can be dispatched in the current cycle.

[0037] In this embodiment, after the instructions are renamed, according to the decoding information of the decoder, the dependency relationships between the instructions are judged, that is, there is data dependence between the instructions. After the false dependencies between the instructions are renamed, the instructions can be executed simultaneously on the same type of execution components. Only the instructions with true dependencies need to be dispatched according to the dependencies.

[0038] This embodiment enters the instruction dispatch stage. According to the instruction dispatch information obtained in the renaming stage, the instructions are dispatched to the reservation stations. After the instructions enter the reservation stations, if the source operands of the instructions are ready, the instructions are dispatched to the execution units for execution, as Figure 3 shown.

[0039] The reservation station in this embodiment is a multi-port queue structure. When the instructions enter the reservation station, the instructions monitor the Forward paths of the execution units in real time.

[0040] The execution unit in this embodiment generates an early wake-up control signal for waking up the instructions in the reservation station in advance according to the execution cycle of the instructions, as Figure 4 shown. When the source operands of the instructions in the reservation station are woken up to the ready state, the instructions do not need to wait until the execution unit finishes execution, but the instructions are dispatched in advance. In the EX0 stage, the instructions can Forward to the data returned by the execution unit.

[0041] When the instructions in the execution unit are executed, the destination registers and the instruction execution structures of the instructions are broadcast. All the instructions with dependencies in the reservation station can obtain the source operands. When each source operand of the instructions in the reservation station is ready, the instructions can enter the dispatch state, and the reservation station dispatches the instructions in the dispatch state to the execution units.

[0042] In this embodiment, in each clock cycle of the high-performance processor, as many instructions as possible are dispatched from the reservation station to the execution units, and the theoretical bandwidth is the number of execution units. According to different application scenarios, the instruction ratios in different scenarios are analyzed. According to the ratios occupied by the instructions, the number of execution units is determined. Especially for the same type of instructions, after the ratio exceeds a certain threshold, multiple identical types of execution units will be placed repeatedly.

[0043] In this embodiment, in order to improve the processing performance of instructions, multiple execution units of the same type are placed inside the CPU to execute the same type of instructions simultaneously.

[0044] For example, the ratio of integer instructions in the instruction stream exceeds 50% on average. Therefore, multiple integer execution units need to be repeatedly placed in the execution component. An integer instruction can be dispatched to any one of the integer execution units for execution. When an instruction is dispatched to an execution unit, it first checks whether there is free space in the reservation station RV corresponding to the execution unit. If there is no free space in the reservation station RV, then the current instruction cannot be dispatched to this reservation station. At the same time, it checks whether there is free space in the reservation stations RV corresponding to other integer execution units. If there is no free space in the RVs corresponding to all integer execution units, then the current instruction cannot be dispatched to the reservation station until there is free space in the reservation station, and then it can be dispatched to the reservation station.

[0045] On the contrary, when there is free space in the reservation stations of multiple integer execution units, then the integer instruction can be dispatched to the reservation station RV of any one of the integer execution units. The instruction dispatched to the reservation station may have a dependency relationship with an older instruction, that is, when the dependent instruction has not been executed yet, the dependent instruction cannot be executed either, but needs to obtain the execution result of the dependent instruction.

[0046] Embodiment 2

[0047] This embodiment provides a multi-core CPU in which N physical cores share L3 and memory. Referring to Figure 1 as shown, each physical core can be a single-threaded or multi-threaded architecture. Each core is applicable to all instruction sets, architectures, and processes.

[0048] This embodiment provides a single physical core. Referring to Figure 2 as shown, the physical core can be a single-threaded or multi-threaded architecture. The module division of this core is described in terms of functions in Table 1.

[0049] Table 1 CPU Pipeline Description

[0050]

[0051] Embodiment 3

[0052] At the specific implementation level, this embodiment provides a dependency relationship between instructions, that is, there is a true dependency between instructions. Let's assume that the window for instruction detection is 3 instructions. The instructions start from 0 as the first instruction, and so on for other cases. The instructions in the first detection window are 0, 1, and 2. The first instruction checks the dependency relationship with the instructions before the current clock cycle; the second instruction checks the dependency relationship with the first instruction; the third instruction checks the dependency relationship with the first and second instructions simultaneously;

[0053] This embodiment covers all the dependency relationships between source registers and destination registers of instructions, as well as the number of source registers and destination registers. To facilitate the explanation of the principle of the present invention: 1. Assume that each instruction has only 2 source registers rs1 and rs2. 2. Assume that each instruction has only one destination register.

[0054] In this embodiment, when the source operand of the second instruction depends on the destination register of the first instruction, then their dependency relationship is:

[0055] The first source register Dep_1_0_rs1 = (r_x1_1 == r_x3_0)

[0056] The second source register Dep_1_0_rs2 = (r_x2_1 == r_x3_0)

[0057] The second instruction depends on the first instruction Dep_1_0 = {Dep_1_0_rs2, Dep_1_0_rs1}.

[0058] In this embodiment, when the source operand of the third instruction depends on the destination register of the first instruction, then their dependency relationship is:

[0059] The first source register Dep_2_0_rs1 = (r_x1_2 == r_x3_0)

[0060] The second source register Dep_2_0_rs2 = (r_x2_2 == r_x3_0)

[0061] The third instruction depends on the first instruction Dep_2_0 = {Dep_2_0_rs2, Dep_2_0_rs1}.

[0062] In this embodiment, when the source operand of the third instruction depends on the destination register of the second instruction, then their dependency relationship is:

[0063] The first source register Dep_2_1_rs1 = (r_x1_2 == r_x3_1)

[0064] The second source register Dep_2_1_rs2 = (r_x2_2 == r_x3_1)

[0065] The third instruction depends on the second instruction Dep_2_1 = {Dep_2_1_rs2, Dep_2_1_rs1}.

[0066] In this embodiment, it is assumed that the 3 instructions in the first instruction window are all integer instructions, and they are dispatched to the reservation stations RV_0 and RV_1 corresponding to 2 identical execution units, and RV_0 and RV_1 can dispatch 2 instructions and issue 1 instruction each time.

[0067] This embodiment mainly proposes instruction scheduling according to the correlation of instructions, without limiting the specific dispatching and scheduling method. For the convenience of explaining the principle of the present invention, it is assumed that the first instruction is dispatched to RV_0, and the dispatching and scheduling examples of other instructions are shown in the following table:

[0068] Table 2 Instruction Dispatching and Scheduling

[0069]

[0070]

[0071]

[0072] In this embodiment, RV_V0 being 001 means that the first instruction is dispatched to RV_0; RV_V0 being 010 means that the second instruction is dispatched to RV_0; RV_V0 being 100 means that the third instruction is dispatched to RV_0; RV_V0 being 011 means that the first instruction and the second instruction are both dispatched to RV_0; RV_V0 being 101 means that the first instruction and the third instruction are both dispatched to RV_0; RV_V0 being 110 means that the second instruction and the third instruction are both dispatched to RV_0. Similarly, RV_V1 represents the instruction situation dispatched to RV_1. In the above text, assuming that the first instruction is dispatched to RV_0, whether the first instruction is dispatched to RV_0 or RV_1 mainly depends on whether the instruction on which the first instruction depends is in RV_0 or RV_1.

[0073] In the process of judgment in this embodiment, the dependencies of rs1 and rs2 of each instruction on the destination registers of other instructions are considered separately. The rs1 and rs2 of the same instruction may have dependencies on different instructions or the same instruction. Multiple instructions may have dependencies on the same instruction.

[0074] Embodiment 4

[0075] On other levels, this embodiment provides a method for instruction dispatching and scheduling according to the correlation of instructions, so that dependent instructions are dispatched to the same reservation station, independent instructions are dispatched to different reservation stations, and irrelevant instructions are parallelly issued to different execution units for execution in different reservation stations.

[0076] This embodiment further elaborates the principle and preferably selects an example of an actual RISC V instruction for supplementary explanation.

[0077] In this preferred embodiment, when the instruction sequence is as shown in the following table, assume that there are 2 integer instruction reservation stations, each reservation station can dispatch 2 instructions each time, and each reservation station can issue one instruction to the execution unit for execution per clock cycle. The instruction pipeline width is 4 instructions, and at most 4 instructions need to be dispatched to the reservation stations per clock cycle.

[0078] Table 3 RISC V Instruction Sequence Example

[0079]

[0080] In this preferred embodiment, for the 4 instructions in the first clock cycle, addiw, slli, addi, and add. The source operands of the fourth instruction add are t5 and t6. The second instruction slli writes to the destination register t5. The third instruction addi writes to the destination register t6. The fourth instruction is dependent on the second and third instructions simultaneously. Since each integer reservation station can only receive 2 instructions per clock cycle, for the convenience of explaining the principle of the present invention, and assuming that the free space of each reservation station is not less than 2, in actual situations when the free space of the reservation station is 1, only one instruction can be dispatched; or when the space of the reservation station is 0, no instruction can be dispatched. For the 4 instructions in the second clock cycle, slli, srli, add, and sd.

[0081] In this preferred embodiment, while dispatching instructions in the second clock cycle, the reservation stations issue the third instruction addi and the first instruction addiw. The source operand t1 of the sixth instruction srli depends on the destination register t1 of the fifth instruction. The source operand ra of the eighth instruction depends on the destination register ra of the fourth instruction add.

[0082] In this preferred embodiment, for the 4 instructions in the third clock cycle, c.addi, bne, sllw, and and. c.addi and bne are dispatched to the reservation station RV_1. and and sllw are dispatched to the reservation station RV_0. Meanwhile, the second instruction slli is issued to the execution unit for execution. The seventh instruction add is also issued to the execution unit for execution.

[0083] In this preferred embodiment, for the 3 instructions in the fourth clock cycle, addiw, slli, and addi. addiw is dispatched to the reservation station RV_0. slli and addi are dispatched to the reservation station RV_2.

[0084] Table 4 Reservation Station Status I

[0085]

[0086] In this preferred embodiment, in the 5th clock cycle, RV_0 issues the instruction "and", and RV_1 issues the instruction "c.addi". "c.addi" writes to the destination register a1, so the instruction "sllw" in reservation station RV_0 is awakened, and the instructions "slli" and "bne" in reservation station RV_1 are awakened.

[0087] In this preferred embodiment, in the 6th clock cycle, RV_0 issues the instruction "sllw", and RV_1 issues the instruction "bne".

[0088] In this preferred embodiment, in the 7th clock cycle, RV_0 issues the instruction "addiw", and RV_1 issues the instruction "slli".

[0089] In this preferred embodiment, in the 8th clock cycle, RV_1 issues the instruction "andi".

[0090] Table 5 Reservation Station Status II

[0091]

[0092] In summary, all RISC V instructions are integer-type instructions and are dispatched to two identical execution units for execution. Instruction dispatching is performed according to the dependencies between instructions.

[0093] Embodiment 5

[0094] This embodiment provides an instruction dispatching device based on instruction dependencies, on which instructions are stored. When the instructions are executed, an instruction dispatching method based on instruction dependencies is implemented.

[0095] In summary, the present invention can dispatch instructions with dependencies to the same reservation station according to the dependencies between instructions, so as not to block the parallel issue and execution of instructions without dependencies.

[0096] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An instruction dispatch method based on instruction correlation, characterized in that Including the following steps: After instruction renaming, judge the dependency relationship between instructions according to the decoding information of the decoder; Enter the instruction dispatch stage, dispatch instructions with dependency relationships to the same reservation station, and dispatch instructions without dependency relationships to different reservation stations; After the source operands of the instruction are ready, dispatch instructions without dependency relationships to different execution units for execution; After the instruction execution is completed, broadcast the destination register of the instruction and the instruction execution structure; When an instruction is dispatched to an execution unit, first judge whether there is free space in the reservation station RV corresponding to the execution unit. If there is no free space in the reservation station RV, the current instruction cannot be dispatched to this reservation station; When an instruction is dispatched to an execution unit, simultaneously judge whether there is free space in the reservation stations RV corresponding to other integer execution units. If there is no free space in the RVs corresponding to all integer execution units, the current instruction cannot be dispatched to the reservation station until there is free space in the reservation station, and then it can be dispatched to the reservation station.

2. The instruction dispatch method based on instruction dependency according to claim 1, wherein, In the method, after the false dependencies between instructions are renamed, the instructions can be executed simultaneously on the same type of execution components.

3. The instruction dispatch method based on instruction correlation according to claim 1, wherein In the method, the reservation station is a multi-port queue structure. When an instruction enters the reservation station, the instruction monitors the Forward path of the execution unit in real time.

4. A method for dispatching instructions based on instruction dependencies according to claim 3, characterized in that In the method, when the source operands of the instruction in the reservation station are awakened to the ready state, the instruction is dispatched in advance. In the EX0 stage, the instruction can Forward to the data returned by the execution unit.

5. A method for dispatching instructions based on instruction dependencies according to claim 1, characterized in that In the method, the execution unit generates an early wake-up control signal for waking up the instructions in the reservation station in advance according to the execution cycle of the instruction.

6. The instruction dispatch method based on instruction correlation according to claim 1, wherein In the method, when an instruction is dispatched to an execution unit, when there is free space in the reservation stations of multiple integer execution units, the integer instruction can be dispatched to the reservation station RV of any integer execution unit.

7. An instruction dispatching device based on instruction correlation, characterized in that Stored thereon are instructions, wherein when the instructions are executed, the instruction dispatch method based on instruction correlation according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Superscale pipeline reservation station processing instruction method and device

    CN104714780A

  • Method and system of distributed instruction execution unit

    CN112214241A