Register look-ahead renaming and static resource allocation method and device

By monitoring and adjusting resource configuration in the superscalar processor, the instruction blocking problem caused by insufficient queues after register renaming level is solved, and more efficient resource utilization and performance improvement is achieved.

CN120353586APending Publication Date: 2025-07-22XINGAOQIAO (SHANGHAI) SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432101.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing superscalar processors have insufficient idle entries in the register rename level after register rename level, resulting in instruction blocking, increasing pipeline bandwidth pressure and dynamic power consumption.

Method used

Monitor blocking events on the pipeline through embedded performance counters, identify blocking hazards after renaming the level in advance, and adjust resource configuration according to the blocking probability distribution, optimize resource allocation to reduce blocking.

Benefits of technology

It alleviates the bandwidth pressure of the later stage pipeline, reduces unnecessary dynamic power consumption, and improves system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353586A_ABST
    Figure CN120353586A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an out-of-order superscale processor look-ahead renaming and static resource allocation method and a related device. The method is suitable for various out-of-order superscale processors. According to the embodiment of the invention, the blocking drive of the renaming level is expanded; whether the distribution level and the emission level respectively have enough resources to complete distribution and emission is monitored in advance for the head instruction of the renaming buffer area to be renamed; and if the distributed or transmitted resources are insufficient, renaming of the head instruction of the renaming buffer area is prohibited. The flow line blocking preposition measure relieves the bandwidth pressure of each stage after the renaming stage, and meaningless dynamic power consumption is reduced. On the other hand, according to the embodiment of the invention, on the basis of the probability distribution, obtained through real-time monitoring and statistics, of the blocking sources, the resources are pertinently and continuously adjusted, and the regression test is iterated until out-of-order scheduling does not become the performance bottleneck of the whole out-of-order superscalar processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of superscalar processors, and in particular, to a method and device for register look-ahead renaming and static resource allocation. Background Art

[0002] The register renaming stage of a superscalar processor is located after the decode stage and is responsible for renaming the architecture register (AR) into a physical register (PR) to solve the write-after-write (WAW) or write-after-read (WAR) data dependency conflicts during out-of-order execution.

[0003] The register renaming stage usually undertakes three responsibilities: First, similar to each pipeline stage before the commit stage, when the processor's speculative execution fails or an interrupt / exception occurs, after receiving the pipeline flush signal squash broadcast forward by the commit stage, the renaming stage needs to empty the instructions younger than the squash instruction in each internal buffer queue of this stage, and roll back the register alias table (RAT) and the ROB to the state before speculative execution; Second, check whether there is an available physical register in the physical register free list (FreeList). If there is, map the AR to the PR of the head entry of the FreeList and maintain the latest AR-PR mapping relationship with the RAT; Third, write the latest AR-PR mapping relationship into the reordering buffer (ROB).

[0004] However, even if the renaming is successful, insufficient free entries in the subsequent related queues will also block the instructions. If it is possible to monitor in advance whether the resources required for dispatching or issuing the instructions to be renamed are sufficient and bring forward the possible pipeline blockage to the renaming stage, the bandwidth pressure of the subsequent pipeline can be alleviated and the unnecessary dynamic power consumption can be reduced.

[0005] Furthermore, according to the probability distribution of the detected blockage sources in advance, resources can be increased in a targeted manner, thereby reducing the probability of related blockages and improving the system throughput. Summary of the Invention

[0006] In view of this, the embodiments of the present invention provide a method and device for register look-ahead renaming and static resource allocation, which can identify in advance the potential blockages in the pipeline after the renaming stage.

[0007] The present invention pre - embeds performance counters to monitor blocking events on the pipeline, including: the ROB is full, the FreeList is full, the dispatch queue (Dispatch Buffer: DB) is full, the load queue (Load Queue: LQ) is full, the store queue (Store Queue: SQ) is full, and the total count of the above - mentioned blocking events.

[0008] Monitoring whether the ROB is full or the FreeList is full targets the entire reordering buffer (Reordering Buffer: RB).

[0009] Monitoring whether the DB is full or the exclusive queue (LQ / SQ) of memory - access instructions is full targets the head instruction of the RB.

[0010] When any of the above - mentioned situations occurs, rename blocking is performed.

[0011] When any of the above - mentioned situations occurs, increment the count of this performance event by one and increment the total count by one.

[0012] Formulate a combination of hardware parameters to configure the CPU, execute the CPU standard test set SPEC 2006, and obtain the probability distribution of the above - mentioned blocking events from the regression test results.

[0013] For high - blocking events, increase the relevant resources and perform a new regression test.

[0014] Among them, the method of increasing resources and performing a new regression test includes:

[0015] Continuously increase the scarce resources pointed to by the relevant blocking events at a certain step until the bottleneck of the processor no longer lies on the out - of - order scheduling data link, and fix the design parameters.

[0016] Preferably, seek the minimum resource combination to reduce blocking within the resource range defined by the product definition.

[0017] The method of determining that the performance bottleneck of the processor no longer lies on the out - of - order scheduling link includes:

[0018] The attenuation of the bandwidth at the rename stage / dispatch stage / issue stage compared to the theoretical bandwidth is less than the attenuation of the bandwidth at the remaining pipeline stages compared to the theoretical bandwidth.

[0019] The regression test method of gradually increasing resources at a certain step includes:

[0020] After uniformly configuring various resources, restart a new round of regression test.

[0021] Preferably, only adjust the scarce resources pointed to by the highest blocking probability in the results of the previous round of regression test, and restart a new round of regression test. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 Schematic diagram of the main data path of the speculative renaming provided in this embodiment;

[0024] Figure 2 Schematic diagram of the response process of the renaming stage to the pipeline flush instruction;

[0025] Figure 3 Schematic diagram of the logic for the renaming buffer head instruction in this embodiment to monitor in advance whether the dispatch buffer is full;

[0026] Figure 4 Schematic diagram of the logic for the memory access type instruction at the head of the renaming buffer in this embodiment to monitor whether the memory access queue is full;

[0027] Figure 5 Description of the state transition and data path inside the renaming stage;

[0028] Figure 6 Schematic diagram of the blocking distribution of the renaming stage after regression testing using a recommended hardware resource combination configuration provided in this embodiment.

[0029] Figure 7 Schematic diagram of the blocking distribution of the renaming stage after re-regression testing for implementing the static resource allocation strategy. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0031] As an optional example of the disclosed content of the embodiments of the present invention, Figure 1 Exemplarily shows the main data path of the speculative renaming. It should be noted that this block diagram is shown for the convenience of understanding the disclosed content of the embodiments of the present invention, and the embodiments of the present invention are not limited to Figure 1 the shown architecture.

[0032] Reference Figure 1 The pipeline architecture of an out-of-order superscalar processor may include: Fetch stage, Decode stage, Rename stage, Dispatch stage, Issue stage, Execute stage, Writeback stage, and Commit stage.

[0033] An out-of-order superscalar processor often adopts speculative execution to improve the performance of the processor. When speculative execution fails or an interrupt / exception occurs, the instruction commit stage generates a pipeline flush signal and broadcasts it forward to all levels of the pipeline. Each level of the pipeline first rolls back the state in response to the flush signal, and then executes the core logic of its own level. Refer to Figure 1 The processing flow of speculative renaming includes:

[0034] Step S10: Receive the pipeline flush signal passed back from the commit stage and respond.

[0035] The response process refers to Figure 2 and includes:

[0036] Step S11: Clear the instructions in each internal buffer queue of this level that are younger than the instruction pointed to by the flush signal;

[0037] Step S12: Roll back the Register Alias Table (RAT) and the Reordering Buffer (ROB) to the state before speculative execution.

[0038] Continue to refer to Figure 1 and traverse the entire RB to monitor whether the resources required for renaming are satisfied. The specific monitoring methods include:

[0039] Step S21: Monitor whether the ROB is full;

[0040] Step S31: Monitor whether the FreeList is full.

[0041] Continue to refer to Figure 1 The early monitoring logic of the speculative renaming includes:

[0042] Step S41: Monitor whether there are free entries in the Dispatch Buffer (DB) to which the instruction at the head of the RB belongs for its dispatch. Refer to Figure 3 Specifically, it includes:

[0043] Step S42: Parse the micro-operation (uOp). Specifically, if it is a floating-point instruction, i.e., the isFloating flag bit is high, it is associated with the FPU, and the corresponding dispatch buffer number dispIdx = 4; if it is a memory access instruction (isMemRef) or a read / write barrier instruction (isReadBarrier / isWriteBarrier), it is associated with the LSU, and the corresponding dispatch buffer number dispIdx = 3; if it is a jump control or branch prediction instruction, it is associated with the BPU, and the corresponding dispatch buffer number dispIdx = 2; if it is an integer multiplication or division, it is associated with the IntMDU, and the corresponding dispatch buffer number dispIdx = 1; if it is an integer arithmetic logic operation, it is associated with the IntAlu, and the corresponding dispatch buffer number dispIdx = 0.

[0044] Optionally, considering that the proportion of operations in the IntMDU is relatively small compared to the IntAlu, step S42 merges a certain proportion of integer arithmetic logic operation instructions into the IntMDU unit to balance the load of the execution units.

[0045] Step S43: Map the target DB according to the parsed dispatch buffer number dispIdx.

[0046] Step S44: Monitor whether there is an idle entry in the target DB, i.e., DB[dispIdx], for dispatching. If DB[dispIdx] is full, renaming of the instruction at the head of the RB is prohibited.

[0047] Step S51: If the instruction at the head of the RB is a memory access instruction, monitor whether the memory access queue is full. According to the detailed instructions, refer to Figure 4 , including:

[0048] Step S52: If the instruction at the head of the rename buffer is a load instruction, monitor whether the load queue is full;

[0049] Step S53: If the instruction at the head of the rename buffer is a store instruction or an atomic instruction, monitor whether the store queue is full.

[0050] Continue to refer to Figure 1 , when the blocking monitoring of the rename stage itself described in steps S21 and S31 returns true or the early monitoring of the blocking of the subsequent pipeline described in steps S41 and S51 returns true, the blocking drive signal of the rename stage is pulled high, and the rename stage enters the blocking state.

[0051] The state transition after the rename stage enters the blocking state refers to Figure 5 , including:

[0052] Step S70: At this time, the state machine inside the rename stage is in the locked state 1. Through the chip select circuit, the instructions in the input buffer insts are transferred to the blocking buffer queue skidBuffer. It should be noted that the buffer path also gives higher priority to obeying the pipeline flush signal sent back from the subsequent commit stage. When the flush signal squash is pulled high, all the instructions in the rename stage input buffer insts and the blocking buffer queue skidBuffer that are younger than the squash instruction need to be cleared.

[0053] Step S80: After the blocking is lifted, the state machine inside the rename stage jumps to the unlocked state 2. Through the chip select circuit, the instructions temporarily stored due to blocking are transferred from skidBuffer to RB. It should be noted that after all the instructions in skidBuffer are transferred to RB, the state machine returns to the running state 0, and the instructions directly enter RB from the input buffer insts.

[0054] In an optional implementation, relevant resources are pre-allocated.

[0055] Specifically, the number of integer physical registers numPhysIntRegs can be set to 128, the number of floating-point physical registers numPhysFloatRegs to 64, the total number of ROB entries numROBEntries to 192, the total number of entries in the load queue LQEntries to 56, the total number of entries in the store queue SQEntries to 48, and the depth of the dispatch buffer DBDepth to {IntAlu:16,IntMDU:8,BPU:8,LSU:16,FPU:16}. In this embodiment, this configuration combination is called Configuration 1.

[0056] Preferably, considering that the Load instruction appears frequently in the SPEC 2006 CPU standard performance test and the penalty delay after a load miss is relatively high, LQEntries is expanded to 64, and the other parameters remain unchanged. This configuration combination is called Configuration 2.

[0057] Select the above two sets of configuration parameters and perform regression on 12 types of SPEC 2006 CPU standard performance test templates including astar, bzip2, gcc, gobmk, h264ref, hmmer, libquantum, mcf, omnetpp, perlbench, sjeng, and xalancbmk.

[0058] Refer to Figure 6, the horizontal axis corresponds to the above-mentioned 12 types of standard test sets. The total number of blocking events during renaming, totalStallsDuringRenaming, is represented by a broken line, and the specific value is measured by the left semi-axis. Among them, the black broken line with rectangular nodes reflects the total blocking situation of Configuration 1 above, and the red broken line with circular nodes reflects the total blocking situation of Configuration 2 above. The proportion of various blocking sources is depicted by a bar chart, and the corresponding proportion is distributed on the right semi-axis. Among them, the bar chart filled with star symbols on the left of each group of test sets represents the blocking distribution of Configuration 1, and the bar chart filled with circles on the right of each group of test sets represents the blocking distribution of Configuration 2.

[0059] Continue to refer to Figure 6 , according to the current blocking distribution, the fullness of the DB or the fullness of the LQ / SQ accounts for the main proportion.

[0060] Therefore, in this embodiment, the relevant resources are expanded, including:

[0061] Set the depth of each dispatch buffer to 24, set LQEntries to 192, and set SQEntries to 128, and the remaining parameters are the same as those of Configuration 2. This parameter configuration combination is Configuration 3.

[0062] Select Configuration 3 and return to the test again. The blocking distribution refers to Figure 7 .

[0063] Among them, the situation where LQEntries or SQEntries are full basically disappears, indicating that the settings of LQEntries = 192 and SQEntries = 128 have met the emission requirements;

[0064] The situation where the dispatch buffer is full still accounts for a considerable proportion;

[0065] Most of the blocking at the renaming level comes from the lack of free entries in the ROB.

[0066] Optionally, the depth of the dispatch buffer can be further increased to alleviate the situation where the dispatch buffer is full.

[0067] Those of ordinary skill in the art can realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, they can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the embodiments of this application.

[0068] An embodiment of the present invention further provides a processor, which includes the look-ahead monitoring circuit and the static resource allocation device described in the above embodiment.

[0069] An embodiment of the present invention further provides a chip, which includes the processor described in the above embodiment.

[0070] Although the embodiments of the present invention are disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.

Claims

1. A method and related device for speculative renaming and static resource allocation of an out-of-order superscalar processor, characterized in that The method includes: The blocking-driven logic of the classic renaming stage is extended.

2. The method according to claim 1, wherein In addition to real-time monitoring of the situation where the reordering buffer (ROB) that directly causes renaming failure and the physical register free list (FreeList) are full, the situation where the queue resources of the subsequent dispatch stage and issue stage are occupied is also monitored, so as to realize the pre-positioning of pipeline blocking, relieve the bandwidth pressure of each stage after the renaming stage, and reduce unnecessary dynamic power consumption.

3. The method of blocking in front according to claim 2, characterized in that, Before renaming the head instruction of the rename buffer (RenameBuffer: RB), it is judged whether there are free entries in the dispatch buffer (Dispatch Buffer: DB) to support the dispatch of this instruction. Specifically, it includes: Parse the micro-operations (uOps) of the head instruction in the RB, and check whether the dispatch buffer dedicated to this uOp, that is, DB[uOp], is full. If it is full, the renaming of the head instruction in the RB is prohibited.

4. The method of blocking in front according to claim 2, characterized in that, If the head instruction in the RB is a memory access instruction, before renaming it, it is judged whether there are free entries in the dedicated queue of the memory access instruction to carry the instruction issue task, including: If the head instruction in the RB is a load instruction, judge whether the load queue (Load Queue: LQ) is full. If it is full, the renaming of the head instruction in the RB is prohibited; If the head instruction in the RB is a store instruction, judge whether the store queue (Store Queue: SQ) is full. If it is full, the renaming of the head instruction in the RB is prohibited.

5. The method according to any one of claims 2-4, characterized in that, Introduce performance counters to count blocking events, including: Introduce LackROBEntries to represent that the renaming fails directly due to the lack of free entries in the ROB. If it is monitored that the ROB is full, increment the LackROBEntries count by one; Introduce NoFreeList to represent that the renaming fails directly due to the lack of physical register free list entries. If it is monitored that the FreeList is full, increment the NoFreeList count by one; Introduce a performance counter for each discrete DB[uOp], including: Introduce DispBuf0Full to represent that the DB dedicated to the integer arithmetic logic unit (IntAlu), that is, DB[IntAlu], is full, and it is quantified by incrementing the DispBuf0Full count by one; Introduce DispBuf1Full to represent that the DB dedicated to the integer multiply-division unit (IntMDU), that is, DB[IntMDU], is full, and it is quantified by incrementing the DispBuf1Full count by one; Introduce DispBuf2Full to represent that the dedicated DB of the Branch Prediction Unit (BPU), i.e., DB[BPU], is full, and it is quantified by incrementing the count of DispBuf2Full by one; Introduce DispBuf3Full to represent that the dedicated DB of the Load Store Unit (LSU), i.e., DB[LSU], is full, and it is quantified by incrementing the count of DispBuf3Full by one; Introduce DispBuf4Full to represent that the dedicated DB of the Floating-point Unit (FPU), i.e., DB[FPU], is full, and it is quantified by incrementing the count of DispBuf4Full by one.

6. Introduce LQFull to represent that the RB head instruction is a load instruction and it is monitored that LQ is full, then increment the count of LQFull by one; Introduce SQFull to represent that the RB head instruction is a store instruction and it is monitored that SQ is full, then increment the count of SQFull by one; Introduce the total blocking count totalStallsDuringRenaming. When any of the above counts is incremented, increment the count of totalStallsDuringRenaming by one.

7. The method according to claim 5, characterized in that After each round of regression testing, count the proportion of each of the above blocking sources in all blocking events.

8. The method according to claim 6, characterized in that, According to the probability distribution of each blocking source obtained from the statistics, increase the relevant resources and conduct regression testing again.

9. The method according to claim 7, wherein Continuously increase the scarce resources pointed to by the blocking events at a certain step until the bottleneck of the processor no longer lies on the out-of-order scheduling data link, and then fix the design parameters.

10. The method according to claim 8, wherein The method of continuously increasing the scarce resources pointed to by the blocking events at a certain step includes: After uniformly configuring various resources, restart a new round of regression testing; Preferably, only adjust the resources involved in the item with the highest blocking probability in the results of the previous round of regression testing, and restart a new round of regression testing.

11. The method according to claim 8, wherein The method of determining that the performance bottleneck of the processor no longer lies on the out-of-order scheduling data link includes: The bandwidth attenuation of the rename stage / dispatch stage / issue stage compared to the theoretical bandwidth is less than the bandwidth attenuation of the remaining pipeline stages compared to the theoretical bandwidth.

12. The method according to claim 7, wherein If the product definition clearly stipulates the resource upper limits of each component, then seek the minimum resource combination for reducing blocking within the limited range.