Two-stage reservation station
Through the cooperation of the two-stage RSV design and LLC predictor, the cycle time challenges and poor CPU performance caused by increasing the RSV size are solved, and the instruction throughput and cycle time are improved, which enhances the overall performance of the CPU.
Patent Information
- Application Number
- CN202280101424.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-05-27
AI Technical Summary
When increasing the reserve station (RSV) size to improve instruction throughput and frequency, cycle time challenges and tight restrictions on wake-up-selecting timing paths result in poor overall CPU performance.
Using a two-stage RSV design, the RSV is divided into a wait buffer (WB) and an RSV cluster, and the LLC predictor actively predicts the instructions that miss LLC, and directs its dependencies to the WB instead of passing it directly to the RSV cluster.
Through this design, the instruction throughput and cycle time can be improved without increasing the RSV size and the overall performance of the CPU.
Smart Images

Figure CN120051766A_ABST
Abstract
Description
Background Art
[0001] This specification relates to an apparatus including one or more reservation stations (RSVs) that may assist out-of-order (OoO) instruction execution on a computing device.
[0002] In modern out-of-order (OoO) processors, instruction throughput (e.g., instructions per cycle (IPC)) is typically increased by increasing the OoO window size. A reservation station (RSV) is one of the components that typically limits the window size. A larger RSV can extract instruction and memory level parallelism, which helps increase IPC.
[0003] However, increasing the RSV size poses cycle time challenges and limits frequency. The wakeup-selection timing path, which is closely related to the RSV size, is one of the most stringent timing paths in modern OoO CPUs and typically limits the overall CPU frequency. Increasing the RSV size puts stress on each component in the wakeup-selection path. For example, the wakeup latency increases with increasing RSV size due to increased load on the tag broadcast line. The selection latency increases because more instructions participate in the selection process and it takes longer to determine the selection priority among them. Since both IPC and frequency contribute to overall performance, simply increasing the RSV size may not result in higher overall performance. Summary of the Invention
[0004] This specification describes systems and methods for implementing an RSV with multiple levels to increase the "effective capacity" in a cycle-time friendly manner.
[0005] "RSV clustering" is the process of dividing an RSV into smaller "clusters" or "groups", each of which is designed to handle a specific type of instruction. In the case where all instructions are fed to all clusters, this structure is referred to as a "fully unified" or "monolithic" RSV. On the other hand, in a "fully distributed" or "fragmented" RSV, instructions are only fed to their designated clusters. Each of these two structures has advantages and disadvantages for IPC and cycle time. The two-level RSV organization presented in this specification depicts an apparatus that does not rely on a "fully distributed" or "fully unified" design to obtain similar performance benefits provided by each structure.
[0006] This two - level RSV design takes advantage of the fact that, most commonly, RSVs are filled with instructions that miss in the last - level cache (LLC) and their dependencies. The LLC is defined as the last cache before the CPU accesses memory. Typically, it may take many (e.g., more than 100) cycles to service an instruction or dependency that misses in the LLC. This can lead to a chain reaction where the RSV is prevented from providing newer instructions. In other words, the "queue" of the RSV may be occupied by these instructions that miss in the LLC and their dependencies, and thus prevent the RSV from storing instructions that could be processed more efficiently.
[0007] To address this issue, the system attempts to proactively predict which instructions will miss in the LLC and direct the dependencies of these instructions to a separate cycle - friendly structure. In some implementations, this structure can be referred to as a wait buffer (WB). The WB is a structure separate from the normal RSV clusters. In some implementations, the RSV is divided into two levels; the first level consists of the WB, and the second level contains one or more RSV clusters. In some implementations, instructions predicted to miss in the LLC will direct their dependencies to the WB in the first level rather than passing them directly to the RSV clusters in the second level. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is an overview of an example system implementation.
[0009] Figure 2 is a detailed view of an example system implementation.
[0010] Figure 3 is an example process in which instructions are processed by an example system implementation.
[0011] Figure 4 is an example process in which the estimation logic of the LLC predictor is improved. DETAILED DESCRIPTION
[0012] Figure 1 is an overview of an example system. System 100 has an extraction module 102, a decoding module 104, a wait buffer (WB) 105, a distribution module 106, one or more RSVs 108, a re - order buffer 110, a commit module 112, and a store buffer 114. The various "modules" mentioned above can be implemented using various logic circuit components, including AND gates, OR gates, NOT gates, NAND gates, or XOR gates. Other implementations may choose to use other circuit components.
[0013] The extraction module 102 retrieves incoming instructions for decoding. The decoding module 104 analyzes the incoming function to determine its consumer. In some implementations, the output of the decoding module 104 is used to determine whether the decoded instruction corresponds to an instruction that may miss the LLC. Once it is determined that the instruction may miss the LLC, a bank in WB 105 is allocated to the dependencies of the instruction. If the instruction is unlikely to be an LLC miss, or if the instruction has met the requirements to leave WB 105, the instruction is sent to the distribution module 106, which sends the instruction to RSV 108. RSV 108 can take various forms between a fully distributed system and a fully unified system. RSV 108 can also take various forms of clustering, where components of instruction types can be assigned to a certain number of RSVs 108. After being invoked from RSV 108, the instruction is then processed by the reorder buffer 110 and the commit module 112, and then reaches the store buffer 114.
[0014] Figure 2 is a detailed view of an example system implementation 200. The system 200 includes a decoding module 202, an LLC predictor 204, and a WB free list 206, a rename module 208, a WB 210, a WB bank 212, a "WB BankID (WB bank ID)" 214, a WB multiplexer 216, one or more RSV multiplexers 218, one or more RSV clusters 220, and one or more execution channels 222.
[0015] The LLC predictor 204 is used to determine which instructions will likely miss the LLC. In some implementations, this prediction can occur during the decoding module 202 of the RSV. Before use, the LLC predictor 204 can undergo initial training on what instructions have a high likelihood of missing the LLC. When the LLC predictor 204 detects that an instruction will miss the LLC, if the open BankID 214 as identified by the WB free list 206 is available, the instruction will request a bank 212 in WB 210. At this time, additional identifiers can be assigned to the instruction, for example, tags for physical register number (PRN) and "BankIDValid (bank ID valid)". In some implementations, this process can be handled by the rename module 208.
[0016] In some implementations, WB 210 can be split into a certain number of banks 212. The number of banks 212 can be further divided into multiple entries, each entry can be occupied by a single instruction, and is constructed such that they are First-In-First-Out (FIFO). The FIFO structure allows bank-level wake-up (i.e., all instructions in the same bank 212), which reduces design complexity. Other implementations within the scope of the claims can use other WB 210 structures or processes.
[0017] When leaving WB 210, the instruction chain will leave in a specific format (e.g., in the FIFO-based allocation order). In some implementations, this process can be handled by the WB multiplexer 216. Then the instructions leaving WB 210 can be provided to the RSV cluster 220. In implementations where multiple instruction types are handled by the same RSV cluster 220, the RSV multiplexer 218 can be used to dispatch the instructions to the appropriate RSV. Then, the RSV dispatches an execution channel 222 for the instructions ready for execution. In some implementations, each RSV cluster 220 is configured to handle different instruction types. For example, each RSV cluster can be configured to handle different types of instructions or instruction categories. For example, different RSV clusters 220 can be dispatched to handle load, store, functional operations, basic mathematical operations, and complex mathematical operations respectively. Other alternative implementations can use any appropriate arrangement of RSV clusters with instruction types or categories.
[0018] Figure 3 is a flowchart of an example process for using a wait buffer upon a predicted load miss. The example process can be executed by any suitable processor configured to operate in accordance with this specification.
[0019] The LLC predictor can undergo initial training (310) on what instructions have a high likelihood of missing the LLC. In some implementations, this training can be based on known program counter (PC) data. Training is described in more detail below with reference to Figure 4 The predictions used by the LLC predictor 204 can also be improved along with RSV operations to better predict which instructions will miss the LLC. In some implementations, the LLC predictor 204 can include a table with multiple entries, each entry having an N-bit saturating counter. The table can be indexed by various means, including the load instruction address, the hash value of the load instruction address, the Global Load Hit / Miss History (GLHR), the load path history, or other parameters that can be easily obtained from the PC. The table can also be indexed by a combination of the above parameters.
[0020] In an implementation using GLHR, the GLHR may include an "N"-bit shift register that is updated at the time of the instruction LLC miss prediction. If an instruction is predicted to miss the LLC, a value of "1" is assigned to the GLHR. Alternatively, if an instruction is expected to hit in the LLC, a value of "0" is assigned to the GLHR. In another implementation using the load path history to update the LLC predictor 204, the operation may include a hash value of the PC bits from "N" previous load instructions.
[0021] Additionally, in a multi-core system where LLC hit / miss information is not easily available, a proxy can be used to train the LLC predictor 204. In this case, the number of cycles taken by the instruction at the head of the reorder buffer (ROB) can be used to train the LLC predictor 204. A certain number of cycles, such as 50, can be dispatched, beyond which the instruction is considered a miss and the corresponding counter is incremented. Otherwise, the instruction is considered a hit and the corresponding counter is decremented.
[0022] After initial training, the LLC predictor 204 decodes instructions during operation and makes a prediction (320) as to which instructions will miss the LLC. If the LLC predictor 204 predicts that an instruction will miss the LLC, the dependencies of that instruction are moved to the memory bank 212 within the WB 210 (330). Once in the WB 210, the dependent instruction can be dispatched a "BankID" 214 corresponding to its target logical register number (LRN). This information can be arranged in a format that is easy to reference, for example, a lookup table. When a dependent instruction corresponding to the same LRN is detected, the BankID 214 can then be used to place the dependent instruction in the same memory bank within the WB 210 as the previous instruction. If a dependent has more than one previous instruction assigned to a unique memory bank in the WB, the system can follow a predetermined response, for example, assigning the dependent instruction to the memory bank 212 with the lowest occupancy within the WB 210.
[0023] The BankID 214 for each instruction can also be shared with the Load Store Unit (LSU). When it is detected that an instruction is ready to leave the WB 210, for example, because a load instruction has been completed, the LSU sends a "wake-up" (340) to all dependent instructions in the same memory bank 212. Additionally, the LSU can take other actions. For example, the LSU can send an early warning to the LRN or other components that a wake-up is in progress. The LSU can also trigger an instruction chain to leave the WB 210 early. When leaving the WB 210, the instruction chain leaves in a specific format (e.g., in the FIFO-based allocation order) (350). In this method, there is no limit to the number of instruction chains that can be woken up per cycle. Other implementations within the scope of the claims can utilize different wake-up methods or can have the LSU perform different actions.
[0024] If multiple instruction chains are woken up in the same cycle, then the system can follow a specific arbitration process to control how the instructions leave the WB 210 (360). In some implementations, round-robin can be implemented to determine the order. In other implementations, an age-based method can be preferred, where older instruction memory banks 212 have priority. Additionally, this age preference can be extended to other instructions not dispatched to the WB 210, such that instruction chains exiting the WB 210 have priority over instructions directly exiting the decode 202.
[0025] Figure 4 is a flowchart of an example process 400 for improving the LLC predictor. The example process can be executed by any suitable processor configured according to this specification. For example, the processor can execute the example process during instruction execution to continuously improve the LLC predictor.
[0026] Following the initial LLC predictor training (410) as described in Figure 3 it can be desirable to improve aspects of the LLC predictor's estimation logic to better identify problem instructions. In some implementations, a counter can be assigned to each load instruction (420) to form an entry table. In some cases, the table can initially be indexed based on the hash value of the load instruction's PC. Other implementations can choose to use a "tagged" LLC predictor that utilizes a Content Addressable Memory (CAM) structure that performs comparisons of load instruction tags. During execution, load instructions are monitored to determine if any LLC misses occur (430).
[0027] When it is detected that a load instruction has missed the LLC (430), the counter assigned to that load instruction is incremented by a fixed number (440). In some implementations, this number can be an integer (e.g., "1"). In the case where the load instruction hits without missing the LLC, the counter assigned to that load instruction is decremented by a fixed number (450). In some implementations, this number can be an integer (e.g., "1"). After the counter for the load instruction has been updated, the system then continues execution using the updated counter (460).
[0028] The above describes an example implementation for updating LLC predictor logic. Other implementations may choose to use different variations of the described process, e.g., incrementing the counter in a different way. Other implementations may choose to use an entirely different process, including using other data from the computing system available to the LLC predictor 204.
[0029] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus.
[0030] The term "data processing apparatus" refers to data processing hardware and includes all kinds of devices, apparatus, and machines for processing data, such as including programmable processors, computers, or multiple processors or computers. The apparatus can also be or further include special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to the hardware, the apparatus may optionally include code for creating an execution environment for the computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0031] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages); and it can be deployed in any form, which includes as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may or may not correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (such as files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0032] For a system of one or more computers that is configured to perform particular operations or actions, it means that the system has installed thereon software, firmware, hardware, or a combination thereof that, in operation, causes the system to perform these operations or actions. For one or more computer programs that are configured to perform particular operations or actions, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0033] As used in this specification, an "engine" or "software engine" refers to a software-implemented input / output system that provides an output different from the input. An engine can be a functional coding block, such as a library, platform, software development kit ("SDK"), or object. Each engine can be implemented on any suitable type of computing device (e.g., a server, mobile phone, tablet computer, notebook computer, music player, e-book reader, laptop or desktop computer, PDA, smart phone, or other fixed or portable device) that includes one or more processors and a computer-readable medium. Additionally, two or more of the engines can be implemented on the same computing device or on different computing devices.
[0034] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, a special logic circuit system such as an FPGA or ASIC, or by a combination of a special logic circuit system and one or more programmed computers.
[0035] A computer suitable for executing a computer program can be based on a general or special purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for carrying out or executing instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or will be operatively coupled to receive data from, or transfer data to, one or more mass storage devices, or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0036] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0037] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse, a trackball, or a presence-sensitive display or other surface) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user can be in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a user device in response to a request received from the web browser. Further, a computer can interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smart phone running a messaging application) and receiving responsive messages from the user in response.
[0038] In addition to the embodiments described above, the following embodiments are also innovative:
[0039] Example 1 is a computing device, including:
[0040] a plurality of processing cores; and
[0041] a reservation station, the reservation station including a waiting buffer, a plurality of clusters, and circuitry configured to coordinate the selection of instructions for out-of-order execution on the plurality of processing cores, wherein the reservation station is configured to:
[0042] predict that a load instruction will result in a cache miss,
[0043] when predicting that a load instruction will result in a cache miss, i) execute the load instruction using a cluster in the plurality of clusters, and ii) store one or more dependency instructions of the load instruction in the waiting buffer, and
[0044] when completion of execution of the load instruction, i) retrieve one or more of the dependency instructions from the waiting buffer, and ii) execute the one or more dependency instructions using the plurality of clusters.
[0045] Example 2 is the computing device as described in Example 1, wherein the waiting buffer includes a plurality of banks, and wherein storing the one or more dependency instructions of the load instruction includes storing all the dependency instructions of the load instruction in the same bank of the waiting buffer.
[0046] Example 3 is the computing device as described in Example 2, wherein each bank entry of the waiting buffer includes a logical register number, a physical register number, a bank id, and a validity value.
[0047] Example 4 is the computing device as described in Example 3, wherein each bank is organized as a first-in-first-out queue.
[0048] Example 5 is the computing device as described in any one of Examples 1 to 4, wherein the reservation station further includes prediction circuitry configured to generate a prediction as to whether the load instruction will result in a cache miss.
[0049] Example 6 is the computing device as described in Example 5, wherein the prediction circuitry includes a counter that increments based on a global load hit / miss history.
[0050] Example 7 is the computing device as described in Example 5, wherein the prediction circuitry includes the number of cycles taken by an instruction at the head of a reorder buffer.
[0051] Example 8 is the computing device as described in Example 5, wherein the prediction circuitry includes a hash load.
[0052] Embodiment 9 is a computing device as described in any one of Embodiments 1 to 8, wherein the cache miss is a miss in the last-level cache of the computing device.
[0053] Embodiment 10 is a computing device as described in any one of Embodiments 1 to 9, wherein two or more of the plurality of clusters are dedicated to executing a mix of different instruction types.
[0054] Embodiment 11 is a computing device as described in Embodiment 10, wherein the first cluster is dedicated to executing branch instructions and simple instructions that are executed in a single cycle.
[0055] Embodiment 12 is a computing device as described in Embodiment 11, wherein the second cluster is dedicated to executing simple instructions and multi-cycle instructions.
[0056] Embodiment 13 is a computing device as described in any one of Embodiments 1 to 12, wherein if multiple memory banks are activated in the same clock cycle, the reservation station is configured to perform memory bank-level arbitration.
[0057] Embodiment 14 is a method performed by a computing device including a plurality of processing cores, a reservation station including a waiting buffer, a plurality of clusters, and circuitry configured to coordinate selection of instructions for out-of-order execution on the plurality of processing cores, the method including:
[0058] Predicting, by the reservation station, that a load instruction will result in a cache miss,
[0059] When predicting that the load instruction will result in a cache miss, i) executing the load instruction using a cluster among the plurality of clusters, and ii) storing one or more dependency instructions of the load instruction in the waiting buffer, and
[0060] When completion of execution of the load instruction, i) fetching one or more of the dependency instructions from the waiting buffer, and ii) executing the one or more dependency instructions using the plurality of clusters.
[0061] Embodiment 15 is the method as described in Embodiment 14, wherein the waiting buffer includes a plurality of memory banks, and wherein storing the one or more dependency instructions of the load instruction includes storing all dependency instructions of the load instruction in the same memory bank of the waiting buffer.
[0062] Embodiment 16 is the method as described in Embodiment 15, wherein each memory bank entry of the waiting buffer includes a logical register number, a physical register number, a memory bank id, and a validity value.
[0063] Example 17 is the method as described in Example 16, wherein each bank is organized as a first-in first-out queue.
[0064] Example 18 is the method as described in any one of Examples 14 to 17, wherein the reservation station further includes a prediction circuitry configured to generate a prediction as to whether the load instruction will result in a cache miss.
[0065] Example 19 is the method as described in Example 18, wherein the prediction circuitry includes a counter that is incremented based on a global load hit / miss history.
[0066] Example 20 is one or more non-transitory computer storage media encoded with computer program instructions that, when executed by one or more computers, cause the one or more computers to perform operations that include:
[0067] predicting, by a reservation station, that a load instruction will result in a cache miss,
[0068] when predicting that the load instruction will result in a cache miss, i) executing the load instruction using a cluster among a plurality of clusters, and ii) storing one or more dependency instructions of the load instruction in a waiting buffer, and
[0069] when execution of the load instruction is completed, i) fetching one or more of the dependency instructions from the waiting buffer, and ii) executing the one or more dependency instructions using the plurality of clusters.
[0070] Although this specification contains many specific implementation details, these details should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be deleted from the combination, and the claimed combination may cover a sub-combination or a variant of a sub-combination.
[0071] Similarly, although the operations are depicted in the drawings in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In some circumstances, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0072] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As one example, the processes depicted in the figures do not necessarily need the particular order shown or an ordered order to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A computing device, comprising: a plurality of processing cores; and a reservation station, the reservation station including a waiting buffer, a plurality of clusters, and circuitry configured to coordinate the selection of instructions for out-of-order execution on the plurality of processing cores, wherein the reservation station is configured to: predict that a load instruction will result in a cache miss, when predicting that a load instruction will result in a cache miss, i) execute the load instruction using a cluster among the plurality of clusters, and ii) store one or more dependency instructions of the load instruction in the waiting buffer, and when execution of the load instruction is completed, i) fetch one or more of the dependency instructions from the waiting buffer, and ii) execute the one or more dependency instructions using the plurality of clusters.
2. The computing device according to claim 1, wherein the waiting buffer includes a plurality of banks, and wherein storing the one or more dependency instructions of the load instruction includes storing all the dependency instructions of the load instruction in the same bank of the waiting buffer.
3. The computing device according to claim 2, wherein each bank entry of the waiting buffer includes a logical register number, a physical register number, a bank id, and a validity value.
4. The computing device according to claim 3, wherein each bank is organized as a first-in-first-out queue.
5. The computing device according to any one of claims 1 to 4, wherein the reservation station further includes prediction circuitry configured to generate a prediction as to whether the load instruction will result in a cache miss.
6. The computing device according to claim 5, wherein the prediction circuitry includes a counter that is incremented based on a global load hit / miss history.
7. The computing device according to claim 5, wherein the prediction circuitry includes the number of cycles taken by an instruction at the head of a reorder buffer.
8. The computing device according to claim 5, wherein the prediction circuitry includes a hash load.
9. The computing device according to any one of claims 1 to 8, wherein the cache miss is a miss in the last-level cache of the computing device.
10. The computing device according to any one of claims 1 to 9, wherein two or more of the plurality of clusters are dedicated to executing a mix of different instruction types.
11. The computing device according to claim 10, wherein a first cluster is dedicated to executing branch instructions and simple instructions that execute in a single cycle.
12. The computing device according to claim 11, wherein a second cluster is dedicated to executing simple instructions and multi-cycle instructions.
13. The computing device according to any one of claims 1 to 12, wherein if multiple banks are activated in the same clock cycle, the reservation station is configured to perform bank-level arbitration.
14. A method performed by a computing device, the computing device comprising: a plurality of processing cores; Reservation station, the reservation station includes a waiting buffer, a plurality of clusters, and circuitry configured to coordinate selection of instructions for out-of-order execution on the plurality of processing cores, the method comprising: Predicting, by the reservation station, that a load instruction will result in a cache miss, When predicting that a load instruction will result in a cache miss, i) executing the load instruction using a cluster among the plurality of clusters, and ii) storing one or more dependent instructions of the load instruction in the waiting buffer, and When completion of execution of the load instruction, i) obtaining one or more of the dependent instructions from the waiting buffer, and ii) executing the one or more dependent instructions using the plurality of clusters.
15. The method according to claim 14, wherein, the waiting buffer includes a plurality of banks, and storing the one or more dependent instructions of the load instruction includes storing all dependent instructions of the load instruction in the same bank of the waiting buffer.
16. The method according to claim 15, wherein, each bank entry of the waiting buffer includes a logical register number, a physical register number, a bank id, and a validity value.
17. The method according to claim 16, wherein, each bank is organized as a first-in-first-out queue.
18. The method according to any one of claims 14 to 17, wherein, the reservation station further includes a prediction circuitry configured to generate a prediction as to whether the load instruction will result in a cache miss.
19. The method according to claim 18, wherein, the prediction circuitry includes a counter that is incremented based on a global load hit / miss history.
20. One or more non-transitory computer storage media encoded with computer program instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising: Predicting, by a reservation station, that a load instruction will result in a cache miss, When predicting that a load instruction will result in a cache miss, i) executing the load instruction using a cluster among a plurality of clusters, and ii) storing one or more dependent instructions of the load instruction in a waiting buffer, and When completion of execution of the load instruction, i) obtaining one or more of the dependent instructions from the waiting buffer, and ii) executing the one or more dependent instructions using the plurality of clusters.