Apparatus and method for operating an emission queue

By introducing the second segment in the issuing queue and adopting allocation standards and postponement mechanisms, the problem that modern OOO processor performance is limited by issuing queue size is solved, and the effect of increasing the instruction window size and improving processor performance is achieved.

CN112416244BActive Publication Date: 2025-06-20ARM LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010812801.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-21
Filing Date
2020-08-13
Publication Date
2025-06-20
Estimated Expiration
2040-08-13

AI Technical Summary

Technical Problem

The performance of modern OOO processors is limited by the size of the emitter queue, resulting in a decrease in frequency at which the processor can be operated, which in turn affects performance.

Method used

By introducing a second segment into the issuance queue and adopting allocation criteria and deferred mechanisms, the effective capacity of the issuance queue is extended while maintaining the performance of the scheduler loop.

Benefits of technology

Increases the effective capacity of the issuing queue and expands the instruction window size, thereby improving processor performance without affecting the frequency of scheduler cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112416244B_ABST
    Figure CN112416244B_ABST
Patent Text Reader

Abstract

Apparatus and methods for operating an issue queue are provided. The issue queue has a first section and a second section, each of these sections including a number of entries, and each entry therein being used to store operation information identifying an operation to be performed by a processing unit. An allocation circuit determines, for each item of received operation information, whether to allocate the operation information to an entry in the first section or an entry in the second section. The operation information identifies not only the associated operation, but also each source operation object required by the associated operation and the availability of each source operation object. A selection circuit selects, during a given selection iteration, an operation to be issued to the processing unit from the issue queue, and selects the operation from among the operations for which the required source operation objects are available. An availability update circuit is used to update the source operation object availability for each entry whose operation information identifies the target operation object of the selected operation in the given selection iteration as a source operation object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to apparatuses and methods for operating an issue queue. Background Art

[0002] Instructions fetched from memory for execution by a processing unit are decoded to identify operations that the processing unit is to perform in order to execute the instructions. Sometimes the operations are broken down into one or more micro-operations (also referred to as micro-ops). Here, operations and micro-operations will be collectively referred to as operations hereinafter.

[0003] An issue queue is typically used to temporarily buffer operations identified by decoding instructions before the operations are issued to relevant execution units within the processing unit. An execution unit will not be able to execute an operation until the source operands required for the operation are available, and thus the operations can be temporarily buffered in the issue queue until the source operands are available.

[0004] To improve performance, many modern processors support out-of-order (OOO) execution of instructions, where instructions are executed out of order relative to the original program order in order to strive to increase the throughput of the processing unit, and the retirement of instructions then occurs in order. In such a system, an issue queue is one of the data structures that can be used to support OOO execution.

[0005] However, the performance of modern OOO processors is constrained by the depth of the instruction window from which instruction-level parallelism (ILP) and memory-level parallelism (MLP) can be extracted. Often, the instruction window size is constrained by the size of the issue queue, because the larger the number of entries in the issue queue, the larger the pool of operations that can be considered when determining whether reordering of operations can be performed to strive to improve throughput.

[0006] However, the issue queue architecture is typically a critical speed path in processor design, and thus, increasing the size of the issue queue can lead to a decrease in the frequency at which the processor can be operated, which in itself will affect performance, and this generally limits the extent to which the issue queue capacity can be increased. Therefore, it would be desirable to provide an improved mechanism for operating an issue queue, with the goal of further improving processor performance. Summary of the Invention

[0007] In one example arrangement, an apparatus is provided that includes: an issue queue including a first section and a second section, each of the first section and the second section including a plurality of entries, and each entry being for storing operation information identifying an operation to be performed by a processing unit; an allocation circuit that receives operation information for a plurality of operations and applies an allocation criterion to determine for each operation whether to allocate the operation information of the operation to an entry in the first section or an entry in the second section, the operation information being arranged to identify each source operation object required by the associated operation and the availability of each source operation object; a selection circuit that selects, during a given selection iteration, an operation to be issued to the processing unit from the issue queue, the selection circuit being arranged to select the operation from among those operations for which the required source operation objects are available; an availability update circuit that updates the source operation object availability for each entry whose operation information identifies the target operation object of the selected operation in the given selection iteration as a source operation object; and a deferral mechanism that, during at least the next selection iteration after the given selection iteration, prohibits the selection circuit from selecting any operation associated with an entry in the second section whose required source operation object is now available because the operation uses the target operation object of the selected operation in the given selection iteration as a source operation object.

[0008] In another example arrangement, a method of operating an issue queue is provided that includes: arranging the issue queue to have a first section and a second section, each of the first section and the second section including a plurality of entries, and each entry being for storing operation information identifying an operation to be performed by a processing unit; receiving operation information for a plurality of operations and applying an allocation criterion to determine for each operation whether to allocate the operation information of the operation to an entry in the first section or an entry in the second section, the operation information being arranged to identify each source operation object required by the associated operation and the availability of each source operation object; selecting, during a given selection iteration, an operation to be issued to the processing unit from the issue queue, the selected operation being picked from among those operations for which the required source operation objects are available; updating the source operation object availability for each entry whose operation information identifies the target operation object of the selected operation in the given selection iteration as a source operation object; and employing a deferral mechanism to prohibit, during at least the next selection iteration after the given selection iteration, the selection of any operation associated with an entry in the second section whose required source operation object is now available because the operation uses the target operation object of the selected operation in the given selection iteration as a source operation object.

[0009] In another example arrangement, an apparatus is provided that includes: an issue queue apparatus including a first section and a second section, each of the first section and the second section including a number of entries, and each entry for storing operation information identifying an operation to be performed by a processing unit; an allocation apparatus for receiving operation information for a plurality of operations and for applying an allocation criterion to determine for each operation whether to allocate the operation information of the operation to an entry in the first section or an entry in the second section, the operation information being arranged to identify each source operation object required by the associated operation and the availability of each source operation object; a selection apparatus for selecting, during a given selection iteration, an operation to be issued to the processing unit from the issue queue apparatus, the selection apparatus for selecting the operation from among those operations for which the required source operation objects are available; an availability update apparatus for updating the source operation object availability for each entry such that the operation information of the entry identifies the target operation object of the selected operation in the given selection iteration as a source operation object; and a deferment apparatus for prohibiting, during at least the next selection iteration after the given selection iteration, the selection apparatus from selecting any operation associated with an entry in the second section for which the required source operation object is now available because the operation has the target operation object of the selected operation in the given selection iteration as a source operation object. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The present technique will be further described by way of illustration only with reference to examples of the present technique illustrated in the drawings, in which:

[0011] Figure 1 schematically illustrates a data processing apparatus according to one example;

[0012] Figure 2 is a block diagram illustrating the arrangement of an issue queue used in one example implementation;

[0013] Figure 3 is a flowchart illustrating the operation of an allocation circuit in one example implementation Figure 2 thereof;

[0014] Figure 4 is a flowchart illustrating how operation information may be migrated from an entry in the second section to the first section of the issue queue according to one example implementation;

[0015] Figure 5 is a flowchart illustrating, according to one example arrangement, Figure 2 the operation of a picker within a selection circuit thereof;

[0016] Figure 6 illustrates one implementation in which the picker circuit includes separate pickers for the first and second sections, and the arrangement of a target operation object determination circuit that may be used in one such implementation;

[0017] Figure 7 An example arrangement illustrates how the utilization of an issue queue can be changed in the presence of a latency-critical indication event; and

[0018] Figure 8 An example implementation schematically illustrates the fields that can be provided within each item of operation information stored within an entry of the issue queue. DETAILED DESCRIPTION

[0019] The logical functions for selecting and issuing operations from an issue queue form a scheduler loop, the timing of which can be affected as the size of the issue queue increases. Specifically, the scheduler loop includes the following functions:

[0020] a) Pick an operation to issue from all operations within the issue queue identified as ready to issue;

[0021] b) Multiplex out the picked operation;

[0022] c) Issue the selected operation to the execution pipeline;

[0023] d) Update source availability information for the (one or more) dependent operations of the issued operation (a dependent operation is an operation having a source operation object corresponding to the target operation object of the selected operation), and repeat.

[0024] The scheduler loop generally forms a critical loop from a timing perspective, and the size of the issue queue can affect the performance of the above steps (a), (b), (d) and is thus critical to the overall frequency capability of the scheduler function. The latency and bandwidth of the issue queue are also critical to performance, and thus it is important that the scheduler loop can be executed quickly, typically within a single clock cycle. Thus, while it is desirable to increase the capacity of the issue queue to increase the instruction window size as described above, the ability to increase the size of the issue queue is generally constrained by the requirement to efficiently execute the above scheduler loop. As will be discussed in more detail herein, the techniques described herein enable the effective capacity of the issue queue to be increased without adversely affecting the performance of the above scheduler loop.

[0025] Specifically, in one exemplary arrangement, an apparatus is provided that has an issue queue including a first section and a second section. Each of the first section and the second section includes a number of entries, and each entry is employed to store operation information identifying an operation to be performed by a processing unit. An allocation circuit is arranged to receive operation information for a plurality of operations and apply an allocation criterion to determine for each operation whether to allocate the operation's operation information to an entry in the first section or an entry in the second section. The operation information is arranged to identify each source operation object required by an associated operation and the availability of each source operation object.

[0026] A selection circuit is then used to select, during a given selection iteration, an operation to be issued to the processing unit from the issue queue, the selection circuit being arranged to select the operation from among those operations for which the required source operation objects are available.

[0027] An availability update circuit is then used to update the source operation object availability for each entry whose operation information identifies the destination operation object of the selected operation in the given selection iteration as a source operation object. There are several ways to identify source operation objects and destination operation objects, but they are typically identified with reference to the physical register values used to store data values. Thus, if the destination operation object is identified by a particular physical register and the same physical register is identified as a source operation object of an operation in one of the entries of the issue queue, then the source operation object availability indication for that source operation object can be updated to identify that the source operation object is now available because it is known that the selected operation will generate the required value.

[0028] Additionally, according to the techniques described herein, a deferral mechanism is used to prohibit the selection circuit from selecting, during at least the next selection iteration after a given selection iteration, any operation associated with an entry in the second section for which the required source operation object is now available because the operation uses the destination operation object of the selected operation in the given selection iteration as a source operation object.

[0029] With this scheme, the capacity of the issue queue is extended by providing a second section in addition to the first section, but due to the use of the deferral mechanism, the entries in the second section are removed from the critical timing path. This can increase the time available to perform one or more functions within the scheduler loop for the entries in the second section. Thus, this allows the size of the first section to be set such that the selection circuit and the availability update circuit can operate quickly enough for the entries in the first section to meet the required timing of the scheduler loop described above. However, since the entries in the second section are removed from the critical path by the use of the deferral mechanism, the selection circuit and the availability update circuit have more relaxed timing available when processing the entries in the second section. Thus, the capacity of the issue queue can be increased to increase the effective instruction window size without having an adverse effect on the frequency at which scheduling operations within the issue queue can be performed.

[0030] There are several ways to implement the deferral mechanism. However, in one example arrangement, the deferral mechanism is arranged to defer providing to the selection circuit updated source operand availabilities determined by the availability update circuit for any such operations associated with the entries in the second section: the required source operands for the operation are now available because the operation has the target operand of the selected operation in a given selection iteration as a source operand. This scheme relaxes the timing constraints on the availability update circuit when processing the entries in the second section because it does not need to update the source operand availability information for the entries in the second section in time for them to be considered by the selection circuit in the same selection iteration. Additionally, in a subsequent selection iteration in which the updated source operand availability information will be made available, it can be made available earlier in that selection iteration because it has been determined by the availability update circuit during a previous selection iteration.

[0031] The deferral mechanism can take many forms, but in one example implementation includes buffer storage. The buffer storage can take many forms and can be formed, for example, by latch circuits that latch the values on certain signal paths at the end of each clock cycle and can thus, for example, latch signals propagating on the path between the availability update circuit and the selection circuit such that the selection circuit receives the signals one clock cycle later than they are generated by the availability update circuit.

[0032] The allocation criteria applied by the allocation circuit to determine whether an item of received operation information is to be allocated to an entry in the first section or an entry in the second section can take various forms. However, in one example implementation, the allocation circuit is arranged to apply the following criteria as the allocation criteria: the criteria ensure an age ordering between the operations where the operation information is stored in entries in the first section and the operations where the operation information is stored in entries in the second section, such that all operations where the operation information is stored in entries in the first section are older than all operations where the operation information is stored in entries in the second section. Specifically, the fetched instructions will typically be decoded in the original program order, and thus the operation information will be received by the allocation circuit in the original program order. Once the operation information has been allocated to the issue queue, the operations can be issued out of order to the execution units of the processing unit to support OOO execution. Thus, the allocation circuit will have an implicit understanding of the relative ages of the items of received operation information, as these items will be received in age order. By applying the criteria that ensure all operations where the operation information is stored in entries in the first section are older than all operations where the operation information is stored in entries in the second section, this can assist the steps taken by the selection circuit. Specifically, the selection circuit can be arranged to preferentially select the older operations whose source operands are available, and thus in this case can be arranged to preferentially select operations from the first section rather than the second section.

[0033] When adopting the allocation criteria in the above form, the allocation circuit can be arranged to allocate the operation information to the first section when there are no occupied entries in the second section, but once the entries in the first section are full and it becomes necessary to start allocating the operation information to the entries in the second section, the allocation circuit needs to take this fact into account when allocating more items of the received operation information to ensure that all the operation information in the first section is older than all the operation information in the second section.

[0034] In one example arrangement, the allocation circuit is also arranged to migrate the operation information from the entries in the second section to the entries in the first section in order to maintain the age ordering. Thus, the allocation circuit can be arranged to migrate the operation information from the entries in the second section to the entries in the first section when there are available entries in the first section. Specifically, once at least one entry in the second section is occupied, the items of newly received operation information will not be able to be directly provided to the entries in the first section until the entries in the second section have been migrated to the entries in the first section.

[0035] In one example implementation, the allocation circuit is arranged to allocate received operation information to available entries in the first section when applying the allocation criteria and there are no active entries in the second section, where an active entry is an entry that stores operation information for an operation awaiting issue to a processing unit. Thus, when there are no active entries in the second section, items of received operation information can be directly allocated to entries in the first section, assuming available entries exist. However, the allocation circuit is also arranged to allocate received operation information to available entries in the second section when applying the allocation criteria and there is at least one active entry in the second section. By this scheme, this maintains the overall age ordering between the entries of the first section and the entries of the second section.

[0036] Although in the above implementation, an age ordering is maintained between the operation information stored in the entries of the first section and the operation information stored in the entries of the second section, age ordering constraints may or may not be applied between individual entries in any particular section. Thus, in one example implementation, age ordering constraints may be provided, but this may require significant movement of operation information between entries within a particular section. Accordingly, according to an alternative implementation, at least one of the first section and the second section is capable of storing allocated operation information to any available entry without being constrained by an age ordering between the entries in that section, and the apparatus is arranged to provide age ordered storage to identify an age order for the operation information stored in the entries of that section. Thus, information can be freely allocated within the entries of a particular section and a separate structure can be used to keep track of the relative ages of the items of operation information maintained within any particular entry.

[0037] In one example arrangement, the selection circuit is arranged to apply an age ordering criterion when selecting an operation from among those operations available from a required source operand, so as to preferentially select the oldest operation from among those operations available from the required source operand. It will be appreciated that although the selection circuit may make its decision primarily based on age ordering, the selection circuit may also take into account one or more other factors when deciding which operation to select, such as the availability of relevant functional units within the processing unit, the availability of result buses for propagating the results generated by the operations performed by the functional units, and so on.

[0038] When the selection circuit applies the above age ordering criterion, it will be appreciated that in an implementation where the allocation criteria ensure that all operations stored in the entries of the first section are older than all operations stored in the entries of the second section, this will mean that the selection circuit is arranged to preferentially select operations whose operation information is stored in the entries of the first section from among those operations available from the required source operand.

[0039] The selection circuit can be arranged in a number of ways. For example, a single selection mechanism can be arranged to review operation information for all operations available to a required source operation object, regardless of whether the operation information is stored in the first section or the second section, and then apply the age sorting criteria described above to determine which one to select. However, in a particular example implementation, the selection circuit includes separate picker circuits associated with the first and second sections. Specifically, in such an arrangement, the selection circuit can include a first picker for selecting a first candidate operation from among the operations for which the operation information is stored in the first section and which are available to the required source operation object, and a second picker for selecting a second candidate operation from among the operations for which the operation information is stored in the second section and which are available to the required source operation object. The final selection circuit is then used to pick the first candidate operation as the selected operation, unless no valid first candidate operation is available, in which case the final selection circuit is arranged to pick the second candidate operation as the selected operation.

[0040] By employing separate first and second pickers as described above, certain implementation benefits can be achieved. For example, as previously mentioned, since the deferral mechanism can defer providing the selection circuit with the updated source operation object availability determined by the availability update circuit for certain operations associated with entries in the second section, this means that the updated source operation object availability information is determined in a selection iteration prior to the selection iteration in which the second picker receives the operation information it uses to select the second candidate operation. Thus, during any given selection iteration, the second picker does not need to wait for the result of the operation performed by the availability update circuit.

[0041] Accordingly, in an example arrangement, the second picker is arranged to perform the selection of the second candidate operation in the next selection iteration before the availability update circuit has produced the updated source operation object availability, while the first picker is arranged to wait for the updated source operation object availability from the availability update circuit for any entry in the first section before performing the selection of the first candidate operation in the next selection iteration. Thus, the output of the second picker can be produced earlier, and the final selection circuit is then able to pick the selected operation as soon as the first picker has selected the first candidate operation.

[0042] The earlier availability of the output from the second selector can also result in other performance improvements within other components of the device. For example, the device may further include target determination circuitry to determine the target operand of the selected operation in each selection iteration. The target determination circuitry may include an initial evaluation circuit to determine the target operand for the second candidate operation and thereby exclude the target operand for any other operation whose operation information is stored in an entry in the second section, and a final evaluation circuit to determine the target operand for the selected operation when the final selection circuit has picked the selected operation, the final evaluation circuit ignoring any target operand excluded by the initial evaluation circuit. Thus, by the time the final evaluation circuit operates, several possible target operands have been excluded, thus improving the performance of the final evaluation circuit.

[0043] The initial and final evaluation circuits can be formed in a variety of ways. However, in one example implementation, the initial evaluation circuit can be formed as a first-stage multiplexing circuit to select a target operand for the second candidate operation from among the possible target operands of the entries in the second section. Similarly, the final evaluation circuit can be formed as a second-stage multiplexing circuit to select a target operand for the selected operation from among the possible target operands of the entries in the first section and the target operand of the second candidate operation output by the first-stage multiplexing circuit. Although the multiplexing circuits can be arranged in a variety of ways, in one example implementation, each multiplexing circuit is hardwired such that it receives target operand information from each entry in the relevant section, regardless of whether these entries store valid information, and regardless of whether the operations in these entries are still selectable, and thus regardless of whether their required source operands are available. This means that the multiplexing circuits do not need to be reconfigured during each selection iteration, and the operation of the selection circuit ensures that the selected operation is an operation ready to be issued. Since the initial evaluation circuit can perform its multiplexing function before the final evaluation circuit is able to perform its multiplexing function, it can be seen that the overall performance of the target determination circuitry can be improved because the use of the initial evaluation circuit reduces the size of the required final evaluation circuit. Specifically, the second-stage multiplexing circuit implementing the final evaluation circuit will have fewer inputs compared to the case where the initial evaluation circuit is not used.

[0044] In one example implementation, the issue queue includes a number of initial entries, and the received operation information is initially stored in these initial entries before the allocation circuit determines whether to allocate the received operation information to an entry in the first section or an entry in the second section. This allows items of the received operation information to be buffered before the allocation circuit determines whether the information should be stored in an entry in the first section or an entry in the second section, and thus the use of a certain number of initial entries can reduce the timing constraints that would otherwise be imposed on the allocation operation. The number of initial entries is a matter of design choice, for example depending on the number of write ports provided to the issue queue. For example, if two write ports are provided to the issue queue, it may be considered appropriate to provide two initial entries, and the items of operation information that can be received in a single cycle can be stored in these two initial entries.

[0045] In one example implementation, the selection circuit is also capable of selecting from the initial entries in some cases. Specifically, the selection circuit can be arranged to select an operation from the initial entries when the required source operand of the operation is available and no entry in the first and second sections is available for storing the operation information of the operation for which the required source operand is available.

[0046] If desired, the use of the first and second sections can be made configurable. For example, in response to at least one latency-critical indication event, the issue queue can be arranged to prohibit the use of the second section. Specifically, in a situation where the latency of an operation is determined to be critical, it may be considered inappropriate to allocate entries to the second section because it is known that there will be at least a one-cycle delay when these operations are awakened by the action of the availability update circuit (assuming these operations are still in the entries of the second section at this stage). Instead, it may be considered better to operate the issue queue with a reduced size in this situation.

[0047] As another example of why it may be desirable to make the use of the first and second sections configurable, the selective prohibition of the use of the second section can be arranged to occur when it is detected that there is little or no parallelism available during the execution of operations. In cases where there is more parallelism, a deeper queue is beneficial and thus enabling a slower second section is beneficial. However, in cases where there is little or no parallelism, the power consumed in moving operations into and out of the slower section may be inappropriate because they may not be picked until later after they move to the faster first section. The movement of operations can thus consume power unnecessarily when instead they could just be deferred. Therefore, the parts of the instruction stream that would not benefit from a deeper queue can be identified, and the slower section can then be disabled during the execution of these parts to save power.

[0048] Specific examples will now be described with reference to the accompanying drawings.

[0049] Figure 1 An example of a data processing apparatus 2 having a processing pipeline including a number of pipeline stages is schematically illustrated. The pipeline includes a branch predictor 4 for predicting the outcome of a branch instruction and generating a series of fetch addresses for the instructions to be fetched. A fetch stage 6 fetches the instructions identified by the fetch addresses from an instruction cache 8. A decode stage 10 decodes the fetched instructions to generate control information for controlling subsequent stages of the pipeline. An out-of-order processing component 12 is provided in the next stage to handle out-of-order execution of instructions. These components can take various forms, for example including a reorder buffer (ROB) and a renaming circuit. The ROB is used to keep track of the progress of instructions and ensure that instructions are committed in order, although they are executed out of order. A rename stage 12 performs register renaming to map the architectural register specifiers identified by the instructions to physical register specifiers identifying registers 14 provided in the hardware. Register renaming can be useful for supporting out-of-order execution, as this can allow hazards between instructions to be eliminated by mapping instructions that specify the same architectural register to different physical registers in the hardware register file, to increase the likelihood that instructions can be executed in an order different from their program order in which they were fetched from the cache 8, which can improve performance by allowing later instructions to execute while earlier instructions are waiting for operands to become available. The ability to map architectural registers to different physical registers can also facilitate the rollback of architectural state in the case of a branch misprediction. An issue stage 16 queues the instructions waiting to be executed until the operands required to process those instructions are available in the registers 14. An execution stage 18 executes the instructions to effect the corresponding processing operations. A write-back stage 20 writes the results of the executed instructions back to the registers 14.

[0050] The execution stage 18 can include a number of execution units, such as a branch unit 21 for evaluating whether a branch instruction has been correctly predicted, an ALU (arithmetic logic unit) 22 for performing arithmetic or logical operations, a floating-point unit 24 for performing operations with floating-point operands, and a load / store unit 26 for performing load operations to load data from the memory system into the registers 14 or store operations to store data from the registers 14 to the memory system. In this example, the memory system includes a first-level instruction cache 8, a first-level data cache 30, a second-level cache 32 shared between data and instructions, and a main memory 34, but it will be appreciated that this is just one example of a possible memory hierarchy, and other implementations may have more levels of cache or different arrangements. The load / store unit 26 can use a translation lookaside buffer 36 and the fetch unit 6 can use a translation lookaside buffer 37 to map the virtual addresses generated by the pipeline to physical addresses identifying locations within the memory system. It will be appreciated that, Figure 1The pipeline shown is only an example, and other examples may have different sets of pipeline stages or execution units.

[0051] As previously described, the techniques described herein specifically relate to the operation of the issue queue, and in particular provide a mechanism that enables the effective size of the issue queue to be increased without reducing the operating frequency, while still enabling the scheduler loop to be executed at a desired rate, typically within a single clock cycle (i.e., enabling the selection iteration to occur every clock cycle, if desired). By increasing the effective size of the issue queue, this can increase the instruction window size, thereby improving the performance of the processor. However, it should be noted that increasing the number of entries in the issue queue is not necessary. For example, for a particular number of entries that make up the issue queue, then by employing the techniques described herein, it may be possible to increase the operating frequency because the scheduler loop can be executed more quickly, and this in turn will also improve performance. As another example, the techniques described herein can be used to reduce power consumption by having fewer fast (and thus more power-consuming) entries. Thus, the techniques described herein can be used to obtain frequency, performance, or power benefits, or a combination of these.

[0052] Figure 2 is a more detailed illustration of the components provided in association with the issue queue provided in Figure 1 the issue stage 16 of. As Figure 2 shown, the issue queue 100 includes a first section 112 and a second section 108, where in one example implementation, each section includes a plurality of entries, and where each entry can be used to store operation information for an operation to be executed by one of the execution units 21, 22, 24, 26 within the execution stage 18. In some implementations, a separate issue queue 100 can be provided for each of the execution units 21, 22, 24, 26, but in alternative implementations, a single issue queue 100 can be provided for a plurality of the execution units 21, 22, 24, 26.

[0053] Here, the entries within the first section 112 will be referred to as fast entries, while the entries within the second section 108 will be referred to as slow entries. From the following discussion, it will become clear that the terms fast and slow for these entries are a description of the behavior of the entry with respect to wake-up events from a performance perspective. Specifically, when all of the required source operands for the operation stored in a particular entry become available, then the entry is considered to be woken up because it then becomes a candidate entry from which the selection circuit 130 can select an operation to be issued to the execution stage 18. From the discussion here, it will become clear that after such a wake-up event for an entry in the second section, there is a delay in terms of the associated operation becoming available for the selection circuit 130 to select as an issued operation.

[0054] The issue queue may include a number of write ports and is shown in Figure 2 as having two write ports such that the operation information of two operations can be received into the issue queue simultaneously. The allocation circuit 120 is used to determine whether to allocate an item of newly received operation information to an entry in the first section or an entry in the second section, and the details of the process employed by the allocation circuit in one example implementation will be referred to later in Figure 3 connection with the discussion of the details of the process employed by the allocation circuit in one example implementation. Assuming that the allocation circuit can perform its allocation analysis quickly enough, an item of received operation information can be immediately directed to the relevant entry in either the first section 112 or the second section 108. However, as Figure 2 shown in

[0055] one example, the issue queue 100 includes a number of initial entries 106, and in this particular example two initial entries 102, 104 are provided, i.e., one initial entry for each write port. Thus, each new item of operation information is initially buffered within one of the initial entries 102, 104 and is routed from there to either the first section 112 or the second section 108, depending on the analysis performed by the allocation circuit 120. Figure 3 The details of the allocation process performed to achieve this age ordering in one particular implementation will be referred to later in

[0056] However, in Figure 2 the example implementation shown in

[0057] The operation information is received by the issue queue 100 from the rename stage 12 and will identify the operation to be performed and the source operand objects required for that operation. It will typically also identify the destination operand object, such as the physical register to which the result should be written. While some source operand objects may be immediate values, it is often the case that one or more of the source operand objects are specified with reference to one of the physical registers, and the logic 115 can be used to perform an initial determination of the availability of such source operand objects and provide that information as part of the operation information received by the issue queue. Thus, at the time of initial allocation, one or more of the possible source operand objects may already be available for the associated operation, and this availability information can be captured within the operation information. However, there may be one or more source operand objects that are not yet available, and thus the availability information associated with such source operand objects will indicate that fact.

[0058] As previously described, an operation can only become a candidate for selection by the selection circuit 130 when its source operand objects are considered available, and for any operation that is initially assigned to the issue queue for which not all of its source operand objects are available, the source operand object availability of the associated entry will need to be updated to account for operations subsequently issued by the selection circuit. This function can be performed by Figure 2 the availability update circuit 155 shown in the following process which will be discussed in more detail below.

[0059] Figure 8 Schematically illustrates the fields that can be provided associated with each item of the operation information 180 stored within an entry of the issue queue 100. The first field 182 is used to identify the operation associated with the entry, while the field 184 is used to identify the various source operand objects. In the case where an immediate value is provided as one of the source operand objects, the immediate value can be provided within this field, and for operand object values specified with reference to registers, the register identifier can be used within this field to identify those source operand objects. The field 184 can be considered to consist of a number of sub-fields, one sub-field for each source operand object.

[0060] The field 186 provides source operand object availability information to identify whether the required source operand objects are available. In one example implementation, the source operand object availability field 186 can be considered to include a number of sub-fields, where each sub-field provides a status flag for the associated source operand object. Only when the status flags indicate that all of the source operand objects are available will the operation identified by the operation information 180 be considered a candidate for selection by the selection circuit 130.

[0061] As Figure 8Also shown therein, the target operand field 188 is used to provide an indication of the target operand associated with the operation. This information is extracted by the target operand determination circuit 145 for the selected operation so that it can be forwarded to the availability update circuit to update the source operand availability information of the compliant operations stored in the issue queue.

[0062] As previously described, a series of logic functions are to be performed to implement the scheduler loop for the issue queue, and these functions form the critical loop. This scheduler loop is shown in Figure 2 by the path from the selection circuit 130 through the target operand determination circuit 145 and via the latch 150 to the availability update circuit 155 which in turn provides its output to the selection circuit 130. In the Figure 2 example shown, at the end of the current selection iteration (i.e., the current iteration of the scheduler loop), the target operand of the operation selected to be issued is latched in the register 150. In the next selection iteration, this value is output to the availability update circuit 155 so that the source operand availability information of the active entries in the issue queue can be updated, and this updated information is forwarded to the selection circuit 130 so that the next operation can be selected and issued, and the target operand determination circuit 145 then determines the target of the selected operation so that it can be stored in the register 150 at the end of this next selection iteration. Generally, it is desired that this entire iteration of the loop occur in a single cycle, and it has been found that this scheduler loop becomes the critical loop when determining the overall frequency at which the specified device can operate so that all required functions can be performed within a single clock cycle.

[0063] As Figure 2 shown, this single-cycle scheduling still executes for the fast entries, i.e., the entries within the first section 112, but the deferral mechanism 170 can be used to decouple the slow entries from the critical path of the scheduler loop. Thus, at the start of the current selection iteration (i.e., the iteration of the scheduler loop), the target operand latched in the register 150 (i.e., the target of the operation issued in the previous selection iteration) is broadcast to the components forming the availability update circuit 155. In Figure 2 this, this target operation is referred to as the PTAG value. Specifically, in an example implementation, renaming is used within stage 12 of the device to map the architecturally specified registers of the instructions to the physical registers within the register set 14, and these physical registers are identified by the PTAG values.

[0064] As Figure 2As shown, the availability update circuit 155 includes a lookup / update circuit 160 for fast entries (i.e., entries in the first section 112) and a lookup / update circuit 165 for slow entries (i.e., entries in the second section 108). The fast entry update process is on the critical path, and thus the lookup / update circuit 160 for fast entries needs to perform a lookup operation for the entries in the first section 112 to update the source operand availability for any compliant entries (i.e., any entries that will store operation information using the target operand broadcast from register 150 as the source operand), and then pass the updated information to the selection circuit so that the remaining part of the scheduler loop can be executed in the same clock cycle. The lookup / update circuit can be arranged in various ways, but in one example implementation, a content addressable memory (CAM) lookup process is used to access and update the various entries in the first section. The updated availability information is then forwarded to the selection circuit 130, and specifically to the readiness determination circuit 135 within the selection circuit.

[0065] However, as Figure 2 shown, the lookup / update circuit 165 for slow entries is not on the critical path because it does not need to produce its output in time for it to be available to the selection circuit 130 during the same selection iteration. Instead, the updated availability information from the lookup / update circuit 165 for slow entries is passed to a deferral mechanism 170, which takes the form of a buffer circuit in one example. Specifically, the buffer circuit can be used to defer the provision of the updated availability information to the selection circuit 130 by at least one clock cycle. In a particular implementation, latches are used to implement the buffer circuit, resulting in a single cycle delay in the provision of the updated availability information to the selection circuit. This means that there is much looser timing for the operations of the lookup / update circuit 165 for slow entries. Similarly, a CAM lookup process can be used to access the entries in the second section 108, and all that is required is to perform this in time so that the updated availability information is buffered within the buffer 170 at the end of the current selection iteration.

[0066] Therefore, through this process, it will be understood that the information provided to the readiness determination circuit 135 is delayed by one cycle with respect to the entries in the second section 108, and thus provides an indication of the source availability of the operations identified by the entries in the second section in the cycle before the current selection iteration being considered by the selection circuit 130.

[0067] The readiness determination circuit 135 analyzes the source operand availability information provided by the availability update circuit 155, particularly the information provided by the lookup / update circuit 160 for fast entries and the information about slow entries forwarded from the buffer 170, and determines which entries identify operations that are candidates for selection during the current selection iteration. As previously mentioned, for an operation to be a candidate for selection, all of its source operands must be available. Additionally, the readiness determination circuit may consider certain other criteria, such as the availability of any other shared resources, and any other prerequisites.

[0068] Based on the analysis performed by the readiness determination circuit 135, the picker circuit 140 is provided with an indication of any operations available for selection, and then applies selection criteria in order to select one of these operations as the next operation to issue. For example, the picker circuit may apply an age sorting criterion in order to attempt to select the oldest operation among those indicated by the readiness determination circuit 135 as being available for selection.

[0069] In addition to an operation being issued by the picker circuit 140 to the execution stage 18, an indication of the selected operation is also provided to the destination operand determination circuit 145, which then determines the destination operand for the selected operation. Specifically, from the foregoing Figure 8 it will be apparent that the operation information 180 will include a field 188 that provides an indication of the destination operand, and the destination operand determination circuit 145 is arranged to capture the information in this field and cause it to be latched in the register 150 at the end of the current selection iteration. The foregoing process may then be repeated in the next selection iteration in order to select the next operation to issue.

[0070] As previously mentioned, a single issue queue may be provided for all execution units in the execution stage 18, or a separate issue queue may be provided for each such execution unit. In the case of using a single issue queue, the components in the scheduler loop may be modified so that more than one instruction may be issued in each selection iteration, for example, one operation may be issued to one execution unit, another operation may be issued to another execution unit, and so on. In an implementation where a separate issue queue is maintained for each execution unit, then in one example implementation, the Figure 2 circuitry may be replicated for each issue queue, where the selection circuit 130 picks at most a single operation to issue during any particular selection iteration, and the selected operation is routed to the associated execution unit within the execution stage 18.

[0071] Figure 3 is illustrated in one example implementation Figure 2Flowchart of the operation of the distribution circuit 120. Specifically, the distribution circuit 120 monitors the information in the initial entries 102, 104 received by the issue queue 100. When it is determined in step 200 that new operation information has been received, it is determined in step 205 whether there are any active entries in the second section, that is, whether the second section contains any pending operations waiting to be issued. If not, it is determined in step 210 whether there are available entries in the first section, that is, whether there are idle entries that do not currently store information about pending operations. If so, in step 215, the operation information can be assigned to the available entry in the first section.

[0072] However, if it is determined in step 205 that there is at least one active entry in the second section, the process proceeds to step 220, where it is determined whether there are available entries in the second section, that is, whether the second section is not yet full. Assuming there are available entries, the process proceeds to step 225, where the operation is assigned to the available entry in the second section. If it is determined in step 210 that there are no available entries in the first section, that is, the first section is currently full, the process also proceeds to step 225.

[0073] In addition, if it is determined in step 220 that there are no available entries in the second section, the process returns to step 225, and in particular at this point, the new operation information remains in the initial entries until the time when it can be moved to the first section or the second section. This may for example mean that the issue queue cannot accept new operation information during the next cycle until at least one of the initial entries is available.

[0074] By adopting the Figure 3 scheme shown in, it will be appreciated that the age ordering is maintained between the entries in the first section and the entries in the second section, and in fact for the entries in the initial section 106 in the following sense: the operation information needs to be held therein rather than being immediately moved to the first section or the second section.

[0075] When the operation information is assigned to the entries in the first section or the second section, the associated age matrices 127, 125 will be updated to keep track of the relative ages of the operations stored in the various entries of the associated section. This information can be provided to the selection circuit 130 so that the picker circuit 140 can apply the age ordering criterion when selecting the operations to be issued.

[0076] In one example implementation, the distribution circuit 120 is also used to control the movement of operation information from the entries in the second section to the entries in the first section. Specifically, from Figure 3It will be appreciated that operation information is assigned to the second section only because there is already operation information in the active entries of the second section or because the first section is full. As space becomes available within the first section, the allocation circuitry endeavors to migrate operation information from the second section to the first section in order to maintain the age ordering between the sections and to free up space so that new items of operation information can be moved out of the initial entries into the first or second section.

[0077] Figure 4 A process is shown that can be used to control the movement of operation information from the second section 108 to the first section 112. At step 250, it is determined whether there are any active entries in the second section, and if not, no action is required. However, whenever there is at least one active entry in the second section, then at step 255 it is determined whether there is at least one available entry in the first section. Assuming there is at least one available entry in the first section, then at step 260 at least the oldest operation information among the active entries of the second section is moved to the available entry in the first section.

[0078] Depending on the number of write ports provided and the number of available entries in the first section, it is possible to move more than one item of operation information in a particular clock cycle. In Figure 2 the example shown, there are two write ports and thus two information items from the second section can be moved to the first section in a single cycle.

[0079] By using the processes discussed above with reference to Figure 3 and Figure 4 it will be appreciated that operations placed within slow entries will migrate down to fast entries as availability permits, unless they are picked by the picker 140 and pre-deallocated. It should also be noted that the movement between the second section and the first section, and the movement from the initial section 106 to the relevant parties in the first and second sections, can be optimized in the described implementation so that the required movement is minimized. Specifically, in the absence of active entries in the second section, operation information will be moved directly from the initial entry to the first section. Since multiple entries can be moved in a single cycle, it will be appreciated that the various movements can be optimized to strive to minimize the required movement while still ensuring the age ordering. Thus, as a very specific example, if two new items of operation information are received in the initial entries 102, 104 and there is an existing active entry in the second section 108, then the operation information in the second section 108 can be moved to a fast entry in the first section, the older operation in the initial entry can be moved directly to a fast entry in the first section, and the younger of the two entries in the initial entry can be moved to a slow entry in the second section 108.

[0080] Figure 5is a flowchart showing the operation of the picker circuit 140 in an example implementation. At step 300, it is determined whether there is at least one active entry of an operation that identifies a source operation object in the first section. If so, then at step 305, the picker circuit 140 selects an operation to issue from the first section. As previously described, in the case where there are more than one such entries in the first section, the age sorting information from the age matrix 127 can be considered by the picker to select the oldest one of these operations to issue. Figure 2 If no operation selection is available from the first section, the process proceeds to step 310, where it is determined whether there is at least one active entry of an operation that identifies a source operation object in the second section. As previously described, due to the deferral mechanism 170, there is at least a one-cycle delay in providing updated source availability information to the picker 140 for the entries in the second section 108, and thus during the current selection iteration, the picker is considering the availability that existed for the second section entries during the previous selection iteration.

[0081] If there is at least one entry of a source operation object available in the second section, the process proceeds to step 315, where an operation is selected from the second section 108 to issue to the execution unit. Similarly, if there are more than one available operations in the second section, the age sorting information from the age matrix 125 can be considered by the picker to strive to select the oldest one of these operations to issue.

[0082] If it is determined at step 310 that there are no available entries for selection in the second section, then at step 320 it is determined whether there is at least one initial entry of an operation that identifies a source operation object, and if so, at step 325 an operation is selected from the initial section 106. In the case where both entries in the initial section 106 contain pickable operations, then at step 325 the oldest one of these operations will be selected. If it is determined at step 320 that there are no selectable operations in the initial section 106, then at step 330 it is determined not to select an operation during the current selection iteration.

[0083]

[0084] Figure 6 ​Illustrated is an example implementation where the picker circuit 140 is actually formed by a number of individual picker circuits, including at least a first picker 400 and a second picker 410, and in this case including a third picker 412. The first picker 400 selects among any ready operations identified by entries in the first section, while the second picker 410 selects among any ready operations identified by entries in the second section. If the first picker 400 is able to select a first candidate operation, this will take precedence over a second candidate operation selected by the second picker, and this functionality is implemented by a final selection circuit 425 which, in one implementation, may take the form of a multiplexer that selects the first candidate operation if a valid first candidate operation is selected, or otherwise selects the second candidate operation if a valid second candidate operation is selected by the second picker. The third picker 412 may be used to select an operation from an initial section if at least one of the initial entries identifies an operation available to the source action object. In the example described in reference Figure 2 there are two initial entries sorted by age (i.e., the content of one is always older than the content of the other), so if the older operation is ready it is picked, otherwise if the younger operation is ready it is picked. If neither the first picker 400 nor the second picker 410 indicates a valid operation ready to be selected, the multiplexer 425 may select the operation indicated by the third picker 412.

[0085] By employing a separate second picker, certain implementation benefits can be achieved. Specifically, from Figure 2 it will be clear that information about the availability of the source action object for entries in the second section is directly available from the buffer 170 during a given selection iteration, without having to wait for the result of an operation performed by the availability update circuit, and thus this second picker 410 can be activated before the first picker 400 because the first picker needs to wait for updated availability information from the lookup / update circuit 160 for the fast entries. The second picker 410 can also be activated before the third picker 412.

[0086] This can also result in downstream performance improvements in the operation of the target action object determination circuit 145, as Figure 6as shown in the specific example implementation method in. Specifically, in this implementation method, the target operation object determination circuit includes an initial multiplexer 415 to select between various target operation object identifiers of the slow purpose based on the second candidate operation generated by the second picker. This results in a single signal being output as the input to the final multiplexer 420, and the final multiplexer also receives the target operation object identifiers for all fast entries and initial entries. Therefore, by the time the output from the final selection circuit 425 is available, the final multiplexer 420 can be directly driven by this output signal to output the target operation object identifier. Thus, compared with the case where the final multiplexer has to receive inputs for each entry in both the first section and the second section, the final multiplexer can be made smaller.

[0087] Specifically, in Figure 6 In the example shown, the total number of inputs to the final multiplexer 420 can be set to be equal to the number of entries in the first section, plus the number of entries in the initial section 106, plus one input for all entries in the second section, because the initial multiplexer 415 will have completed initial filtering for these second section entries to generate only a single target operation object identifier for the entries in the second section. Since the size of the final multiplexer is thus reduced, this can improve the speed of operation of the target operation object determination circuit 145.

[0088] In one example implementation, the use of the first and second sections can be made configurable. Specifically, in a normal usage scenario, both the first and second sections will be used, and the issue queue will operate in the manner described earlier. However, when one or more identified events occur, it can be decided to disable the second section, thereby reducing the effective size of the issue queue. This can be used, for example, in certain latency-critical scenarios. Figure 7 is a flowchart illustrating the process that can be adopted to implement this function. Thus, at step 500, it is determined whether a latency-critical indication event has occurred, and when such an event is detected, the process proceeds to step 505, in which any active entries in the second section are migrated to the first section as space becomes available in the first section, and at the same time, any further allocation to the second section is prevented.

[0089] As mentioned earlier, instead of making the selective disabling of the second section based on the detection of a latency-critical event, the selective disabling of the use of the second section can be arranged to occur when it is detected that there is little or no parallelism available during the execution of an operation, and thus step 500 will be the detection of low levels of parallelism available in this case.

[0090] After step 505, then at step 510, when there are no longer any active entries in the second section, the second section is disabled. Then, it will be appreciated that new items of operation information can still be received in the initial entry 106, but once there is space available in the first section, they will be directly migrated down to the first section 112. The lookup / update circuitry 165 for the slowpoke purpose is thus no longer necessary when in this reduced-size mode of operation, but the main timing-critical scheduler path continues to operate in the same manner as discussed earlier with reference to the fast entries.

[0091] At step 515, it is determined whether there is a latency-critical end event, and when it is determined that such an event has occurred, then at step 520 the second section can be re-enabled.

[0092] According to the above example implementation, it will be appreciated that the techniques described herein enable the effective capacity of the issue queue to be increased, thereby increasing the instruction window size and enabling an increase in the performance of the OOO processor. Additionally, this increase in size can be achieved without adversely affecting the timing of the scheduler loop, which is a timing-critical function within the processor. As an alternative to increasing the size of the issue queue, the size of the issue queue can remain the same, but the scheduler function will be able to be executed at a higher frequency, and thus an increase in performance can be achieved in that manner if desired, rather than increasing the overall size of the issue queue.

[0093] As described herein, the movement between the second and first sections of the issue queue can be managed to maintain the desired age ordering, and this movement can be performed independently of any picking and deallocation of operations.

[0094] In this application, the phrase "configured to" is used to mean that an element of a device has a configuration capable of performing a defined operation. In this context, "configuration" refers to the arrangement or interconnection of hardware or software. For example, a device can have dedicated hardware that provides the defined operation, or a processor or other processing device can be programmed to perform the function. "Configured to" does not mean that the device element needs to be changed in any way in order to provide the defined operation.

[0095] Although the illustrative embodiments of the present invention have been described in detail herein with reference to the accompanying drawings, it is to be understood that the present invention is not limited to these exact embodiments, and various changes, additions, and modifications can be made by those skilled in the art without departing from the scope and spirit of the present invention as defined by the appended claims. For example, the features of the dependent claims can be combined with the features of the independent claims in various ways without departing from the scope of the present invention.

Claims

1. An apparatus for operating an issue queue, comprising: An issue queue including a first section and a second section, each of the first section and the second section including a number of entries, and each entry being used to store operation information identifying an operation to be performed by a processing unit; An allocation circuit that receives operation information for a plurality of operations and applies an allocation criterion to determine for each operation whether to allocate the operation information of the operation to an entry in the first section or an entry in the second section, the operation information being arranged to identify each source operation object required by the associated operation and the availability of each source operation object; A selection circuit that selects, during a given selection iteration, an operation to be issued to the processing unit from the issue queue, the selection circuit being arranged to select the operation from among those operations for which the required source operation objects are available; An availability update circuit that updates the source operation object availability for each entry such that the operation information of the entry identifies the target operation object of the selected operation in the given selection iteration as a source operation object; And A postponement mechanism that prohibits the selection circuit from selecting, during at least the next selection iteration after the given selection iteration, any operation associated with an entry in the second section for which the required source operation object is now available because the operation uses the target operation object of the selected operation in the given selection iteration as a source operation object.

2. The apparatus as claimed in claim 1, wherein: The postponement mechanism is arranged to postpone providing the selection circuit with the updated source operation object availability determined by the availability update circuit for any operation associated with an entry in the second section for which the required source operation object is now available because the operation uses the target operation object of the selected operation in the given selection iteration as a source operation object.

3. The apparatus as claimed in claim 2, wherein the deferral mechanism includes buffer storage.

4. The apparatus as claimed in claim 1, wherein: The allocation circuit is arranged to apply the following criterion as the allocation criterion: the criterion ensures an age ordering between operations whose operation information is stored in entries in the first section and operations whose operation information is stored in entries in the second section, such that all operations whose operation information is stored in entries in the first section are older than all operations whose operation information is stored in entries in the second section.

5. The apparatus as claimed in claim 4, wherein: The allocation circuit is further arranged to migrate operation information from entries in the second section to entries in the first section in order to maintain the age ordering.

6. The apparatus as claimed in claim 4, wherein the allocation circuit is arranged to migrate operation information from an entry in the second section to an entry in the first section when there is an available entry in the first section.

7. The apparatus as claimed in claim 4, wherein: The allocation circuit is arranged to allocate the received operation information to available entries in the first section when applying the allocation criterion and the second section has no active entries, where an active entry is an entry that stores operation information for an operation waiting to be issued to the processing unit.

8. The apparatus as claimed in claim 7, wherein: The allocation circuit is arranged to allocate the received operation information to available entries in the second section when applying the allocation criterion and the second section has at least one active entry.

9. The apparatus as claimed in claim 4, wherein at least one of the first section and the second section is capable of storing allocated operation information into any available entry without being constrained by the age ordering among the entries in that section, and the apparatus is arranged to provide age-ordered storage to identify the age order of the operation information stored in the entries of that section.

10. The apparatus as claimed in claim 1, wherein: The selection circuit is arranged to apply an age ordering criterion when selecting an operation from among those operations for which the required source operation objects are available, so as to preferentially select the oldest operation from among those operations for which the required source operation objects are available.

11. The apparatus as claimed in claim 4, wherein: The selection circuit is arranged to preferentially select, from among the operations available for the required source operation object, the operations whose operation information is stored in the entries of the first section.

12. The apparatus as claimed in claim 11, wherein the selection circuit includes: A first picker for selecting a first candidate operation from among the operations whose operation information is stored in the entries of the first section and which are available for the required source operation object; A second picker for selecting a second candidate operation from among the operations whose operation information is stored in the entries of the second section and which are available for the required source operation object; And A final selection circuit that picks the first candidate operation as the selected operation, unless no valid first candidate operation is available, in which case the final selection circuit is arranged to pick the second candidate operation as the selected operation.

13. The device as claimed in claim 12, wherein: The second picker is arranged to perform the selection of the second candidate operation in the next selection iteration before the availability update circuit has generated the updated source operation object availability, while the first picker is arranged to wait for the updated source operation object availability from the availability update circuit for any entry in the first section before performing the selection of the first candidate operation in the next selection iteration.

14. The device as claimed in claim 13, further comprising: A target determination circuit that determines the target operation object of the selected operation in each selection iteration; Wherein the target determination circuit includes: An initial evaluation circuit that determines the target operation object for the second candidate operation and thereby excludes the target operation object for any other operation whose operation information is stored in the entries of the second section; And A final evaluation circuit that determines the target operation object for the selected operation when the final selection circuit has picked the selected operation, the final evaluation circuit ignoring any target operation object excluded by the initial evaluation circuit.

15. The device as claimed in claim 14, wherein: The initial evaluation circuit is formed as a first-stage multiplexing circuit to select the target operation object for the second candidate operation from among the possible target operation objects of the entries in the second section; and The final evaluation circuit is formed as a second-stage multiplexing circuit to select the target operation object for the selected operation from among the possible target operation objects of the entries in the first section and the target operation object of the second candidate operation output by the first-stage multiplexing circuit.

16. The device as claimed in claim 1, wherein: The issue queue includes a number of initial entries, and the received operation information is initially stored in the initial entries before the allocation circuit determines whether to allocate the received operation information to an entry in the first section or an entry in the second section.

17. The device as claimed in claim 16, wherein the selection circuit is arranged to select an operation from the initial entries when the following occurs: the required source operation object for the operation is available and no entry in the first section and the second section stores operation information for an operation for which the required source operation object is available.

18. The device as claimed in claim 1, wherein, In response to at least one latency-critical indication event, the issue queue is arranged to prohibit the use of the second section.

19. A method for operating an issue queue, comprising: The issue queue is arranged to have a first section and a second section, each of the first section and the second section including a number of entries, and each entry being used to store operation information identifying an operation to be executed by a processing unit; Receive operation information for a plurality of operations, and apply an allocation criterion to determine for each operation whether to allocate the operation information of the operation to an entry in the first section or an entry in the second section, the operation information being arranged to identify each source operation object required by the associated operation and the availability of each source operation object; Select, during a given selection iteration, an operation to be issued to the processing unit from the issue queue, the selected operation being selected from among those operations for which the required source operation objects are available; Update the availability of the source operation objects for each entry whose operation information identifies the target operation object of the selected operation in the given selection iteration as a source operation object; And Adopt a deferral mechanism to prohibit, during at least the next selection iteration after the given selection iteration, the selection of any operation associated with an entry in the second section whose required source operation object is now available because the operation uses the target operation object of the selected operation in the given selection iteration as a source operation object.

20. A device for operating an issue queue, comprising: An issue queue device including a first section and a second section, each of the first section and the second section including a plurality of entries, and each entry being for storing operation information identifying an operation to be executed by a processing unit; An allocation device for receiving operation information for a plurality of operations, and for applying an allocation criterion to determine for each operation whether to allocate the operation information of the operation to an entry in the first section or an entry in the second section, the operation information being arranged to identify each source operation object required by the associated operation and the availability of each source operation object; A selection device for selecting, during a given selection iteration, an operation to be issued to the processing unit from the issue queue device, the selection device being for selecting the operation from among those operations for which the required source operation objects are available; An availability update device for updating the availability of the source operation objects for each entry whose operation information identifies the target operation object of the selected operation in the given selection iteration as a source operation object; And A deferral device for prohibiting, during at least the next selection iteration after the given selection iteration, the selection device from selecting any operation associated with an entry in the second section whose required source operation object is now available because the operation uses the target operation object of the selected operation in the given selection iteration as a source operation object.

Citation Information

Patent Citations

  • System and method for load and store queue allocations at address generation time

    CN109564510A

  • Integrated Scheduling of General Messages and Time-Critical Messages

    US20170054828A1